Articles

Towards Nursing Activity Recognition Using Large Language Models Based on Visual Information

Kodai IWAMOTO, Sho MITARAI, Goshiro YAMAMOTO, Chang LIU, Kazumasa KISHIMOTO, Yukiko MORI, Tomohiro KURODA
Vol. 15 (2026) p. 425-433

Many ongoing initiatives in medical digital transformation use historical data stored in electronic health records (EHRs). However, real-time support is essential because clinical situations can change rapidly and unpredictably. This study investigated the potential of recognizing nursing activities from visual information using an inference task to identify corresponding nursing record names, in which a multimodal large language model (MLLM) and expert nurses analyzed videos. The primary outcome was the accuracy of the MLLM under a visual-only condition, with nurse judgments and audio-enhanced MLLM outputs used as complementary references. Scenario videos were generated by sampling the most frequently documented nursing records from hospital databases and reproducing situations that closely resembled real clinical settings. Under the visual-only condition, the approximate accuracy rates were 32% for the MLLM and 85% for nurses, but the MLLM accuracy improved to 85% when audio information was added. Based on these results, we grouped nursing record names into three patterns according to the ease of inference from image information, and summarized the characteristics and challenges of each pattern. The inferable records were grouped into two categories: those requiring professional interpretation based on nurse-specific actions and prior contextual information, and those not requiring specialized expertise, such as the use of standard medical devices. In contrast, records that were difficult to infer primarily involved patient interviews, which lacked observable variations in appearance. These findings suggest the potential utility of the MLLM for drafting nursing records and the potential value of incorporating audio information, and delineate barriers to deployment in clinical settings.

READ FULL ARTICLE ON J-STAGE