LMOD+

LMOD+

A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating
Multimodal Large Language Models in Ophthalmology

ACM ACM Transactions on Computing for Healthcare
An extension of LMOD, Findings of NAACL 2025, with a 50% larger dataset, broader clinical tasks, and 24 MLLMs evaluated
32,633 Instances
12 Conditions
5 Modalities
24 MLLMs Evaluated
Zhenyue Qin1* Yang Liu2* Yu Yin3 Jinyu Ding1 Haoran Zhang1 Anran Li1 Dylan Campbell2 Xuansheng Wu4 Ke Zou5 Tiarnan D. L. Keenan6 Emily Y. Chew6 Zhiyong Lu7 Yih Chung Tham5 Ninghao Liu4 Xiuzhen Zhang8 Qingyu Chen1†
1USASchool of Medicine, Yale University 2AustraliaAustralian National University 3UKImperial College London 4USAUniversity of Georgia 5SingaporeNational University of Singapore 6USANational Eye Institute, NIH 7USANational Library of Medicine, NIH 8AustraliaRMIT University
Abstract

The rising prevalence of vision-threatening eye diseases poses a major global health and economic burden, yet timely diagnosis remains limited by workforce shortages, diagnostic delays, and restricted access to specialized care. While multimodal large language models (MLLMs) have shown promise in medical image interpretation and automated clinical documentation, advancing MLLMs for ophthalmology is hindered by the lack of unified, comprehensive benchmark datasets.

We present LMOD+, a large-scale multimodal ophthalmology benchmark comprising 32,633 instances with multi-granular annotations across 12 common ophthalmic conditions and 5 imaging modalities. The dataset integrates imaging, anatomical structures, demographics, and free-text annotations, supporting anatomical recognition, disease screening, disease staging, and demographic prediction.

LMOD+ extends our preliminary LMOD benchmark (Findings of NAACL 2025) with three major enhancements: a ~50% dataset expansion, broadened task coverage including binary and multi-class diagnosis with internationally adopted grading standards, and systematic evaluation of 24 state-of-the-art MLLMs from the InternVL, Qwen, and DeepSeek families. Top models achieve ~58% accuracy in zero-shot disease screening, a considerably more challenging paradigm than traditional supervised fine-tuning.

0
Total Instances
0
Ophthalmic Conditions
0
Imaging Modalities
0
MLLMs Evaluated
Dataset

Five Ophthalmic Imaging Modalities

LMOD+ integrates nine publicly accessible ophthalmology datasets into a unified benchmark, curated with multi-granular annotations spanning anatomical regions, disease labels, demographics, and free-text clinical reports.

SS

Surgical Scenes

2,256 images

Cataract surgery frames with detailed procedural and anatomical region annotations. Average 3.3 bounding boxes per image.

Cataract-1K
OCT

Optical Coherence Tomography

3,859 images

High-resolution cross-sectional retinal scans enabling macular hole staging and retinal layer analysis. Average 2.4 boxes per image.

OIMHS
SLO

Scanning Laser Ophthalmoscopy

10,000 images

The largest modality in LMOD+. Fundus images supporting equitable glaucoma diagnosis research across demographic subgroups.

Harvard FairSeg
LP

Lens Photographs

2,432 images

Eye-region and cataract images covering pupil, iris, and lens clarity assessment. Average 1.9 bounding boxes per image.

CAU001 · CatDet2
CFP

Color Fundus Photography

3,386 images

Retinal photographs covering glaucoma, diabetic retinopathy, and optic disc segmentation from four public datasets.

REFUGE · IDRiD · ORIGA · G1020
Benchmark

Multi-Granular Clinical Evaluation

LMOD+ evaluates MLLMs across four clinically grounded task categories, from anatomical structure recognition to demographic bias analysis.

01

Anatomical Recognition

Identifies and localizes structures within ophthalmic images, from retinal layers in OCT to optic disc boundaries in fundus photographs.

PrecisionRecallF1Hallucination Resistance
02

Disease Screening

Binary detection of 12 ophthalmic conditions including diabetic retinopathy, glaucoma, AMD, and retinal vein occlusion.

AccuracyF1 Score
03

Disease Staging

Severity classification using internationally adopted grading standards, including international clinical DR and Scottish DR grading schemes.

Stage AccuracyCohen's Kappa
04

Demographic Prediction

Predicts patient age and sex from ophthalmic images to quantify potential model bias and enable subgroup fairness analysis.

Age AccuracySex Accuracy
Six Identified Failure Modes in Current MLLMs
Misclassification Failure to Abstain Inconsistent Reasoning Hallucination Unjustified Assertions Insufficient Domain Knowledge
Results

MLLM Performance on LMOD+

Comprehensive evaluation of 24 MLLMs reveals persistent gaps between general-domain and ophthalmology-specific performance, with staging tasks remaining particularly challenging.

Model Anatomical Recognition Disease Screening
Precision F1 Hallucination Resistance Glaucoma Acc. MH Stage Acc.
GPT-4o 0.561 0.575 0.951 54.09% 19.71%
Qwen-7B 0.724 0.554 0.963 58.26% 7.30%
InternVL-2B 0.603 0.481 0.981 57.83% 30.26%
InternVL-4B 0.028 0.037 0.963 50.00% 18.42%
LLaVA-13B 0.060 0.170 0.599 50.00% 5.26%
Avg. (24 models) 0.269 0.219 0.754 ~50% ~15%
Substantial performance gap

MLLMs exhibit consistent drops in ophthalmic tasks. Disease screening accuracy hovers near chance for most models.

Zero-shot screening is hard

Best single-model accuracy is ~58%, far below supervised fine-tuned baselines that achieve near-perfect results.

InternVL & Qwen lead open-source

InternVL-2B achieves the best MH staging (30.26%); Qwen-7B achieves the best single-model glaucoma screening (58.26%).

Staging remains unsolved

Disease staging performance is often at or below random baselines, the hardest challenge for current MLLMs.

Citation

BibTeX

If you find our work useful, please consider citing both LMOD+ and the preliminary LMOD benchmark.

LMOD+ ACM Transactions on Computing for Healthcare · 2026
@article{10.1145/3801746,
  author    = {Qin, Zhenyue and Liu, Yang and Yin, Yu and Ding, Jinyu
               and Zhang, Haoran and Li, Anran and Campbell, Dylan
               and Wu, Xuansheng and Zou, Ke and Keenan, Tiarnan D. L.
               and Chew, Emily Y. and Lu, Zhiyong and Tham, Yih Chung
               and Liu, Ninghao and Zhang, Xiuzhen and Chen, Qingyu},
  title     = {LMOD \(\boldsymbol{+}\): A Comprehensive Multimodal
               Dataset and Benchmark for Developing and Evaluating
               Multimodal Large Language Models in Ophthalmology},
  year      = {2026},
  issue_date = {July 2026},
  publisher = {Association for Computing Machinery},
  address   = {New York, NY, USA},
  volume    = {7},
  number    = {3},
  url       = {https://doi.org/10.1145/3801746},
  doi       = {10.1145/3801746},
  journal   = {ACM Trans. Comput. Healthcare},
  month     = jun,
  articleno = {52},
  numpages  = {38},
  keywords  = {Multimodal large language models, ophthalmology,
               medical AI, benchmark dataset, healthcare computing}
}
LMOD Findings of NAACL · 2025
@inproceedings{qin-etal-2025-lmod,
  title     = {{LMOD}: A Large Multimodal Ophthalmology Dataset and
               Benchmark for Large Vision-Language Models},
  author    = {Qin, Zhenyue and Yin, Yu and Campbell, Dylan and
               Wu, Xuansheng and Zou, Ke and Liu, Ninghao and
               Tham, Yih Chung and Zhang, Xiuzhen and Chen, Qingyu},
  editor    = {Chiruzzo, Luis and Ritter, Alan and Wang, Lu},
  booktitle = {Findings of the Association for Computational
               Linguistics: NAACL 2025},
  month     = apr,
  year      = {2025},
  address   = {Albuquerque, New Mexico},
  publisher = {Association for Computational Linguistics},
  url       = {https://aclanthology.org/2025.findings-naacl.135/},
  doi       = {10.18653/v1/2025.findings-naacl.135},
  pages     = {2501--2522},
  ISBN      = {979-8-89176-195-7}
}