A Comprehensive Multimodal Dataset and Benchmark for Developing and Evaluating
Multimodal Large Language Models in Ophthalmology
The rising prevalence of vision-threatening eye diseases poses a major global health and economic burden, yet timely diagnosis remains limited by workforce shortages, diagnostic delays, and restricted access to specialized care. While multimodal large language models (MLLMs) have shown promise in medical image interpretation and automated clinical documentation, advancing MLLMs for ophthalmology is hindered by the lack of unified, comprehensive benchmark datasets.
We present LMOD+, a large-scale multimodal ophthalmology benchmark comprising 32,633 instances with multi-granular annotations across 12 common ophthalmic conditions and 5 imaging modalities. The dataset integrates imaging, anatomical structures, demographics, and free-text annotations, supporting anatomical recognition, disease screening, disease staging, and demographic prediction.
LMOD+ extends our preliminary LMOD benchmark (Findings of NAACL 2025) with three major enhancements: a ~50% dataset expansion, broadened task coverage including binary and multi-class diagnosis with internationally adopted grading standards, and systematic evaluation of 24 state-of-the-art MLLMs from the InternVL, Qwen, and DeepSeek families. Top models achieve ~58% accuracy in zero-shot disease screening, a considerably more challenging paradigm than traditional supervised fine-tuning.
LMOD+ integrates nine publicly accessible ophthalmology datasets into a unified benchmark, curated with multi-granular annotations spanning anatomical regions, disease labels, demographics, and free-text clinical reports.
Cataract surgery frames with detailed procedural and anatomical region annotations. Average 3.3 bounding boxes per image.
High-resolution cross-sectional retinal scans enabling macular hole staging and retinal layer analysis. Average 2.4 boxes per image.
The largest modality in LMOD+. Fundus images supporting equitable glaucoma diagnosis research across demographic subgroups.
Eye-region and cataract images covering pupil, iris, and lens clarity assessment. Average 1.9 bounding boxes per image.
Retinal photographs covering glaucoma, diabetic retinopathy, and optic disc segmentation from four public datasets.
LMOD+ evaluates MLLMs across four clinically grounded task categories, from anatomical structure recognition to demographic bias analysis.
Identifies and localizes structures within ophthalmic images, from retinal layers in OCT to optic disc boundaries in fundus photographs.
Binary detection of 12 ophthalmic conditions including diabetic retinopathy, glaucoma, AMD, and retinal vein occlusion.
Severity classification using internationally adopted grading standards, including international clinical DR and Scottish DR grading schemes.
Predicts patient age and sex from ophthalmic images to quantify potential model bias and enable subgroup fairness analysis.
Comprehensive evaluation of 24 MLLMs reveals persistent gaps between general-domain and ophthalmology-specific performance, with staging tasks remaining particularly challenging.
| Model | Anatomical Recognition | Disease Screening | |||
|---|---|---|---|---|---|
| Precision | F1 | Hallucination Resistance | Glaucoma Acc. | MH Stage Acc. | |
| GPT-4o | 0.561 | 0.575 | 0.951 | 54.09% | 19.71% |
| Qwen-7B | 0.724 | 0.554 | 0.963 | 58.26% | 7.30% |
| InternVL-2B | 0.603 | 0.481 | 0.981 | 57.83% | 30.26% |
| InternVL-4B | 0.028 | 0.037 | 0.963 | 50.00% | 18.42% |
| LLaVA-13B | 0.060 | 0.170 | 0.599 | 50.00% | 5.26% |
| Avg. (24 models) | 0.269 | 0.219 | 0.754 | ~50% | ~15% |
MLLMs exhibit consistent drops in ophthalmic tasks. Disease screening accuracy hovers near chance for most models.
Best single-model accuracy is ~58%, far below supervised fine-tuned baselines that achieve near-perfect results.
InternVL-2B achieves the best MH staging (30.26%); Qwen-7B achieves the best single-model glaucoma screening (58.26%).
Disease staging performance is often at or below random baselines, the hardest challenge for current MLLMs.
If you find our work useful, please consider citing both LMOD+ and the preliminary LMOD benchmark.
@article{10.1145/3801746,
author = {Qin, Zhenyue and Liu, Yang and Yin, Yu and Ding, Jinyu
and Zhang, Haoran and Li, Anran and Campbell, Dylan
and Wu, Xuansheng and Zou, Ke and Keenan, Tiarnan D. L.
and Chew, Emily Y. and Lu, Zhiyong and Tham, Yih Chung
and Liu, Ninghao and Zhang, Xiuzhen and Chen, Qingyu},
title = {LMOD \(\boldsymbol{+}\): A Comprehensive Multimodal
Dataset and Benchmark for Developing and Evaluating
Multimodal Large Language Models in Ophthalmology},
year = {2026},
issue_date = {July 2026},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
volume = {7},
number = {3},
url = {https://doi.org/10.1145/3801746},
doi = {10.1145/3801746},
journal = {ACM Trans. Comput. Healthcare},
month = jun,
articleno = {52},
numpages = {38},
keywords = {Multimodal large language models, ophthalmology,
medical AI, benchmark dataset, healthcare computing}
}
@inproceedings{qin-etal-2025-lmod,
title = {{LMOD}: A Large Multimodal Ophthalmology Dataset and
Benchmark for Large Vision-Language Models},
author = {Qin, Zhenyue and Yin, Yu and Campbell, Dylan and
Wu, Xuansheng and Zou, Ke and Liu, Ninghao and
Tham, Yih Chung and Zhang, Xiuzhen and Chen, Qingyu},
editor = {Chiruzzo, Luis and Ritter, Alan and Wang, Lu},
booktitle = {Findings of the Association for Computational
Linguistics: NAACL 2025},
month = apr,
year = {2025},
address = {Albuquerque, New Mexico},
publisher = {Association for Computational Linguistics},
url = {https://aclanthology.org/2025.findings-naacl.135/},
doi = {10.18653/v1/2025.findings-naacl.135},
pages = {2501--2522},
ISBN = {979-8-89176-195-7}
}