LMOD+ Benchmark

Leaderboard

Zero-shot evaluation of multimodal large language models on the LMOD+ subset of 1,076 curated images, covering anatomical recognition, disease diagnosis, and staging assessment with consistent prompt templates.

Proprietary Open-Source Click a metric header to sort
Model Release Anatomical Recognition Diagnosis Analysis Staging Acc
Prec. Rec. F1 Halluc. Resist. Binary Acc Multi-class Acc
Random Baseline N/AN/AN/AN/A 0.50000.25000.2500
GPT-4o 2024-05 0.31160.24240.25770.9786 N/AN/A0.1053
LLaVA-Med-v1.5-Mistral-7B 2023-06 0.20900.25580.20980.6240 0.50000.25000.2368
Yi-VL-6B 2024-05 0.15090.03790.04510.8135 0.41170.25250.4000
Med-Flamingo InvalidInvalidInvalidInvalid InvalidInvalidInvalid
InternVL 1.5-2B 2024-05 0.23770.23860.22040.7722 0.50000.25250.2632
InternVL 1.5-4B 2024-05 0.24990.25410.24140.8096 0.50000.25000.2500
InternVL 2.0-2B 2024-07 0.37180.39120.37220.9069 0.50000.21750.2500
InternVL 2.0-4B 2024-07 0.42000.42940.40960.8951 0.50000.25250.2237
InternVL 2.0-8B 2024-07 0.28650.28720.27700.9533 0.50170.35000.2500
InternVL 2.5-2B 2024-12 0.37510.36810.35240.9744 0.50170.30000.2308
InternVL 2.5-4B 2024-12 0.26850.18930.18900.9854 0.50000.32250.2500
InternVL 2.5-8B 2024-12 0.28510.25190.25740.9864 0.50000.32750.2500
InternVL 2.5-2B-MPO 2025-04 0.31230.29920.28350.9646 0.50000.23750.3077
InternVL 2.5-4B-MPO 2025-04 0.31260.22610.23200.9918 0.50000.32000.2500
InternVL 2.5-8B-MPO 2025-04 0.29860.25360.26240.9838 0.50000.26500.1053
LLaVA-1.5-7B 2023-10 0.09650.05300.05570.4629 0.49370.24750.2500
LLaVA-Mistral-7B 2024-01 0.08060.07730.06960.5788 0.50000.25000.2500
LLaVA-Vicuna-7B 2024-01 0.04370.03530.03420.1929 0.50000.3693N/A
LLaVA-Vicuna-13B 2024-01 0.11230.00710.01050.6612 0.50000.2725N/A
Qwen-VL-Chat 2023-08 0.10400.01200.01780.7790 0.50000.26750.2763
Qwen-3B 0.26110.14890.15090.7468 0.50000.25000.2368
Qwen-7B 2023-08 0.25060.22610.22510.7556 0.50170.25000.2368
DeepSeek VL2-Tiny 2024-12 0.22280.05830.06880.9518 0.50000.25000.2237
DeepSeek VL2-Small 2024-12 0.08050.01300.02180.5917 Invalid0.17000.0667

Comprehensive zero-shot evaluation across anatomical recognition, diagnosis, and stage classification. The best result in each column is bold, the second best is underlined. All models use consistent prompt templates.

Submit your model. To appear on this leaderboard, send the scores and a brief description of your multimodal model to Zhenyue Qin at zhenyue.qin@yale.edu.