Proprietary
Open-Source
Click a metric header to sort
| Model | Release | Anatomical Recognition | Diagnosis Analysis | Staging Acc | ||||
|---|---|---|---|---|---|---|---|---|
| Prec. | Rec. | F1 | Halluc. Resist. | Binary Acc | Multi-class Acc | |||
| Random Baseline | N/A | N/A | N/A | N/A | 0.5000 | 0.2500 | 0.2500 | |
| GPT-4o | 2024-05 | 0.3116 | 0.2424 | 0.2577 | 0.9786 | N/A | N/A | 0.1053 |
LLaVA-Med-v1.5-Mistral-7B |
2023-06 | 0.2090 | 0.2558 | 0.2098 | 0.6240 | 0.5000 | 0.2500 | 0.2368 |
| 2024-05 | 0.1509 | 0.0379 | 0.0451 | 0.8135 | 0.4117 | 0.2525 | 0.4000 | |
Med-Flamingo |
Invalid | Invalid | Invalid | Invalid | Invalid | Invalid | Invalid | |
InternVL 1.5-2B |
2024-05 | 0.2377 | 0.2386 | 0.2204 | 0.7722 | 0.5000 | 0.2525 | 0.2632 |
InternVL 1.5-4B |
2024-05 | 0.2499 | 0.2541 | 0.2414 | 0.8096 | 0.5000 | 0.2500 | 0.2500 |
InternVL 2.0-2B |
2024-07 | 0.3718 | 0.3912 | 0.3722 | 0.9069 | 0.5000 | 0.2175 | 0.2500 |
InternVL 2.0-4B |
2024-07 | 0.4200 | 0.4294 | 0.4096 | 0.8951 | 0.5000 | 0.2525 | 0.2237 |
InternVL 2.0-8B |
2024-07 | 0.2865 | 0.2872 | 0.2770 | 0.9533 | 0.5017 | 0.3500 | 0.2500 |
InternVL 2.5-2B |
2024-12 | 0.3751 | 0.3681 | 0.3524 | 0.9744 | 0.5017 | 0.3000 | 0.2308 |
InternVL 2.5-4B |
2024-12 | 0.2685 | 0.1893 | 0.1890 | 0.9854 | 0.5000 | 0.3225 | 0.2500 |
InternVL 2.5-8B |
2024-12 | 0.2851 | 0.2519 | 0.2574 | 0.9864 | 0.5000 | 0.3275 | 0.2500 |
InternVL 2.5-2B-MPO |
2025-04 | 0.3123 | 0.2992 | 0.2835 | 0.9646 | 0.5000 | 0.2375 | 0.3077 |
InternVL 2.5-4B-MPO |
2025-04 | 0.3126 | 0.2261 | 0.2320 | 0.9918 | 0.5000 | 0.3200 | 0.2500 |
InternVL 2.5-8B-MPO |
2025-04 | 0.2986 | 0.2536 | 0.2624 | 0.9838 | 0.5000 | 0.2650 | 0.1053 |
LLaVA-1.5-7B |
2023-10 | 0.0965 | 0.0530 | 0.0557 | 0.4629 | 0.4937 | 0.2475 | 0.2500 |
LLaVA-Mistral-7B |
2024-01 | 0.0806 | 0.0773 | 0.0696 | 0.5788 | 0.5000 | 0.2500 | 0.2500 |
LLaVA-Vicuna-7B |
2024-01 | 0.0437 | 0.0353 | 0.0342 | 0.1929 | 0.5000 | 0.3693 | N/A |
LLaVA-Vicuna-13B |
2024-01 | 0.1123 | 0.0071 | 0.0105 | 0.6612 | 0.5000 | 0.2725 | N/A |
Qwen-VL-Chat |
2023-08 | 0.1040 | 0.0120 | 0.0178 | 0.7790 | 0.5000 | 0.2675 | 0.2763 |
Qwen-3B |
0.2611 | 0.1489 | 0.1509 | 0.7468 | 0.5000 | 0.2500 | 0.2368 | |
Qwen-7B |
2023-08 | 0.2506 | 0.2261 | 0.2251 | 0.7556 | 0.5017 | 0.2500 | 0.2368 |
| 2024-12 | 0.2228 | 0.0583 | 0.0688 | 0.9518 | 0.5000 | 0.2500 | 0.2237 | |
| 2024-12 | 0.0805 | 0.0130 | 0.0218 | 0.5917 | Invalid | 0.1700 | 0.0667 | |
Comprehensive zero-shot evaluation across anatomical recognition, diagnosis, and stage classification. The best result in each column is bold, the second best is underlined. All models use consistent prompt templates.
Submit your model. To appear on this leaderboard, send the scores and a brief description of your multimodal model to Zhenyue Qin at zhenyue.qin@yale.edu.
LLaVA-Med-v1.5-Mistral-7B
Med-Flamingo
InternVL 1.5-2B
LLaVA-1.5-7B
Qwen-VL-Chat