a Harvard Ophthalmology AI Lab, Schepens Eye Research Institute of Massachusetts Eye and Ear, Harvard Medical School, Boston, MA, USA · b Media Lab, Massachusetts Institute of Technology, Cambridge, MA, USA · c Dept. of CSE, The Hong Kong University of Science and Technology, Hong Kong, China · d School of Computing and Informatics, University of Louisiana at Lafayette, LA, USA · e International School of Medicine, Istanbul Medipol University, Istanbul, Turkey
1 These three authors contributed equally. * Corresponding: mengyu_wang@meei.harvard.edu, ppliang@mit.edu
Glaucoma is a leading cause of irreversible blindness worldwide. Ophthalmologists diagnose it through a structured reasoning process — sequentially evaluating optic-nerve-head characteristics before reaching a final diagnosis — whereas existing AI systems typically perform direct image classification without clinically meaningful reasoning. We present the first clinically annotated fundus reasoning dataset, 1,077 fundus photographs paired with expert-authored six-step diagnostic reports, and a reasoning-driven vision–language framework that explicitly models the ophthalmologist's diagnostic workflow by generating structured clinical reasoning prior to diagnosis. The generated reports are clinically validated, achieving the best performance across all evaluated clinical findings — a cup-to-disc-ratio MAE of 0.070, an ISNT Kendall distance of 1.73, and the highest semantic agreement with expert reports (BERTScore-F1 = 0.874). The framework also improves diagnosis, reaching a balanced accuracy of 94.7% and precision of 94.8%, showing that explicitly modeling expert clinical reasoning simultaneously improves interpretability and diagnostic performance.
Conventional screening maps a fundus image directly to a label. Our framework keeps the same frozen RETFound backbone but inserts explicit clinical reasoning: the backbone first produces structured clinical evidence, a vision–language model then writes a six-step chain-of-thought, and only then is the binary diagnosis made. The improvement therefore comes from added clinical reasoning, not from a stronger visual backbone.
Assess focus, contrast, artifacts and disc-margin clarity for reliable analysis.
Estimate vertical & horizontal cup-to-disc ratio against the glaucomatous cutoff.
Rank inferior–superior–nasal–temporal rim widths and flag violations.
Check notching, RNFL defect, bayoneting, PPA / β-zone atrophy, disc hemorrhage.
Integrate the findings and weigh the automated probability against them.
Reason to a binary glaucoma / non-glaucoma decision with justification.
A Papila test-split fundus photograph is passed to the model, which outputs the six-step report below. Steps 1–5 are the reasoning; step 6 is the final classification — here matching the ground-truth label.
1,077 color fundus photographs from the public LAG and Papila datasets, each paired with a complete expert-authored six-step reasoning report (average length 321.6 ± 46.3 words) and a binary label.
| Characteristic | Value |
|---|---|
| Fundus images | 1,077 |
| Glaucoma | 430 (39.9%) |
| Non-glaucoma | 647 (60.1%) |
| LAG database | 659 (343 / 316) |
| Papila dataset | 418 (87 / 331) |
| Training split | 825 (315 / 510) |
| Validation split | 92 (35 / 57) |
| Test split | 160 (80 / 80) |
| Avg. report length | 321.6 ± 46.3 words |
Six reasoning indicators are annotated per image (steps 2–4 of the CoT):
Every image is paired with an expert-authored reasoning report, making this the first large-scale glaucoma fundus dataset with multi-step reasoning annotations.
Against a fine-tuned Qwen3.5-VL and the proprietary foundation models Claude and GPT-5.5, our framework achieves the strongest performance on every evaluated clinical indicator and the highest semantic agreement with expert reports.
| Metric | Ours | Qwen3.5-VL | Claude | GPT-5.5 |
|---|---|---|---|---|
| CDR MAE ↓ | 0.070 | 0.102 | 0.135 | 0.127 |
| ISNT Kendall ↓ | 1.73 | 2.07 | 2.27 | n/a |
| Rim-level macro-F1 ↑ | 0.656 | 0.512 | 0.510 | 0.373 |
| Signs macro-F1 ↑ | 0.613 | 0.501 | 0.465 | 0.375 |
| BERTScore-F1 ↑ | 0.874 | 0.867 | 0.869 | 0.866 |
| ROUGE-Lsum ↑ | 0.454 | 0.416 | 0.461 | 0.439 |
| Model | Bal. Acc. | 95% CI |
|---|---|---|
| Ours (VLM) | 94.7% | 90.6–98.1 |
| RetiZero | 93.6% | 88.1–98.0 |
| RETFound | 93.1% | 87.6–97.5 |
| DINOv2-L | 84.7% | 77.1–91.2 |
| Qwen3.5-VL | 84.3% | 78.8–89.5 |
| Claude | 74.4% | 67.5–81.0 |
| GPT-5.5 | 71.9% | 65.1–78.2 |
Removing the reasoning stage, or replacing structured evidence with only an initial diagnostic signal, collapses the balanced diagnosis — confirming that the structured clinical evidence and intermediate reasoning, not the backbone alone, drive reliable performance.
| Configuration | Bal. Acc. | Sens. | Spec. |
|---|---|---|---|
| Full framework | 94.69% | 92.88% | 96.50% |
| Abl. 1 — image only | 83.75% | 97.50% | 70.00% |
| Abl. 2 — image + diagnosis | 70.63% | 98.75% | 42.50% |
| Variant | Bal. Acc. |
|---|---|
| Ours (all indicators) | 94.69% |
| − per-quadrant rim status | 94.38% |
| − ISNT rim ordering | 94.37% |
| − glaucomatous signs | 93.75% |
| − cup-to-disc ratio | 92.50% |
The cup-to-disc ratio is the single most informative cue (largest drop when removed), while every indicator contributes complementary diagnostic information.
@article{zhou2026glaucoma,
title = {Learning Ophthalmologist Clinical Reasoning for Glaucoma Diagnosis from Fundus Images},
author = {Zhou, Kaichen and Chen, Yuzhen and Yildiz, Elif and Shi, Min and Dai, David
and Chen, Grace and Zheng, Jiale and Wang, He and Zhan, Fangneng and Saini, Chhavi
and Shen, Lucy Q. and Guo, Yike and Liang, Paul Pu and Wang, Mengyu},
journal = {npj Digital Medicine (under review)},
year = {2026}
}