REVIEW 3 major objections 6 minor 25 references
AI-powered virtual eye: perspective, challenges and opportunities
T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read A next-generation AI platform built from interconnected foundation models could simulate the human eye across every scale, from molecules to the whole organ, and support personalized eye care.
desk verdict A competent perspective that imports the virtual cell blueprint into ophthalmology; useful synthesis, thin on feasibility, and the abstract oversells the near-term promise. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the virtual eye itself: a proposed platform of interconnected foundation models spanning molecular to organ scale, with shared representations, generative simulation, and continuous feedback. The load-bearing mechanisms are: foundation models pre-trained on large ophthalmic datasets that provide generalizable backbones; multimodal contrastive learning that produces shared latent spaces; generative AI that creates synthetic or missing data and reconstructs 3D structures; agent-based architectures that route tasks and update knowledge; and internal plus external feedback loops that allow dynamic recalibration from real-world data. The paper also proposes a 'divide and conquer' construction strategy and a hierarchical evaluation framework across molecular, tissue-organ, clinical, and longitudinal levels.
What would settle it
If a concrete prototype following the roadmap cannot align data from the same eye across modalities (for example, single-cell transcriptomics, OCT images, and genetic variants) into a single spatiotemporal reference without contradictions, or if a zero-shot prediction of disease progression from baseline multimodal data is no better than a single-modality model, the central feasibility claim would be undercut.
Extended reading notes
Core claim
The paper's central claim is that advances in AI, imaging, and multi-omics make it feasible to construct a universal, high-fidelity digital replica of the human eye—the virtual eye. The authors define it as a platform of interconnected foundation models with four hallmarks: multimodal modeling, multi-scale integration, representation of dynamic processes, and complex feedback loops. Unlike earlier stage models, this universal virtual eye would be hybrid: it retains mechanistic insights from optics, biomechanics, fluid dynamics, and pharmacokinetics while using generative AI and foundation models to learn across modalities and simulate untested interventions. The paper argues such a system could become an in silico laboratory and a clinical decision-support tool, shifting ophthalmology toward proactive, personalized care. It also recommends a 'divide and conquer' development path, building modular subsystems and later integrating them, and a hierarchical four-level evaluation strategy spanning molecular, tissue-organ, clinical, and longitudinal levels.
Load-bearing premise
The roadmap depends on the assumption that heterogeneous data—imaging, molecular profiles, clinical records, and environmental streams—can be integrated into a unified, spatiotemporally aligned reference framework spanning molecular to organ scales, and that interconnected foundation models can maintain self-consistency across contexts; the paper presents no pilot demonstration of this.
Editorial extensions
If this is right
- Clinicians could compare a patient's current eye state against a digital twin's predicted trajectory to detect early deviations from healthy baselines and intervene sooner.
- Surgeons could rehearse cataract or refractive procedures on a patient-specific virtual eye before operating, allowing them to test different intervention strategies.
- Researchers could use the virtual eye as an in silico laboratory to investigate causal mechanisms, generate hypotheses, and prioritize wet-lab experiments.
- The platform could incorporate wearable and environmental data streams, such as smart contact lenses and smartwatch light-exposure measures, to update risk predictions in real time.
- A dedicated data-processing AI could autonomously annotate, clean, and harmonize heterogeneous ophthalmic datasets, creating a reusable common data representation.
Reading between the lines
- If the roadmap succeeds, the eye could serve as a proving ground for organ-level digital twins generally, because its accessible imaging and rich data landscape make it one of the easiest organs to reconstruct and validate.
- The modular 'divide and conquer' strategy implies that near-term progress may come from linking existing modality-specific foundation models into pipelines long before a single universal model exists, so incremental clinical tools could arrive first.
- A testable extension would be to benchmark whether a shared multimodal representation trained on paired fundus images, OCT volumes, and genomic data enables zero-shot cross-modal predictions that single-modality models cannot make.
- The emphasis on feedback loops suggests that the most informative evaluation would be longitudinal: compare a continuously updated digital twin against a static model on real clinical data to measure whether adaptive recalibration actually improves prediction.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This perspective paper argues that advances in AI, imaging, and multi-omics make it feasible to construct an "AI-powered virtual eye": a platform of interconnected foundation models that simulate the eye across molecular, cellular, tissue, and organ scales. The paper surveys the historical evolution of eye modeling in three stages (mechanistic, deep-learning-based, and the proposed universal virtual eye), then presents a roadmap organized around data acquisition, modeling architecture, and human/environment interaction. It lists challenges concerning interpretability, ethics, data standardization, and evaluation, and closes with envisioned applications in research and clinical care. The central claim is aspirational: no prototype, pilot study, or quantitative demonstration is provided, and the paper itself acknowledges that a unified cross-scale framework remains elusive.
Significance. If the proposed vision were realized, the virtual eye could meaningfully advance personalized ophthalmology and in silico research, paralleling the virtual cell and digital twin movements. The paper provides a useful, well-referenced synthesis of mechanistic modeling, deep learning, and foundation models in ophthalmology, and its three-stage taxonomy (Tables 1 and 2) clarifies the conceptual landscape. It also names several concrete challenges (cross-context self-consistency, interpretability, evaluation) that are likely to be central to any serious effort. The paper does not contain machine-checked proofs, reproducible code, parameter-free derivations, or falsifiable predictions; its value lies in framing a research agenda rather than validating one. The main risk is that the roadmap's load-bearing premise, a unified spatiotemporally aligned cross-scale reference, is asserted rather than supported, and the paper itself concedes that key ingredients are missing.
major comments (3)
- [§3.1.1, §3.2.1]
- [§4.3]
- [§2.3, Table 2]
minor comments (6)
- [Abstract/§1]
- [§1 and throughout]
- [§3.1.2]
- [Table 1]
- [§6]
- [References]
Circularity Check
No significant circularity: the paper is a perspective with no derived predictions; its self-citations are illustrative, and the central vision is explicitly conceded to be unbuilt.
full rationale
This perspective paper contains no equations, fitted parameters, or performed predictions, so the standard circular failure modes (self-definitional derivation, fitted input called a prediction) do not arise. The central claim is a forward-looking proposal that a universal AI-powered virtual eye is feasible, and the paper's own text repeatedly concedes that the crux of that proposal is unbuilt: Section 3.2.1 states that 'a unified framework linking molecules, pathways, cells, and whole organ remains elusive,' and the Conclusion lists 'cross-context self-consistency' as critical and unresolved. A proposal to build a system, combined with an explicit admission that the system does not yet exist, cannot reduce to its own inputs by construction. The few self-citations (EyeFound, EyeCLIP, Fundus2Globe, and the CFP-to-FFA translation works) appear only as examples of partial, component-level progress, described as 'an early form of an organ-specific foundation model' and as demonstrations that 'reconstructing 3D structures... from planar imaging data' is feasible; they are not cited as evidence that the complete virtual eye functions, so they are not load-bearing. The skeptic's concern that building the spatiotemporally aligned cross-scale reference would require the very integrated biological model the virtual eye is meant to learn is a feasibility objection rather than a definitional circularity, and the manuscript itself flags cross-scale alignment as an open challenge while proposing a hierarchical evaluation strategy (Section 4.4) that validates molecular/cellular accuracy, tissue/organ simulation, clinical outcomes, and longitudinal adaptation at separate levels against independent ground truth, which is precisely the non-circular validation structure. Under the review rules, an acknowledged open problem presented as an open problem, with no result derived from it, does not constitute circularity. Multiple external anchors (mechanistic eye models, RETFound, AlphaFold, Evo, MorphoDiff, and virtual-cell frameworks) provide independent support for the component technologies, and the paper is self-contained in the sense that it derives no result whose conclusion is required by its own assumptions. Verdict: no significant circularity, score 1.
Assumptions & free parameters
assumptions (3)
- domain assumption Advances in AI, imaging, and multiomics are sufficient to build a universal high-fidelity digital replica of the human eye.
- domain assumption Multimodal biological data can be integrated into a unified spatiotemporally aligned reference framework.
- domain assumption Interconnected foundation models can be composed to maintain cross-scale and cross-context consistency.
invented entities (1)
-
AI-powered virtual eye platform
Cite this review
Pith. "Pith review of AI-powered virtual eye: perspective, challenges and opportunities." pith.science (2026). https://pith.science/paper/G743R3RV
@misc{pith2026250505516,
author = {Pith},
title = {Pith review of: AI-powered virtual eye: perspective, challenges and opportunities},
year = {2026},
howpublished = {\url{https://pith.science/paper/G743R3RV}},
note = {Machine review of arXiv:2505.05516}
}
read the original abstract
We envision the "virtual eye" as a next-generation, AI-powered platform that uses interconnected foundation models to simulate the eye's intricate structure and biological function across all scales. Advances in AI, imaging, and multiomics provide a fertile ground for constructing a universal, high-fidelity digital replica of the human eye. This perspective traces the evolution from early mechanistic and rule-based models to contemporary AI-driven approaches, integrating in a unified model with multimodal, multiscale, dynamic predictive capabilities and embedded feedback mechanisms. We propose a development roadmap emphasizing the roles of large-scale multimodal datasets, generative AI, foundation models, agent-based architectures, and interactive interfaces. Despite challenges in interpretability, ethics, data processing and evaluation, the virtual eye holds the potential to revolutionize personalized ophthalmic care and accelerate research into ocular health and disease.
Figures
Reference graph
Works this paper leans on
-
[1]
School of Optometry, The Hong Kong Polytechnic University, Hong Kong SAR, China
-
[2]
Key Laboratory of Carcinogenesis and Cancer Invasion, Liver Cancer Institute, Zhongshan Hospital, Fudan University, Shanghai, China
-
[3]
School of Medicine, Shanghai Jiao Tong University, Shanghai, China
-
[4]
Swiss Federal Institute of Technology Lausanne (EPFL), Lausanne, Switzerland
-
[5]
Intelligent Medicine Institute, Fudan Microbiome Center, Fudan University Shanghai Medical College, Fudan University, Shanghai, China
-
[6]
Department of Genetics, Stanford University School of Medicine, Stanford, CA, USA
-
[7]
Collaboratory on Longitudinal Deep Omics, The Hong Kong Polytechnic University, Hong Kong SAR, China
-
[8]
Department of Ophthalmology, Yong Loo Lin School of Medicine, National University of Singapore, Singapore, Singapore
Show all 25 references
-
[9]
Singapore Eye Research Institute, Singapore National Eye Centre, Singapore, Singapore
-
[10]
Research Centre for SHARP Vision (RCSV), The Hong Kong Polytechnic University, Hong Kong SAR, China
-
[11]
virtual eye
Centre for Eye and Vision Research (CEVR), 17W Hong Kong Science Park, Hong Kong SAR, China #Contributed equally Correspondence Prof. Mingguang He, MD, PhD., Chair Professor of Experimental Ophthalmology, School of Optometry, The Hong Kong Polytechnic University, Hong Kong, Ch...
-
[12]
seeing , interpreting, and predicting
Introduction A computational eye model aims to simulate, generate, predict, and analyze the structural and functional states of the eye. Owing to its rich imaging landscape and well- characterized anatomy, the eye serves as an ideal organ for virtual reconstruction. Traditiona...
-
[13]
virtual eye
Conceptual evolution of the eye model The pursuit of a “virtual eye ” began with early computational models that used mathematics, physics, statistics, and computer science to simulate ocular systems. These models incorporated interdependent variables to enable analysis of how...
-
[14]
Below, we outline three critical stages that have collectively shaped the conceptual architecture of the virtual eye: Figure 1
The development of the virtual eye model progresses as its functionality and complexity increase . Below, we outline three critical stages that have collectively shaped the conceptual architecture of the virtual eye: Figure 1. Evolutions of the virtual eye. 2.1 Stage 1: Mechan...
-
[16]
To properly model such complex behaviors, many approaches should be explored and their merits carefully judged
Roadmap for virtual eye with AI: data, modeling, and interaction Building the virtual eye with AI is an interdisciplinary system engineering challenge. To properly model such complex behaviors, many approaches should be explored and their merits carefully judged. Here, t o bet...
-
[19]
divide and conquer
Challenges and recommendations Although the virtual eye holds enormous potential, realizing its full utility requires addressing a range of technical, ethical, and practical challenges. Many of these issues are shared with traditional deep learning systems but become significa...
-
[22]
The benchmarking framework itself should be adaptive and iterative, co- evolving with ongoing experimental findings and clinical feedback
-
[23]
Figure 3 outlines several envisioned applications
Application and future directions As data volumes grow and model architectures evolve, the Virtual Eye has the potential to revolutionize many aspects of ophthalmology. Figure 3 outlines several envisioned applications. While initial use cases may focus on improving medical ed...
-
[25]
Conclusion The concept of an AI-powered virtual eye embodies a convergence of ophthalmology, computer science, mechanical engineering, and biology. In this perspective, we traced the evolution from early computational eye models to the current landscape shaped by AI, and outli...
2015
-
[28]
By providing a generalizable backbone rather than a narrow single-purpose network, Foundation models achieved high accuracy in disease detection with minimal retraining. Multimodal foundation models like EyeFound, VisionFM and EyeCLIP, further expanded the ophthalmic modalitie...
-
[58]
By combining empirical data with prior knowledge, the Virtual Eye may one day model therapeutic r esponses before treatments are administered, enabling truly personalized medicine
The MorphoDiff framework 59, which generates realistic images of cellular responses to chemical or genetic perturbations, exemplifies how generative AI can simulate “what-if” scenarios. By combining empirical data with prior knowledge, the Virtual Eye may one day model therape...
-
[60]
Additionally, embodied AI expands the Virtual Eye’s capabilities by interfacing with robotics and diagnostic tools 61
Systems trained with heterogenous data will also enable multimodal interactions, linking images, annotations, tabular data, and explanations in both directions. Additionally, embodied AI expands the Virtual Eye’s capabilities by interfacing with robotics and diagnostic tools 6...
-
[63]
We recommend incorporating diversity -aware data curation, ongoing bias audits, and fairness metrics during model development. Additionally, given the sensitive nature of the biological and clinical data involved, robust data privacy protocols, secure federated learning framew...
-
[65]
This approach could pave the way for a common computational language that more effectively links fragmented data
This system, powered by self-supervised learning and context -aware algorithms, can construct a unified, scalable data representation. This approach could pave the way for a common computational language that more effectively links fragmented data. 4.4 Evaluation frameworks Tr...
-
[67]
This capability would not only allow for virtual validation of hypotheses but also foster hypothesis generation, guiding more targeted and efficient experimental designs
By simulating ocular systems at multiple scales, the platform could help identify potential causal relationships underlying observed phenotypes with quantified uncertainty. This capability would not only allow for virtual validation of hypotheses but also foster hypothesis gen...
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.