Pith. sign in

REVIEW 2 major objections 3 minor 62 references

The Philosophy and Physics of Duality

T0 review · 2 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read A unified 'geometric view' of theories aims to make sense of all dualities in physics, from position–momentum to gauge–gravity.

desk verdict A promising philosophy-of-physics monograph that I could not actually review, because the supplied full text was a different paper; judged from the abstract, it deserves peer review, not desk rejection. read the letter →

arxiv 2508.01616 v1 pith:EWAADNVU submitted 2025-08-03 physics.hist-ph cond-mat.stat-mechgr-qchep-th

classification physics.hist-phcond-mat.stat-mechgr-qchep-th MSC 81P0581T30
keywords dualitytheoreticalequivalencegeometricviewoftheoriesscientificrealismunderdeterminationKramers-Wanniergauge-gravityM-theory
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The monograph argues that the bewildering variety of dualities in physics—position–momentum, wave–particle, electric–magnetic, Kramers–Wannier, particle–soliton, string-theoretic, and gauge–gravity—are instances of one phenomenon that a 'geometric view of theories' can capture. On that view, a theory is not a set of axioms but a structured space of models, and a duality is an isomorphism between two such spaces that preserves all physical content while changing the presentation. This single idea is then used to treat several old philosophical problems in one stroke: when two theories really say the same thing, what a scientific realist can believe when dual theories disagree about ontology, whether underdetermination by data is real, why theory succession in the M-theory programme is rational, and how explanation and understanding work across dual descriptions. The book works from elementary examples up to the hole argument and string theory's black-hole microstate counting, and it proposes the geometric view as the unifying lesson.

What carries the argument

The central object is the geometric view itself: a theory is identified with a class of models together with its structure—its state space, quantities, and dynamics—rather than with a set of axioms. The load-bearing identity is the duality map, a bijection between the model spaces of two theories that preserves all physical quantities, which converts 'duality' into a precise mathematical notion: an isomorphism of geometric structures. This machinery is what allows the monograph to connect dualities to the familiar notion of symmetry, to distinguish gauge redundancies from physical content, and to give a unified treatment of theoretical equivalence, realism, underdetermination, theory succession, explanation, and understanding.

What would settle it

The framework would be falsified by exhibiting a case that physicists and philosophers would agree is a genuine duality but that does not correspond to an isomorphism of model spaces preserving all physical quantities, or by finding two theories that satisfy that isomorphism condition yet are not regarded as dual.

Watch

Extended reading notes

Core claim

The central claim is that dualities are a species of theoretical equivalence, and that the geometric view makes that species precise: two theories are dual when their model spaces can be mapped onto one another by an isomorphism that preserves the values of all physical quantities. This analysis is pressed into service on the hardest cases: the gauge–gravity duality, where a gravitational theory and a quantum field theory describe the same physics; the hole argument, where the geometric view supplies a principled distinction between gauge and physical structure; and black-hole entropy, where string theory's microstate counting becomes an illustration of duality at work. Read this way, dueling ontologies are not competing pictures of the world but two vocabularies for one structure, and the book argues that realism should attach to the shared structure rather than to either dressing.

Load-bearing premise

The book's general conclusions depend on the assumption that the specific dualities it studies—from position–momentum up to gauge–gravity duality—are representative enough of dualities in physics as a whole that a framework built from them will cover every future duality.

Editorial extensions

If this is right

  • The equivalence of two dual theories becomes decidable in principle by inspecting whether their model spaces are isomorphic, not by comparing axioms or vocabularies.
  • Scientific realism gains a precise target: the structure preserved by the duality map, so a realist can accept both descriptions without choosing one's ontology.
  • Underdetermination by data is defused for dual theories: since they are the same physics in different dress, no experiment can favor one over the other.
  • The M-theory programme is recast as rational theory succession in which dualities guide extension, rather than as competing paradigms replacing one another.
  • The hole argument's conclusion loses its sting because the geometric view provides a principled line between gauge and physical structure.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the geometric view is accepted, a natural research programme would be to look for new dualities by actively searching for isomorphisms between the model spaces of known theories, rather than discovering them case by case from matching spectra or correlation functions.
  • The framework implies that dual pairs should be treated as a single theory, which would change how the landscape of string vacua is counted and compared.
  • An editorial note: the full text supplied with this record is a different manuscript about a biomedical vision-language model (LLaDA-MedV), so this summary follows the title, abstract, and reader notes as the intended paper; the record's mismatch should be resolved.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 3 minor

Summary. The abstract of arXiv:2508.01616 announces a philosophy-of-physics monograph on dualities. Part I is said to conceptualize dualities and to discuss simple examples such as position-momentum, wave-particle, electric-magnetic, and Kramers-Wannier dualities. Part II covers advanced examples, including particle-soliton dualities, electric-magnetic dualities in quantum field theories, string-theoretic dualities, gauge-gravity duality, the hole argument, and black-hole microstate counting. Part III is described as addressing theoretical equivalence, scientific realism and underdetermination, theory succession and the M-theory programme, explanation, and scientific understanding, while proposing a 'geometric view of theories.' The supplied full text, however, is arXiv:2508.01617v2, a biomedical computer-vision paper on LLaDA-MedV, not the target monograph. As a result, the monograph's definitions, examples, and arguments are not available for review.

Significance. If the monograph substantiates its abstract's claims, it could be a significant contribution to the philosophy of physics: a unified 'geometric view of theories' that covers the main dualities and connects them to standard philosophical debates would be of interest to philosophers and theoretically inclined physicists. The abstract's broad scope, with self-contained treatments of both technical examples and philosophical topics, is a stated strength. However, no part of the argument is visible in the supplied materials, so the significance cannot be confirmed. There are no machine-checked proofs, reproducible code, or parameter-free derivations to evaluate; this is a philosophical monograph whose merits, if any, lie in its definitions and argumentation, which are absent here. The reader's concern about the representativeness of the chosen examples is plausible but cannot be adjudicated without the full text.

major comments (2)
  1. [Full Text / entire manuscript] The supplied full text does not correspond to the target paper; it is LLaDA-MedV, a biomedical computer-vision paper, not the philosophy-of-physics monograph described in the abstract. None of the monograph's definitions, examples, or Part III arguments are present. The central claim—that a single 'geometric view of theories' accommodates all major dualities—is therefore unverifiable from the submitted materials. This is not a judgment about the truth of the claim; it is a request for the actual manuscript before any content-level review can begin.
  2. [Abstract / Part III] The general philosophical conclusions announced for Part III depend on generalizing from the seven example families listed in the abstract, but the abstract does not state how 'duality' is defined, what counts as a 'geometric view of theories,' or how the examples support the general claims. As presented, the representative-sample assumption is undefended; the defense may well occur in Parts I–III, but those parts are not supplied, so the claim cannot be evaluated.
minor comments (3)
  1. [Abstract] The abstract uses both 'under-determination' (with hyphen) and 'underdetermination' (without hyphen); please choose one spelling and use it consistently.
  2. [Abstract] The text opens by calling the work 'This monograph' but does not indicate whether it is a book manuscript, a long-form article, or a preprint of a target article; stating the intended genre and providing a brief chapter outline would help readers navigate the announced structure.
  3. [Abstract] The abstract says the book is aimed at students and researchers with an interest in the physical examples and philosophical questions, but it does not specify any assumed philosophical background; a sentence on prerequisites would be helpful.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: only the monograph abstract is available, and the supplied full text is a different arXiv paper, so there is no derivation chain to audit.

full rationale

The submission for circularity analysis is arXiv:2508.01616, a philosophy of physics monograph whose abstract is quoted, but the body text provided is arXiv:2508.01617v2, an unrelated computer-vision paper. There are therefore no equations, fitted parameters, or derivation steps from the monograph that could be reduced to inputs. The abstract's only constructive claim is that Part III 'proposes a view of scientific theories that it dubs the geometric view of theories'; without the definitions of 'duality' and 'geometric view' and the argument linking the Part I/II examples to the Part III conclusions, no specific circular step can be quoted or exhibited. Under the hard rule that circularity may only be claimed when the paper's own text exhibits the reduction, the correct finding is no significant circularity (score 0). This is a non-finding due to unavailable evidence, not an endorsement of the monograph's philosophical arguments.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

No free parameters or invented entities apply to a philosophical monograph. Axioms cannot be audited from the abstract alone; the full text was not available.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Philosophy and Physics of Duality." pith.science (2026). https://pith.science/paper/EWAADNVU

@misc{pith2026250801616,
  author       = {Pith},
  title        = {Pith review of: The Philosophy and Physics of Duality},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/EWAADNVU}},
  note         = {Machine review of arXiv:2508.01616}
}
read the original abstract

This monograph discusses dualities in physics: what dualities are, their main examples--from quantum mechanics and electrodynamics to statistical mechanics, quantum field theory and string theory--and the philosophical questions they raise. Part I first conceptualises dualities and discusses their main roles and themes, including how they are related to familiar notions like symmetry and interpretation. It also discusses the main simple examples of dualities: position-momentum, wave-particle, electric-magnetic, and Kramers-Wannier dualities. Part II discusses advanced examples and their inter-relations: particle-soliton dualities, electric-magnetic dualities in quantum field theories, dualities in string theory, and gauge-gravity duality. This Part ends with discussions of the hole argument, and how string theory counts the microstates of a black hole. Part III is an in-depth discussion of general philosophical issues on which dualities bear: theoretical equivalence (two theories 'saying the same thing, in different words'), scientific realism and the under-determination of theories by data, theory succession and the M-theory programme, explanation, and scientific understanding. It proposes a view of scientific theories that it dubs 'the geometric view of theories'. The book's treatment of the examples is at the advanced undergraduate and graduate level, starting from elementary and progressing to more advanced examples. The discussions of philosophical topics, such as referential semantics, theoretical equivalence, scientific realism and scientific understanding, are both self-contained and in-depth. Thus the book is aimed at students and researchers with an interest in the physical examples and philosophical questions about dualities, and also in how physics and philosophy can fruitfully interact with each other.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

62 extracted references · 27 canonical work pages

  1. [1]

    Structured denoising dif- fusion models in discrete state-spaces.Advances in neural information processing systems, 34:17981–17993, 2021

    Jacob Austin, Daniel D Johnson, Jonathan Ho, Daniel Tar- low, and Rianne Van Den Berg. Structured denoising dif- fusion models in discrete state-spaces.Advances in neural information processing systems, 34:17981–17993, 2021. 3

  2. [2]

    Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2023

    Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2023. 5, 13

  3. [3]

    Learning to exploit temporal structure for biomed- ical vision-language processing

    Shruthi Bannur, Stephanie Hyland, Qianchu Liu, Fernando Perez-Garcia, Maximilian Ilse, Daniel C Castro, Benedikt Boecking, Harshita Sharma, Kenza Bouzid, Anja Thieme, et al. Learning to exploit temporal structure for biomed- ical vision-language processing. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15016–15...

  4. [4]

    Vision–language model for visual question answering in medical imagery.Bioengineering, 10 (3):380, 2023

    Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Laila Bash- mal, and Mansour Zuair. Vision–language model for visual question answering in medical imagery.Bioengineering, 10 (3):380, 2023. 6

  5. [5]

    Maskgit: Masked generative image transformer

    Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T Freeman. Maskgit: Masked generative image transformer. InProceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 11315–11325, 2022. 4, 12

  6. [6]

    Analog bits: Generating discrete data using diffusion models with self-conditioning.arXiv preprint arXiv:2208.04202, 2022

    Ting Chen, Ruixiang Zhang, and Geoffrey Hinton. Analog bits: Generating discrete data using diffusion models with self-conditioning.arXiv preprint arXiv:2208.04202, 2022. 2

  7. [7]

    Visual prompt engineering for vision language models in radiology

    Stefan Denner, Markus Bujotzek, Dimitrios Bounias, David Zimmerer, Raphael Stock, and Klaus Maier-Hein. Visual prompt engineering for vision language models in radiology. arXiv preprint arXiv:2408.15802, 2024. 2

  8. [8]

    Sedigheh Eslami, Christoph Meinel, and Gerard De Melo. Pubmedclip: How much does clip benefit visual question answering in the medical domain? InFindings of the As- sociation for Computational Linguistics: EACL 2023, pages 1181–1193, 2023. 6

Show all 62 references
  1. [9]

    Diffuseq: Sequence to sequence text generation with diffusion models.arXiv preprint arXiv:2210.08933, 2022

    Shansan Gong, Mukai Li, Jiangtao Feng, Zhiyong Wu, and LingPeng Kong. Diffuseq: Sequence to sequence text generation with diffusion models.arXiv preprint arXiv:2210.08933, 2022. 2

  2. [10]

    Scaling diffusion language models via adaptation from autoregressive models.arXiv preprint arXiv:2410.17891, 2024

    Shansan Gong, Shivam Agarwal, Yizhe Zhang, Jiacheng Ye, Lin Zheng, Mukai Li, Chenxin An, Peilin Zhao, Wei Bi, Jiawei Han, et al. Scaling diffusion language models via adaptation from autoregressive models.arXiv preprint arXiv:2410.17891, 2024. 2

  3. [11]

    The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024

    Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Ab- hinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024. 2

  4. [12]

    Prompting medical large vision-language models to diagnose pathologies by vi- sual question answering.arXiv preprint arXiv:2407.21368,

    Danfeng Guo and Demetri Terzopoulos. Prompting medical large vision-language models to diagnose pathologies by vi- sual question answering.arXiv preprint arXiv:2407.21368,

  5. [13]

    Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948, 2025

    Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948, 2025. 5

  6. [14]

    Pathvqa: 30000+ questions for medical visual question answering.arXiv preprint arXiv:2003.10286, 2020

    Xuehai He, Yichen Zhang, Luntian Mou, Eric Xing, and Pengtao Xie. Pathvqa: 30000+ questions for medical visual question answering.arXiv preprint arXiv:2003.10286, 2020. 4, 13, 14

  7. [15]

    Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

    Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020. 1

  8. [16]

    Gpt-4 as an effec- tive zero-shot evaluator for scientific figure captions.arXiv preprint arXiv:2310.15405, 2023

    Ting-Yao Hsu, Chieh-Yang Huang, Ryan Rossi, Sungchul Kim, C Lee Giles, and Ting-Hao K Huang. Gpt-4 as an effec- tive zero-shot evaluator for scientific figure captions.arXiv preprint arXiv:2310.15405, 2023. 5

  9. [17]

    Med-r1: Reinforcement learning for general- izable medical reasoning in vision-language models.arXiv preprint arXiv:2503.13939, 2025

    Yuxiang Lai, Jike Zhong, Ming Li, Shitian Zhao, and Xi- aofeng Yang. Med-r1: Reinforcement learning for general- izable medical reasoning in vision-language models.arXiv preprint arXiv:2503.13939, 2025. 2

  10. [18]

    A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018

    Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018. 4, 13, 14

  11. [19]

    Llava-med: Training a large language- and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36:28541–28564,

    Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. Llava-med: Training a large language- and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36:28541–28564,

  12. [20]

    Self-supervised vision-language pretraining for me- dial visual question answering

    Pengfei Li, Gang Liu, Lin Tan, Jinying Liao, and Shenjun Zhong. Self-supervised vision-language pretraining for me- dial visual question answering. In2023 IEEE 20th Interna- tional Symposium on Biomedical Imaging (ISBI), pages 1–5. IEEE, 2023. 6

  13. [21]

    Diffusion-lm improves control- lable text generation.Advances in neural information pro- cessing systems, 35:4328–4343, 2022

    Xiang Li, John Thickstun, Ishaan Gulrajani, Percy S Liang, and Tatsunori B Hashimoto. Diffusion-lm improves control- lable text generation.Advances in neural information pro- cessing systems, 35:4328–4343, 2022. 2

  14. [22]

    Pmc-clip: Con- trastive language-image pre-training using biomedical docu- ments

    Weixiong Lin, Ziheng Zhao, Xiaoman Zhang, Chaoyi Wu, Ya Zhang, Yanfeng Wang, and Weidi Xie. Pmc-clip: Con- trastive language-image pre-training using biomedical docu- ments. InInternational Conference on Medical Image Com- puting and Computer-Assisted Intervention, pages 525–5...

  15. [23]

    Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering

    Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma, Yan Yang, and Xiao-Ming Wu. Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering. In 2021 IEEE 18th international symposium on biomedical imaging (ISBI), pages 1650–1654. IEEE, 2021. 4, 13, 14

  16. [24]

    Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023. 3, 5, 13

  17. [25]

    Improved baselines with visual instruction tuning

    Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. InPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 26296–26306, 2024. 3

  18. [26]

    G-eval: Nlg evaluation us- ing gpt-4 with better human alignment.arXiv preprint arXiv:2303.16634, 2023

    Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G-eval: Nlg evaluation us- ing gpt-4 with better human alignment.arXiv preprint arXiv:2303.16634, 2023. 5

  19. [27]

    Q2atransformer: Improving medical vqa via an answer querying decoder

    Yunyi Liu, Zhanyu Wang, Dong Xu, and Luping Zhou. Q2atransformer: Improving medical vqa via an answer querying decoder. InInternational conference on infor- mation processing in medical imaging, pages 445–456. Springer, 2023. 6

  20. [28]

    Discrete diffusion modeling by estimating the ratios of the data distri- bution.arXiv preprint arXiv:2310.16834, 2023

    Aaron Lou, Chenlin Meng, and Stefano Ermon. Discrete diffusion modeling by estimating the ratios of the data distri- bution.arXiv preprint arXiv:2310.16834, 2023. 1, 3

  21. [29]

    Biomedgpt: Open mul- timodal generative pre-trained transformer for biomedicine

    Yizhen Luo, Jiahuan Zhang, Siqi Fan, Kai Yang, Yushuai Wu, Mu Qiao, and Zaiqing Nie. Biomedgpt: Open mul- timodal generative pre-trained transformer for biomedicine. arXiv preprint arXiv:2308.09442, 2023. 2

  22. [30]

    Tess: Text-to-text self-conditioned simplex diffu- sion.arXiv preprint arXiv:2305.08379, 2023

    Rabeeh Karimi Mahabadi, Hamish Ivison, Jaesung Tae, James Henderson, Iz Beltagy, Matthew E Peters, and Arman Cohan. Tess: Text-to-text self-conditioned simplex diffu- sion.arXiv preprint arXiv:2305.08379, 2023. 2

  23. [31]

    Llama-3.2-11B-Vision: A Multimodal Vision– Language LLM

    Meta AI. Llama-3.2-11B-Vision: A Multimodal Vision– Language LLM. Model card via Meta AI, 2024. Version released September 25, 2024; instruction-tuned for image reasoning, captioning, and VQA with 10.6B parameters. 5, 13

  24. [32]

    Med-flamingo: a multimodal medical few-shot learner

    Michael Moor, Qian Huang, Shirley Wu, Michihiro Ya- sunaga, Yash Dalmia, Jure Leskovec, Cyril Zakka, Ed- uardo Pontes Reis, and Pranav Rajpurkar. Med-flamingo: a multimodal medical few-shot learner. InMachine Learn- ing for Health (ML4H), pages 353–367. PMLR, 2023. 2, 5, 13

  25. [33]

    Large language diffusion models.arXiv preprint arXiv:2502.09992, 2025

    Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, and Chongxuan Li. Large language diffusion models.arXiv preprint arXiv:2502.09992, 2025. 1, 2, 3, 4, 8, 12, 13

  26. [34]

    Introducing gpt-4.1 in the api.https : / / openai

    OpenAI. Introducing gpt-4.1 in the api.https : / / openai . com / index / gpt - 4 - 1/, 2025. Includes GPT-4.1, GPT-4.1 mini, and GPT-4.1 nano. 5, 12

  27. [35]

    , and Barret Zoph

    OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, . . . , and Barret Zoph. Gpt-4 technical report.https://arxiv.org/abs/2303. 08774, 2023. arXiv:2303.08774, submitted March 15 2023. 4

  28. [36]

    Your absorbing dis- crete diffusion secretly models the conditional distributions of clean data.arXiv preprint arXiv:2406.03736, 2024

    Jingyang Ou, Shen Nie, Kaiwen Xue, Fengqi Zhu, Jiacheng Sun, Zhenguo Li, and Chongxuan Li. Your absorbing dis- crete diffusion secretly models the conditional distributions of clean data.arXiv preprint arXiv:2406.03736, 2024. 2, 3

  29. [37]

    Training language models to follow instructions with human feedback.Ad- vances in neural information processing systems, 35:27730– 27744, 2022

    Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Car- roll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback.Ad- vances in neural information processing systems, 35:2...

  30. [38]

    Medvlm-r1: Incentivizing medical reasoning ca- pability of vision-language models (vlms) via reinforcement learning.arXiv preprint arXiv:2502.19634, 2025

    Jiazhen Pan, Che Liu, Junde Wu, Fenglin Liu, Jiayuan Zhu, Hongwei Bran Li, Chen Chen, Cheng Ouyang, and Daniel Rueckert. Medvlm-r1: Incentivizing medical reasoning ca- pability of vision-language models (vlms) via reinforcement learning.arXiv preprint arXiv:2502.19634, 2025. 2, 5, 13

  31. [39]

    Learning transferable visual models from natural language supervi- sion

    Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. InInternational conference on machine learning, p...

  32. [40]

    Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning

    Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning. InProceed- ings of the international conference for high performance computing, networking, storage and analysis, pages 1–14,

  33. [41]

    Categorical sdes with simplex diffusion.arXiv preprint arXiv:2210.14784, 2022

    Pierre H Richemond, Sander Dieleman, and Arnaud Doucet. Categorical sdes with simplex diffusion.arXiv preprint arXiv:2210.14784, 2022. 2

  34. [42]

    Simple and effective masked dif- fusion language models.Advances in Neural Information Processing Systems, 37:130136–130184, 2024

    Subham Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin Chiu, Alexander Rush, and V olodymyr Kuleshov. Simple and effective masked dif- fusion language models.Advances in Neural Information Processing Systems, 37:130136–130184, 2024. 1

  35. [43]

    Simplified and generalized masked diffu- sion for discrete data.Advances in neural information pro- cessing systems, 37:103131–103167, 2024

    Jiaxin Shi, Kehang Han, Zhe Wang, Arnaud Doucet, and Michalis Titsias. Simplified and generalized masked diffu- sion for discrete data.Advances in neural information pro- cessing systems, 37:103131–103167, 2024. 1

  36. [44]

    Denoising diffusion implicit models.arXiv preprint arXiv:2010.02502, 2020

    Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models.arXiv preprint arXiv:2010.02502, 2020. 1

  37. [45]

    Score-based generative modeling through stochastic differential equa- tions.arXiv preprint arXiv:2011.13456, 2020

    Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equa- tions.arXiv preprint arXiv:2011.13456, 2020. 1

  38. [46]

    Self- conditioned embedding diffusion for text generation.arXiv preprint arXiv:2211.04236, 2022

    Robin Strudel, Corentin Tallec, Florent Altch ´e, Yilun Du, Yaroslav Ganin, Arthur Mensch, Will Grathwohl, Niko- lay Savinov, Sander Dieleman, Laurent Sifre, et al. Self- conditioned embedding diffusion for text generation.arXiv preprint arXiv:2211.04236, 2022. 2

  39. [47]

    Llama: Open and efficient foundation language models

    Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth´ee Lacroix, Baptiste Rozi`ere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. 3

  40. [48]

    Siglip 2: Multilingual vision-language en- coders with improved semantic understanding, localization, and dense features.arXiv preprint arXiv:2502.14786, 2025

    Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muham- mad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, et al. Siglip 2: Multilingual vision-language en- coders with improved semantic understanding, localization, ...

  41. [49]

    Open- ended medical visual question answering through prefix tun- ing of language models

    Tom Van Sonsbeek, Mohammad Mahdi Derakhshani, Ivona Najdenkoska, Cees GM Snoek, and Marcel Worring. Open- ended medical visual question answering through prefix tun- ing of language models. InInternational Conference on Medical Image Computing and Computer-Assisted Interven- t...

  42. [50]

    A connection between score matching and denoising autoencoders.Neural computation, 23(7):1661– 1674, 2011

    Pascal Vincent. A connection between score matching and denoising autoencoders.Neural computation, 23(7):1661– 1674, 2011. 3

  43. [51]

    How does diverse interpretability of textual prompts impact med- ical vision-language zero-shot tasks?arXiv preprint arXiv:2409.00543, 2024

    Sicheng Wang, Che Liu, and Rossella Arcucci. How does diverse interpretability of textual prompts impact med- ical vision-language zero-shot tasks?arXiv preprint arXiv:2409.00543, 2024. 2

  44. [52]

    Interactive computer-aided diag- nosis on medical image using large language models.Com- munications Engineering, 3(1):133, 2024

    Sheng Wang, Zihao Zhao, Xi Ouyang, Tianming Liu, Qian Wang, and Dinggang Shen. Interactive computer-aided diag- nosis on medical image using large language models.Com- munications Engineering, 3(1):133, 2024. 2

  45. [53]

    Medclip: Contrastive learning from unpaired medi- cal images and text

    Zifeng Wang, Zhenbang Wu, Dinesh Agarwal, and Jimeng Sun. Medclip: Contrastive learning from unpaired medi- cal images and text. InProceedings of the Conference on Empirical Methods in Natural Language Processing. Con- ference on Empirical Methods in Natural Language Process- ...

  46. [54]

    Gpt-4 as evaluator: Evaluating large language models on pest management in agriculture.arXiv preprint arXiv:2403.11858, 2024

    Shanglong Yang, Zhipeng Yuan, Shunbao Li, Ruoling Peng, Kang Liu, and Po Yang. Gpt-4 as evaluator: Evaluating large language models on pest management in agriculture.arXiv preprint arXiv:2403.11858, 2024. 5

  47. [55]

    Dinoiser: Diffused conditional se- quence learning by manipulating noises.arXiv preprint arXiv:2302.10025, 2023

    Jiasheng Ye, Zaixiang Zheng, Yu Bao, Lihua Qian, and Mingxuan Wang. Dinoiser: Diffused conditional se- quence learning by manipulating noises.arXiv preprint arXiv:2302.10025, 2023. 2

  48. [56]

    Llada-v: Large language diffusion models with visual instruction tuning

    Zebin You, Shen Nie, Xiaolu Zhang, Jun Hu, Jun Zhou, Zhiwu Lu, Ji-Rong Wen, and Chongxuan Li. Llada-v: Large language diffusion models with visual instruction tuning. arXiv preprint arXiv:2505.16933, 2025. 2, 3, 4, 5, 12, 13

  49. [57]

    Biomedgpt: A unified and generalist biomed- ical generative pre-trained transformer for vision, language, and multimodal tasks.arXiv e-prints, pages arXiv–2305,

    Kai Zhang, Jun Yu, Eashan Adhikarla, Rong Zhou, Zhiling Yan, Yixin Liu, Zhengliang Liu, Lifang He, Brian Davison, Xiang Li, et al. Biomedgpt: A unified and generalist biomed- ical generative pre-trained transformer for vision, language, and multimodal tasks.arXiv e-prints, pag...

  50. [58]

    Large-scale domain-specific pre- training for biomedical vision-language processing.arXiv preprint arXiv:2303.00915, 2(3):6, 2023

    Sheng Zhang, Yanbo Xu, Naoto Usuyama, Jaspreet Bagga, Robert Tinn, Sam Preston, Rajesh Rao, Mu Wei, Naveen Valluri, Cliff Wong, et al. Large-scale domain-specific pre- training for biomedical vision-language processing.arXiv preprint arXiv:2303.00915, 2(3):6, 2023. 6

  51. [59]

    Toward effective reinforcement learning fine-tuning for med- ical vqa in vision-language models.arXiv e-prints, pages arXiv–2505, 2025

    Wenhui Zhu, Xuanzhao Dong, Xin Li, Peijie Qiu, Xiwen Chen, Abolfazl Razi, Aris Sotiras, Yi Su, and Yalin Wang. Toward effective reinforcement learning fine-tuning for med- ical vqa in vision-language models.arXiv e-prints, pages arXiv–2505, 2025. 5, 13

  52. [60]

    Retinalgpt: A retinal clinical preference conversational assistant powered by large vision- language models.arXiv e-prints, pages arXiv–2503, 2025

    Wenhui Zhu, Xin Li, Xiwen Chen, Peijie Qiu, Vamsi Kr- ishna Vasa, Xuanzhao Dong, Yanxi Chen, Natasha Lepore, Oana Dumitrascu, Yi Su, et al. Retinalgpt: A retinal clinical preference conversational assistant powered by large vision- language models.arXiv e-prints, pages arXiv–2...

  53. [256]

    Unless otherwise specified, this setting is used consistently in all detailed analyses

    However, for downstream biomedical VQA tasks, we disable the semi-autoregressive mechanism by settingL= B=Z= 64. Unless otherwise specified, this setting is used consistently in all detailed analyses. Baseline Model Configuration.We compare LLaDA- MedV against nine baseline vi...

  54. [2023]

    2, 3, 4, 5, 6, 12, 13, 14, 15

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.