REVIEW 2 major objections 3 minor 62 references
The Philosophy and Physics of Duality
T0 review · 2 major / 3 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A unified 'geometric view' of theories aims to make sense of all dualities in physics, from position–momentum to gauge–gravity.
desk verdict A promising philosophy-of-physics monograph that I could not actually review, because the supplied full text was a different paper; judged from the abstract, it deserves peer review, not desk rejection. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the geometric view itself: a theory is identified with a class of models together with its structure—its state space, quantities, and dynamics—rather than with a set of axioms. The load-bearing identity is the duality map, a bijection between the model spaces of two theories that preserves all physical quantities, which converts 'duality' into a precise mathematical notion: an isomorphism of geometric structures. This machinery is what allows the monograph to connect dualities to the familiar notion of symmetry, to distinguish gauge redundancies from physical content, and to give a unified treatment of theoretical equivalence, realism, underdetermination, theory succession, explanation, and understanding.
What would settle it
The framework would be falsified by exhibiting a case that physicists and philosophers would agree is a genuine duality but that does not correspond to an isomorphism of model spaces preserving all physical quantities, or by finding two theories that satisfy that isomorphism condition yet are not regarded as dual.
Extended reading notes
Core claim
The central claim is that dualities are a species of theoretical equivalence, and that the geometric view makes that species precise: two theories are dual when their model spaces can be mapped onto one another by an isomorphism that preserves the values of all physical quantities. This analysis is pressed into service on the hardest cases: the gauge–gravity duality, where a gravitational theory and a quantum field theory describe the same physics; the hole argument, where the geometric view supplies a principled distinction between gauge and physical structure; and black-hole entropy, where string theory's microstate counting becomes an illustration of duality at work. Read this way, dueling ontologies are not competing pictures of the world but two vocabularies for one structure, and the book argues that realism should attach to the shared structure rather than to either dressing.
Load-bearing premise
The book's general conclusions depend on the assumption that the specific dualities it studies—from position–momentum up to gauge–gravity duality—are representative enough of dualities in physics as a whole that a framework built from them will cover every future duality.
Editorial extensions
If this is right
- The equivalence of two dual theories becomes decidable in principle by inspecting whether their model spaces are isomorphic, not by comparing axioms or vocabularies.
- Scientific realism gains a precise target: the structure preserved by the duality map, so a realist can accept both descriptions without choosing one's ontology.
- Underdetermination by data is defused for dual theories: since they are the same physics in different dress, no experiment can favor one over the other.
- The M-theory programme is recast as rational theory succession in which dualities guide extension, rather than as competing paradigms replacing one another.
- The hole argument's conclusion loses its sting because the geometric view provides a principled line between gauge and physical structure.
Reading between the lines
- If the geometric view is accepted, a natural research programme would be to look for new dualities by actively searching for isomorphisms between the model spaces of known theories, rather than discovering them case by case from matching spectra or correlation functions.
- The framework implies that dual pairs should be treated as a single theory, which would change how the landscape of string vacua is counted and compared.
- An editorial note: the full text supplied with this record is a different manuscript about a biomedical vision-language model (LLaDA-MedV), so this summary follows the title, abstract, and reader notes as the intended paper; the record's mismatch should be resolved.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The abstract of arXiv:2508.01616 announces a philosophy-of-physics monograph on dualities. Part I is said to conceptualize dualities and to discuss simple examples such as position-momentum, wave-particle, electric-magnetic, and Kramers-Wannier dualities. Part II covers advanced examples, including particle-soliton dualities, electric-magnetic dualities in quantum field theories, string-theoretic dualities, gauge-gravity duality, the hole argument, and black-hole microstate counting. Part III is described as addressing theoretical equivalence, scientific realism and underdetermination, theory succession and the M-theory programme, explanation, and scientific understanding, while proposing a 'geometric view of theories.' The supplied full text, however, is arXiv:2508.01617v2, a biomedical computer-vision paper on LLaDA-MedV, not the target monograph. As a result, the monograph's definitions, examples, and arguments are not available for review.
Significance. If the monograph substantiates its abstract's claims, it could be a significant contribution to the philosophy of physics: a unified 'geometric view of theories' that covers the main dualities and connects them to standard philosophical debates would be of interest to philosophers and theoretically inclined physicists. The abstract's broad scope, with self-contained treatments of both technical examples and philosophical topics, is a stated strength. However, no part of the argument is visible in the supplied materials, so the significance cannot be confirmed. There are no machine-checked proofs, reproducible code, or parameter-free derivations to evaluate; this is a philosophical monograph whose merits, if any, lie in its definitions and argumentation, which are absent here. The reader's concern about the representativeness of the chosen examples is plausible but cannot be adjudicated without the full text.
major comments (2)
- [Full Text / entire manuscript] The supplied full text does not correspond to the target paper; it is LLaDA-MedV, a biomedical computer-vision paper, not the philosophy-of-physics monograph described in the abstract. None of the monograph's definitions, examples, or Part III arguments are present. The central claim—that a single 'geometric view of theories' accommodates all major dualities—is therefore unverifiable from the submitted materials. This is not a judgment about the truth of the claim; it is a request for the actual manuscript before any content-level review can begin.
- [Abstract / Part III] The general philosophical conclusions announced for Part III depend on generalizing from the seven example families listed in the abstract, but the abstract does not state how 'duality' is defined, what counts as a 'geometric view of theories,' or how the examples support the general claims. As presented, the representative-sample assumption is undefended; the defense may well occur in Parts I–III, but those parts are not supplied, so the claim cannot be evaluated.
minor comments (3)
- [Abstract] The abstract uses both 'under-determination' (with hyphen) and 'underdetermination' (without hyphen); please choose one spelling and use it consistently.
- [Abstract] The text opens by calling the work 'This monograph' but does not indicate whether it is a book manuscript, a long-form article, or a preprint of a target article; stating the intended genre and providing a brief chapter outline would help readers navigate the announced structure.
- [Abstract] The abstract says the book is aimed at students and researchers with an interest in the physical examples and philosophical questions, but it does not specify any assumed philosophical background; a sentence on prerequisites would be helpful.
Circularity Check
No circularity found: only the monograph abstract is available, and the supplied full text is a different arXiv paper, so there is no derivation chain to audit.
full rationale
The submission for circularity analysis is arXiv:2508.01616, a philosophy of physics monograph whose abstract is quoted, but the body text provided is arXiv:2508.01617v2, an unrelated computer-vision paper. There are therefore no equations, fitted parameters, or derivation steps from the monograph that could be reduced to inputs. The abstract's only constructive claim is that Part III 'proposes a view of scientific theories that it dubs the geometric view of theories'; without the definitions of 'duality' and 'geometric view' and the argument linking the Part I/II examples to the Part III conclusions, no specific circular step can be quoted or exhibited. Under the hard rule that circularity may only be claimed when the paper's own text exhibits the reduction, the correct finding is no significant circularity (score 0). This is a non-finding due to unavailable evidence, not an endorsement of the monograph's philosophical arguments.
Assumptions & free parameters
Cite this review
Pith. "Pith review of The Philosophy and Physics of Duality." pith.science (2026). https://pith.science/paper/EWAADNVU
@misc{pith2026250801616,
author = {Pith},
title = {Pith review of: The Philosophy and Physics of Duality},
year = {2026},
howpublished = {\url{https://pith.science/paper/EWAADNVU}},
note = {Machine review of arXiv:2508.01616}
}
read the original abstract
This monograph discusses dualities in physics: what dualities are, their main examples--from quantum mechanics and electrodynamics to statistical mechanics, quantum field theory and string theory--and the philosophical questions they raise. Part I first conceptualises dualities and discusses their main roles and themes, including how they are related to familiar notions like symmetry and interpretation. It also discusses the main simple examples of dualities: position-momentum, wave-particle, electric-magnetic, and Kramers-Wannier dualities. Part II discusses advanced examples and their inter-relations: particle-soliton dualities, electric-magnetic dualities in quantum field theories, dualities in string theory, and gauge-gravity duality. This Part ends with discussions of the hole argument, and how string theory counts the microstates of a black hole. Part III is an in-depth discussion of general philosophical issues on which dualities bear: theoretical equivalence (two theories 'saying the same thing, in different words'), scientific realism and the under-determination of theories by data, theory succession and the M-theory programme, explanation, and scientific understanding. It proposes a view of scientific theories that it dubs 'the geometric view of theories'. The book's treatment of the examples is at the advanced undergraduate and graduate level, starting from elementary and progressing to more advanced examples. The discussions of philosophical topics, such as referential semantics, theoretical equivalence, scientific realism and scientific understanding, are both self-contained and in-depth. Thus the book is aimed at students and researchers with an interest in the physical examples and philosophical questions about dualities, and also in how physics and philosophy can fruitfully interact with each other.
Reference graph
Works this paper leans on
-
[1]
Jacob Austin, Daniel D Johnson, Jonathan Ho, Daniel Tar- low, and Rianne Van Den Berg. Structured denoising dif- fusion models in discrete state-spaces.Advances in neural information processing systems, 34:17981–17993, 2021. 3
work page 2021
-
[2]
Jinze Bai, Shuai Bai, Shusheng Yang, Shijie Wang, Sinan Tan, Peng Wang, Junyang Lin, Chang Zhou, and Jingren Zhou. Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond.arXiv preprint arXiv:2308.12966, 2023. 5, 13
arXiv 2023
-
[3]
Learning to exploit temporal structure for biomed- ical vision-language processing
Shruthi Bannur, Stephanie Hyland, Qianchu Liu, Fernando Perez-Garcia, Maximilian Ilse, Daniel C Castro, Benedikt Boecking, Harshita Sharma, Kenza Bouzid, Anja Thieme, et al. Learning to exploit temporal structure for biomed- ical vision-language processing. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 15016–15...
work page 2023
-
[4]
Yakoub Bazi, Mohamad Mahmoud Al Rahhal, Laila Bash- mal, and Mansour Zuair. Vision–language model for visual question answering in medical imagery.Bioengineering, 10 (3):380, 2023. 6
work page 2023
-
[5]
Maskgit: Masked generative image transformer
Huiwen Chang, Han Zhang, Lu Jiang, Ce Liu, and William T Freeman. Maskgit: Masked generative image transformer. InProceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, pages 11315–11325, 2022. 4, 12
work page 2022
-
[6]
Ting Chen, Ruixiang Zhang, and Geoffrey Hinton. Analog bits: Generating discrete data using diffusion models with self-conditioning.arXiv preprint arXiv:2208.04202, 2022. 2
arXiv 2022
-
[7]
Visual prompt engineering for vision language models in radiology
Stefan Denner, Markus Bujotzek, Dimitrios Bounias, David Zimmerer, Raphael Stock, and Klaus Maier-Hein. Visual prompt engineering for vision language models in radiology. arXiv preprint arXiv:2408.15802, 2024. 2
arXiv 2024
-
[8]
Sedigheh Eslami, Christoph Meinel, and Gerard De Melo. Pubmedclip: How much does clip benefit visual question answering in the medical domain? InFindings of the As- sociation for Computational Linguistics: EACL 2023, pages 1181–1193, 2023. 6
work page 2023
Show all 62 references
-
[9]
Diffuseq: Sequence to sequence text generation with diffusion models.arXiv preprint arXiv:2210.08933, 2022
Shansan Gong, Mukai Li, Jiangtao Feng, Zhiyong Wu, and LingPeng Kong. Diffuseq: Sequence to sequence text generation with diffusion models.arXiv preprint arXiv:2210.08933, 2022. 2
2022 arXiv
-
[10]
Scaling diffusion language models via adaptation from autoregressive models.arXiv preprint arXiv:2410.17891, 2024
Shansan Gong, Shivam Agarwal, Yizhe Zhang, Jiacheng Ye, Lin Zheng, Mukai Li, Chenxin An, Peilin Zhao, Wei Bi, Jiawei Han, et al. Scaling diffusion language models via adaptation from autoregressive models.arXiv preprint arXiv:2410.17891, 2024. 2
-
[11]
The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024
Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Ab- hinav Pandey, Abhishek Kadian, Ahmad Al-Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, et al. The llama 3 herd of models.arXiv preprint arXiv:2407.21783, 2024. 2
2024 arXiv
-
[12]
Prompting medical large vision-language models to diagnose pathologies by vi- sual question answering.arXiv preprint arXiv:2407.21368,
Danfeng Guo and Demetri Terzopoulos. Prompting medical large vision-language models to diagnose pathologies by vi- sual question answering.arXiv preprint arXiv:2407.21368,
-
[13]
Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948, 2025
Daya Guo, Dejian Yang, Haowei Zhang, Junxiao Song, Ruoyu Zhang, Runxin Xu, Qihao Zhu, Shirong Ma, Peiyi Wang, Xiao Bi, et al. Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv preprint arXiv:2501.12948, 2025. 5
2025 arXiv
-
[14]
Pathvqa: 30000+ questions for medical visual question answering.arXiv preprint arXiv:2003.10286, 2020
Xuehai He, Yichen Zhang, Luntian Mou, Eric Xing, and Pengtao Xie. Pathvqa: 30000+ questions for medical visual question answering.arXiv preprint arXiv:2003.10286, 2020. 4, 13, 14
2003 arXiv
-
[15]
Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020. 1
2020
-
[16]
Gpt-4 as an effec- tive zero-shot evaluator for scientific figure captions.arXiv preprint arXiv:2310.15405, 2023
Ting-Yao Hsu, Chieh-Yang Huang, Ryan Rossi, Sungchul Kim, C Lee Giles, and Ting-Hao K Huang. Gpt-4 as an effec- tive zero-shot evaluator for scientific figure captions.arXiv preprint arXiv:2310.15405, 2023. 5
2023 arXiv
-
[17]
Med-r1: Reinforcement learning for general- izable medical reasoning in vision-language models.arXiv preprint arXiv:2503.13939, 2025
Yuxiang Lai, Jike Zhong, Ming Li, Shitian Zhao, and Xi- aofeng Yang. Med-r1: Reinforcement learning for general- izable medical reasoning in vision-language models.arXiv preprint arXiv:2503.13939, 2025. 2
2025
-
[18]
A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018
Jason J Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman. A dataset of clinically generated visual questions and answers about radiology images.Scientific data, 5(1):1–10, 2018. 4, 13, 14
2018
-
[19]
Llava-med: Training a large language- and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36:28541–28564,
Chunyuan Li, Cliff Wong, Sheng Zhang, Naoto Usuyama, Haotian Liu, Jianwei Yang, Tristan Naumann, Hoifung Poon, and Jianfeng Gao. Llava-med: Training a large language- and-vision assistant for biomedicine in one day.Advances in Neural Information Processing Systems, 36:28541–28564,
-
[20]
Self-supervised vision-language pretraining for me- dial visual question answering
Pengfei Li, Gang Liu, Lin Tan, Jinying Liao, and Shenjun Zhong. Self-supervised vision-language pretraining for me- dial visual question answering. In2023 IEEE 20th Interna- tional Symposium on Biomedical Imaging (ISBI), pages 1–5. IEEE, 2023. 6
2023
-
[21]
Diffusion-lm improves control- lable text generation.Advances in neural information pro- cessing systems, 35:4328–4343, 2022
Xiang Li, John Thickstun, Ishaan Gulrajani, Percy S Liang, and Tatsunori B Hashimoto. Diffusion-lm improves control- lable text generation.Advances in neural information pro- cessing systems, 35:4328–4343, 2022. 2
2022
-
[22]
Pmc-clip: Con- trastive language-image pre-training using biomedical docu- ments
Weixiong Lin, Ziheng Zhao, Xiaoman Zhang, Chaoyi Wu, Ya Zhang, Yanfeng Wang, and Weidi Xie. Pmc-clip: Con- trastive language-image pre-training using biomedical docu- ments. InInternational Conference on Medical Image Com- puting and Computer-Assisted Intervention, pages 525–5...
2023
-
[23]
Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering
Bo Liu, Li-Ming Zhan, Li Xu, Lin Ma, Yan Yang, and Xiao-Ming Wu. Slake: A semantically-labeled knowledge- enhanced dataset for medical visual question answering. In 2021 IEEE 18th international symposium on biomedical imaging (ISBI), pages 1650–1654. IEEE, 2021. 4, 13, 14
2021
-
[24]
Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023. 3, 5, 13
2023
-
[25]
Improved baselines with visual instruction tuning
Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning. InPro- ceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 26296–26306, 2024. 3
2024
-
[26]
G-eval: Nlg evaluation us- ing gpt-4 with better human alignment.arXiv preprint arXiv:2303.16634, 2023
Yang Liu, Dan Iter, Yichong Xu, Shuohang Wang, Ruochen Xu, and Chenguang Zhu. G-eval: Nlg evaluation us- ing gpt-4 with better human alignment.arXiv preprint arXiv:2303.16634, 2023. 5
2023 arXiv
-
[27]
Q2atransformer: Improving medical vqa via an answer querying decoder
Yunyi Liu, Zhanyu Wang, Dong Xu, and Luping Zhou. Q2atransformer: Improving medical vqa via an answer querying decoder. InInternational conference on infor- mation processing in medical imaging, pages 445–456. Springer, 2023. 6
2023
-
[28]
Discrete diffusion modeling by estimating the ratios of the data distri- bution.arXiv preprint arXiv:2310.16834, 2023
Aaron Lou, Chenlin Meng, and Stefano Ermon. Discrete diffusion modeling by estimating the ratios of the data distri- bution.arXiv preprint arXiv:2310.16834, 2023. 1, 3
2023 arXiv
-
[29]
Biomedgpt: Open mul- timodal generative pre-trained transformer for biomedicine
Yizhen Luo, Jiahuan Zhang, Siqi Fan, Kai Yang, Yushuai Wu, Mu Qiao, and Zaiqing Nie. Biomedgpt: Open mul- timodal generative pre-trained transformer for biomedicine. arXiv preprint arXiv:2308.09442, 2023. 2
2023 arXiv
-
[30]
Tess: Text-to-text self-conditioned simplex diffu- sion.arXiv preprint arXiv:2305.08379, 2023
Rabeeh Karimi Mahabadi, Hamish Ivison, Jaesung Tae, James Henderson, Iz Beltagy, Matthew E Peters, and Arman Cohan. Tess: Text-to-text self-conditioned simplex diffu- sion.arXiv preprint arXiv:2305.08379, 2023. 2
2023 arXiv
-
[31]
Llama-3.2-11B-Vision: A Multimodal Vision– Language LLM
Meta AI. Llama-3.2-11B-Vision: A Multimodal Vision– Language LLM. Model card via Meta AI, 2024. Version released September 25, 2024; instruction-tuned for image reasoning, captioning, and VQA with 10.6B parameters. 5, 13
2024
-
[32]
Med-flamingo: a multimodal medical few-shot learner
Michael Moor, Qian Huang, Shirley Wu, Michihiro Ya- sunaga, Yash Dalmia, Jure Leskovec, Cyril Zakka, Ed- uardo Pontes Reis, and Pranav Rajpurkar. Med-flamingo: a multimodal medical few-shot learner. InMachine Learn- ing for Health (ML4H), pages 353–367. PMLR, 2023. 2, 5, 13
2023
-
[33]
Large language diffusion models.arXiv preprint arXiv:2502.09992, 2025
Shen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang, Jingyang Ou, Jun Hu, Jun Zhou, Yankai Lin, Ji-Rong Wen, and Chongxuan Li. Large language diffusion models.arXiv preprint arXiv:2502.09992, 2025. 1, 2, 3, 4, 8, 12, 13
2025 arXiv
-
[34]
Introducing gpt-4.1 in the api.https : / / openai
OpenAI. Introducing gpt-4.1 in the api.https : / / openai . com / index / gpt - 4 - 1/, 2025. Includes GPT-4.1, GPT-4.1 mini, and GPT-4.1 nano. 5, 12
2025
-
[35]
, and Barret Zoph
OpenAI, Josh Achiam, Steven Adler, Sandhini Agarwal, Lama Ahmad, Ilge Akkaya, . . . , and Barret Zoph. Gpt-4 technical report.https://arxiv.org/abs/2303. 08774, 2023. arXiv:2303.08774, submitted March 15 2023. 4
2023 arXiv
-
[36]
Your absorbing dis- crete diffusion secretly models the conditional distributions of clean data.arXiv preprint arXiv:2406.03736, 2024
Jingyang Ou, Shen Nie, Kaiwen Xue, Fengqi Zhu, Jiacheng Sun, Zhenguo Li, and Chongxuan Li. Your absorbing dis- crete diffusion secretly models the conditional distributions of clean data.arXiv preprint arXiv:2406.03736, 2024. 2, 3
2024 arXiv
-
[37]
Training language models to follow instructions with human feedback.Ad- vances in neural information processing systems, 35:27730– 27744, 2022
Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Car- roll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, et al. Training language models to follow instructions with human feedback.Ad- vances in neural information processing systems, 35:2...
2022
-
[38]
Medvlm-r1: Incentivizing medical reasoning ca- pability of vision-language models (vlms) via reinforcement learning.arXiv preprint arXiv:2502.19634, 2025
Jiazhen Pan, Che Liu, Junde Wu, Fenglin Liu, Jiayuan Zhu, Hongwei Bran Li, Chen Chen, Cheng Ouyang, and Daniel Rueckert. Medvlm-r1: Incentivizing medical reasoning ca- pability of vision-language models (vlms) via reinforcement learning.arXiv preprint arXiv:2502.19634, 2025. 2, 5, 13
2025 arXiv
-
[39]
Learning transferable visual models from natural language supervi- sion
Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervi- sion. InInternational conference on machine learning, p...
2021
-
[40]
Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning
Samyam Rajbhandari, Olatunji Ruwase, Jeff Rasley, Shaden Smith, and Yuxiong He. Zero-infinity: Breaking the gpu memory wall for extreme scale deep learning. InProceed- ings of the international conference for high performance computing, networking, storage and analysis, pages 1–14,
-
[41]
Categorical sdes with simplex diffusion.arXiv preprint arXiv:2210.14784, 2022
Pierre H Richemond, Sander Dieleman, and Arnaud Doucet. Categorical sdes with simplex diffusion.arXiv preprint arXiv:2210.14784, 2022. 2
2022 arXiv
-
[42]
Simple and effective masked dif- fusion language models.Advances in Neural Information Processing Systems, 37:130136–130184, 2024
Subham Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan, Edgar Marroquin, Justin Chiu, Alexander Rush, and V olodymyr Kuleshov. Simple and effective masked dif- fusion language models.Advances in Neural Information Processing Systems, 37:130136–130184, 2024. 1
2024
-
[43]
Simplified and generalized masked diffu- sion for discrete data.Advances in neural information pro- cessing systems, 37:103131–103167, 2024
Jiaxin Shi, Kehang Han, Zhe Wang, Arnaud Doucet, and Michalis Titsias. Simplified and generalized masked diffu- sion for discrete data.Advances in neural information pro- cessing systems, 37:103131–103167, 2024. 1
2024
-
[44]
Denoising diffusion implicit models.arXiv preprint arXiv:2010.02502, 2020
Jiaming Song, Chenlin Meng, and Stefano Ermon. Denoising diffusion implicit models.arXiv preprint arXiv:2010.02502, 2020. 1
2010 arXiv
-
[45]
Score-based generative modeling through stochastic differential equa- tions.arXiv preprint arXiv:2011.13456, 2020
Yang Song, Jascha Sohl-Dickstein, Diederik P Kingma, Ab- hishek Kumar, Stefano Ermon, and Ben Poole. Score-based generative modeling through stochastic differential equa- tions.arXiv preprint arXiv:2011.13456, 2020. 1
2011 arXiv
-
[46]
Self- conditioned embedding diffusion for text generation.arXiv preprint arXiv:2211.04236, 2022
Robin Strudel, Corentin Tallec, Florent Altch ´e, Yilun Du, Yaroslav Ganin, Arthur Mensch, Will Grathwohl, Niko- lay Savinov, Sander Dieleman, Laurent Sifre, et al. Self- conditioned embedding diffusion for text generation.arXiv preprint arXiv:2211.04236, 2022. 2
2022 arXiv
-
[47]
Llama: Open and efficient foundation language models
Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth´ee Lacroix, Baptiste Rozi`ere, Naman Goyal, Eric Hambro, Faisal Azhar, et al. Llama: Open and efficient foundation language models. arXiv preprint arXiv:2302.13971, 2023. 3
2023 arXiv
-
[48]
Siglip 2: Multilingual vision-language en- coders with improved semantic understanding, localization, and dense features.arXiv preprint arXiv:2502.14786, 2025
Michael Tschannen, Alexey Gritsenko, Xiao Wang, Muham- mad Ferjad Naeem, Ibrahim Alabdulmohsin, Nikhil Parthasarathy, Talfan Evans, Lucas Beyer, Ye Xia, Basil Mustafa, et al. Siglip 2: Multilingual vision-language en- coders with improved semantic understanding, localization, ...
2025 arXiv
-
[49]
Open- ended medical visual question answering through prefix tun- ing of language models
Tom Van Sonsbeek, Mohammad Mahdi Derakhshani, Ivona Najdenkoska, Cees GM Snoek, and Marcel Worring. Open- ended medical visual question answering through prefix tun- ing of language models. InInternational Conference on Medical Image Computing and Computer-Assisted Interven- t...
2023
-
[50]
A connection between score matching and denoising autoencoders.Neural computation, 23(7):1661– 1674, 2011
Pascal Vincent. A connection between score matching and denoising autoencoders.Neural computation, 23(7):1661– 1674, 2011. 3
2011
-
[51]
How does diverse interpretability of textual prompts impact med- ical vision-language zero-shot tasks?arXiv preprint arXiv:2409.00543, 2024
Sicheng Wang, Che Liu, and Rossella Arcucci. How does diverse interpretability of textual prompts impact med- ical vision-language zero-shot tasks?arXiv preprint arXiv:2409.00543, 2024. 2
2024 arXiv
-
[52]
Interactive computer-aided diag- nosis on medical image using large language models.Com- munications Engineering, 3(1):133, 2024
Sheng Wang, Zihao Zhao, Xi Ouyang, Tianming Liu, Qian Wang, and Dinggang Shen. Interactive computer-aided diag- nosis on medical image using large language models.Com- munications Engineering, 3(1):133, 2024. 2
2024
-
[53]
Medclip: Contrastive learning from unpaired medi- cal images and text
Zifeng Wang, Zhenbang Wu, Dinesh Agarwal, and Jimeng Sun. Medclip: Contrastive learning from unpaired medi- cal images and text. InProceedings of the Conference on Empirical Methods in Natural Language Processing. Con- ference on Empirical Methods in Natural Language Process- ...
2022
-
[54]
Gpt-4 as evaluator: Evaluating large language models on pest management in agriculture.arXiv preprint arXiv:2403.11858, 2024
Shanglong Yang, Zhipeng Yuan, Shunbao Li, Ruoling Peng, Kang Liu, and Po Yang. Gpt-4 as evaluator: Evaluating large language models on pest management in agriculture.arXiv preprint arXiv:2403.11858, 2024. 5
2024 arXiv
-
[55]
Dinoiser: Diffused conditional se- quence learning by manipulating noises.arXiv preprint arXiv:2302.10025, 2023
Jiasheng Ye, Zaixiang Zheng, Yu Bao, Lihua Qian, and Mingxuan Wang. Dinoiser: Diffused conditional se- quence learning by manipulating noises.arXiv preprint arXiv:2302.10025, 2023. 2
2023 arXiv
-
[56]
Llada-v: Large language diffusion models with visual instruction tuning
Zebin You, Shen Nie, Xiaolu Zhang, Jun Hu, Jun Zhou, Zhiwu Lu, Ji-Rong Wen, and Chongxuan Li. Llada-v: Large language diffusion models with visual instruction tuning. arXiv preprint arXiv:2505.16933, 2025. 2, 3, 4, 5, 12, 13
2025 arXiv
-
[57]
Biomedgpt: A unified and generalist biomed- ical generative pre-trained transformer for vision, language, and multimodal tasks.arXiv e-prints, pages arXiv–2305,
Kai Zhang, Jun Yu, Eashan Adhikarla, Rong Zhou, Zhiling Yan, Yixin Liu, Zhengliang Liu, Lifang He, Brian Davison, Xiang Li, et al. Biomedgpt: A unified and generalist biomed- ical generative pre-trained transformer for vision, language, and multimodal tasks.arXiv e-prints, pag...
-
[58]
Large-scale domain-specific pre- training for biomedical vision-language processing.arXiv preprint arXiv:2303.00915, 2(3):6, 2023
Sheng Zhang, Yanbo Xu, Naoto Usuyama, Jaspreet Bagga, Robert Tinn, Sam Preston, Rajesh Rao, Mu Wei, Naveen Valluri, Cliff Wong, et al. Large-scale domain-specific pre- training for biomedical vision-language processing.arXiv preprint arXiv:2303.00915, 2(3):6, 2023. 6
2023 arXiv
-
[59]
Toward effective reinforcement learning fine-tuning for med- ical vqa in vision-language models.arXiv e-prints, pages arXiv–2505, 2025
Wenhui Zhu, Xuanzhao Dong, Xin Li, Peijie Qiu, Xiwen Chen, Abolfazl Razi, Aris Sotiras, Yi Su, and Yalin Wang. Toward effective reinforcement learning fine-tuning for med- ical vqa in vision-language models.arXiv e-prints, pages arXiv–2505, 2025. 5, 13
2025
-
[60]
Retinalgpt: A retinal clinical preference conversational assistant powered by large vision- language models.arXiv e-prints, pages arXiv–2503, 2025
Wenhui Zhu, Xin Li, Xiwen Chen, Peijie Qiu, Vamsi Kr- ishna Vasa, Xuanzhao Dong, Yanxi Chen, Natasha Lepore, Oana Dumitrascu, Yi Su, et al. Retinalgpt: A retinal clinical preference conversational assistant powered by large vision- language models.arXiv e-prints, pages arXiv–2...
2025
-
[256]
Unless otherwise specified, this setting is used consistently in all detailed analyses
However, for downstream biomedical VQA tasks, we disable the semi-autoregressive mechanism by settingL= B=Z= 64. Unless otherwise specified, this setting is used consistently in all detailed analyses. Baseline Model Configuration.We compare LLaDA- MedV against nine baseline vi...
1943
-
[2023]
2, 3, 4, 5, 6, 12, 13, 14, 15
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.