Pith. sign in

REVIEW 2 major objections 4 minor 67 references

OralAgent is a dental AI agent that reasons, calls vision tools, and retrieves textbook knowledge to analyze images end-to-end.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · grok-4.5

2026-07-13 00:10 UTC pith:D3UDBJ3Y

load-bearing objection Solid dental systems paper: first full ReAct agent with 22 tools + textbook-scale bilingual RAG, real SOTA gains and public artifacts; tool accuracy is self-reported on internal splits, so the margins need a grain of salt. the 2 major comments →

arxiv 2605.27378 v1 pith:D3UDBJ3Y submitted 2026-04-09 cs.CL cs.CVcs.MA

OralAgent: Integrating Reasoning, Tools, and Knowledge for Interactive Dental Image Analysis

classification cs.CL cs.CVcs.MA
keywords medical agentmultimodal large language modeldental image analysisretrieval-augmented generationReActtool useOralCorpusOralQA-ZH
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

Isolated dental AI models for single tasks and single imaging types do not match real clinical workflows, which require flexible multi-step analysis across modalities. This paper introduces OralAgent, a ReAct-style agent that understands user intent and image modality, then loops through observation, thought, and action: it can call any of 22 specialized vision tools, retrieve passages from 368 dental textbooks, and synthesize an answer with visualizations and citations. The authors also release OralCorpus (134.8M bilingual tokens for dental retrieval) and OralQA-ZH (798 Chinese multiple-choice questions across eleven oral subspecialties). On public dental VQA benchmarks the agent sets new highs, and when different language models serve as its orchestrator, the textbook retrieval module consistently lifts accuracy on multidisciplinary dental knowledge questions. The result is an interpretable, modular system meant to be usable in real clinical settings rather than as another single-purpose detector.

Core claim

A single dental-specialized agent that combines multimodal reasoning, a toolbox of 22 vision experts, and retrieval over 368 classical textbooks can outperform both standalone dental multimodal models and general medical agents on open-ended dental image analysis and on multidisciplinary dental knowledge questions, while producing tool traces, visualizations, and page-level citations.

What carries the argument

The ReAct observation–thought–action loop driven by a configurable orchestrator: after intent and modality classification, the agent selects and runs dental vision tools in parallel, retrieves top-k textbook chunks with sources, and iterates until it can answer, with explicit instructions to critically evaluate tool outputs against its own knowledge.

Load-bearing premise

The specialist vision tools are accurate enough that the orchestrator can safely use or override their outputs, and that the public open-ended VQA benchmarks are a fair stand-in for real multi-step clinical work.

What would settle it

Run the same agent without its dental vision tools or without OralCorpus retrieval on a held-out set of real multi-turn clinical cases; if overall accuracy and clinical usefulness do not drop relative to the reported SOTA numbers, the claim that tools-plus-RAG drive the gains fails.

Watch this falsifier — get emailed when new claim-graph text bears on it.

If this is right

  • Dental image analysis can be delivered as one configurable agent instead of a pile of single-task models for each modality.
  • Adding or swapping a vision tool or a private textbook collection becomes a plug-in change rather than a full retrain.
  • Answers can carry tool visualizations and book-and-page citations, improving auditability for clinical use.
  • Other oral subspecialties and languages can be covered by extending the same toolkit and OralCorpus pipeline.
  • Future multimodal RAG and 3D imaging tools can slot into the same ReAct loop without redesigning the agent.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The same modular pattern—intent/modality front-end, specialist tool cohort, domain corpus RAG—could transfer to other imaging-heavy specialties that currently suffer from fragmented task models.
  • Because many tools were trained and validated on public splits, real-world drift in camera or scanner settings remains an untested risk for the observation quality the loop depends on.
  • OralQA-ZH’s multi-correct multiple-choice format may understate the value of free-text clinical reasoning that the agent produces in multi-turn dialogue.
  • If clinics adopt private knowledge bases as the paper allows, the system’s reliability will hinge as much on local document quality as on the public OralCorpus.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 4 minor

Summary. The paper introduces OralAgent, a ReAct-style dental AI agent that couples a core LLM/MLLM orchestrator with (i) an instruction-comprehension stage (intent recognition + modality classification), (ii) a toolbox of 22 specialist vision models spanning six imaging modalities, and (iii) a RAG branch over OralCorpus (134.8M tokens from 368 textbooks). It also releases OralQA-ZH, a 798-item Chinese MCQ benchmark across 11 oral subspecialties. On MMOral-Uni and MMOral-OPG the system reports new SOTA overall scores of 57.70 and 61.00 (gains of +5.86 and +15.69 over OralGPT-Omni); on OralQA-ZH, attaching the OralCorpus RAG module consistently lifts a range of LLM orchestrators. Ablations isolate the intent and modality modules; tool-use statistics and multi-turn case studies are provided. Code and models are stated to be public.

Significance. If the reported gains hold under independent scrutiny, the work is a useful systems contribution: it is the first end-to-end dental agent that unifies multimodal tools, textbook RAG, and multi-step planning, and it ships two reusable resources (OralCorpus, OralQA-ZH) plus a public codebase. The modular design (plug-and-play tools, private knowledge bases, configurable orchestrator) is practically relevant for clinical deployment. The SOTA margins on public multimodal benchmarks and the consistent RAG lifts on OralQA-ZH are concrete, falsifiable claims that advance dental MLLM evaluation beyond single-pass models.

major comments (2)
  1. Section III-C and Table I: many of the 22 tools (especially the 12 in-house models) report very high metrics (e.g., Acc=99.9%, mAP50=99.0) obtained only from 8:2 random validation splits of each tool’s own training set. There is no external held-out test set, no cross-dataset evaluation, and no measurement of how tool false positives/negatives propagate through the ReAct loop into final VQA scores. Because the headline SOTA margins rest on the premise that these tools supply reliable observations the orchestrator can trust or override, the current evidence is insufficient to secure the central claim.
  2. Tables III–IV and the ablation in Table VI: the ablation removes only intent recognition and modality classification; it does not ablate or replace the specialist tools. Consequently the large gains over pure MLLM baselines (including OralGPT-Omni) cannot be cleanly attributed to orchestration versus simply granting access to stronger detectors that the baselines lack. A controlled experiment that freezes the tool set and varies only the agent loop (or that substitutes weaker tools) is needed to isolate the contribution of the agent framework itself.
minor comments (4)
  1. Figure 5 reports average tool calls (3.67) but does not break down success/failure rates or correlation with final answer correctness; adding this would strengthen the tool-use analysis.
  2. OralQA-ZH construction (Section IV) states that a senior professional reviewed 20% of items; the exact agreement rate and any residual error rate should be reported.
  3. Several baseline citations and model names appear with inconsistent versioning (e.g., GPT-5 vs GPT-5.4, Qwen2.5-VL vs Qwen3-VL); a single consistent naming table would help reproducibility.
  4. The prompt used to instruct the orchestrator to critically evaluate tool outputs is mentioned but not fully reproduced; including it in the appendix would aid replication.

Circularity Check

0 steps flagged

No definitional circularity in the agent claims; SOTA numbers are empirical measurements on (partly self-authored) benchmarks, not forced by construction from fitted inputs or self-definitions.

full rationale

OralAgent is an engineering systems paper whose central claims are empirical SOTA scores (57.70 on MMOral-Uni, 61.00 on MMOral-OPG) obtained by running a ReAct loop that calls 22 specialist tools plus RAG over OralCorpus. These scores are measured against external or newly released VQA items and are not algebraically identical to any free parameter or definition inside the paper. Tool accuracies in Table I are self-reported on 8:2 validation splits of each tool’s training data and several tools/baselines (OralGPT-Omni, OralGPT, MMOral-*) share authors, but the agent’s reported gains are obtained by additional orchestration and multi-tool composition rather than by re-labeling those tool metrics as the final answer. OralQA-ZH questions are extracted from the same textbook collection used to build OralCorpus, so RAG improvements on that benchmark are expected once retrieval succeeds; this is ordinary RAG evaluation design, not a fitted-parameter-as-prediction loop. No uniqueness theorem, ansatz, or first-principles derivation is claimed. The derivation chain therefore remains self-contained against the stated benchmarks; residual concerns about tool generalization or benchmark authorship belong to correctness risk, not circularity.

Axiom & Free-Parameter Ledger

2 free parameters · 3 axioms · 3 invented entities

The central empirical claims rest on standard agent and RAG assumptions plus the reliability of the 22 specialist tools and the quality of the newly constructed corpus and benchmark. No continuous free parameters are fitted to the main evaluation tables; the main modeling choices are discrete (choice of orchestrator, top-k=7, intent taxonomy).

free parameters (2)
  • RAG top-k = 7
    Default top-k set to 7 for retrieval; affects which textbook chunks reach the orchestrator.
  • Intent taxonomy size = 9 categories
    Nine hand-defined intent categories chosen by brainstorming with dentists; used by the 0.6 B intent classifier.
axioms (3)
  • domain assumption ReAct-style observation–thought–action loops with structured tool calling are a valid control paradigm for multi-step dental image analysis.
    Adopted from Yao et al. 2022 and used throughout Algorithm 1 and Section III-A without further justification.
  • domain assumption The 22 specialist vision models produce observations accurate enough that the orchestrator can critically evaluate and override them when needed.
    Stated via prompt engineering in Section III-A and supported only by per-tool validation-set metrics in Table I.
  • domain assumption OralQA-ZH questions extracted from licensing exams and textbooks, after professional spot-check, constitute a valid measure of multidisciplinary dental knowledge.
    Construction described in Section IV; 20 % random review by a senior academic is the only quality control reported.
invented entities (3)
  • OralAgent independent evidence
    purpose: End-to-end dental agent that unifies intent/modality understanding, 22 vision tools, and textbook RAG under a ReAct loop.
    The system itself is the primary contribution; independent evidence is the public code and the three benchmark scores.
  • OralCorpus independent evidence
    purpose: 134.8 M-token bilingual vector store built from 368 dental textbooks for RAG.
    New resource; independent evidence is the reported token counts and the RAG ablation gains on OralQA-ZH.
  • OralQA-ZH independent evidence
    purpose: 798-item Chinese multiple-choice benchmark spanning 11 oral subspecialties.
    New evaluation resource; independent evidence is the category distribution table and the public claim of professional verification.

pith-pipeline@v1.1.0-grok45 · 29670 in / 2805 out tokens · 32695 ms · 2026-07-13T00:10:22.755341+00:00 · methodology

0 comments
read the original abstract

Dental image analysis plays a pivotal role in supporting accurate diagnosis and treatment planning in oral healthcare. Although recent advances have produced dental AI models for specific tasks and individual imaging modalities, their isolated designs limit practical use in real-world clinical workflows. In this paper, we present OralAgent, the first dental-specialized AI agent that unifies multimodal reasoning, tool-based decision-making, and knowledge-grounded retrieval within an end-to-end automated framework. It integrates 22 visual analysis tools and 368 widely-used classical dental textbooks, enabling autonomous reasoning, planning, tool use, knowledge retrieval, and multi-step workflow execution. Furthermore, we introduce OralCorpus, a large-scale, high-quality bilingual textual resource containing 134.8M tokens curated for dental retrieval-augmented generation (RAG). To evaluate models' multidisciplinary dental knowledge, we construct OralQA-ZH, a Chinese multiple-choice question benchmark consisting of 798 items across eleven oral subspecialties. Extensive experiments demonstrate that OralAgent achieves state-of-the-art performance on the MMOral-Uni, MMOral-OPG, and OralQA-ZH benchmarks, highlighting its effectiveness, interpretability, and adaptability in real-world clinical settings. The code and models are publicly available at https://github.com/isjinghao/OralAgent.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

67 extracted references · 29 linked inside Pith

  1. [1]

    A semi-supervised transformer-based deep learning framework for automated tooth segmentation and identification on panoramic radiographs,

    J. Hao, L. M. Wong, Z. Shan, Q. Y . H. Ai, X. Shi, J. K. H. Tsoi, and K. F. Hung, “A semi-supervised transformer-based deep learning framework for automated tooth segmentation and identification on panoramic radiographs,”Diagnostics, vol. 14, no. 17, p. 1948, 2024

  2. [2]

    Semit-sam: Building a visual foundation model for tooth instance segmentation on panoramic radiographs,

    J. Hao, M. Liu, L. He, L. Yao, J. K. H. Tsoi, and K. F. Hung, “Semit-sam: Building a visual foundation model for tooth instance segmentation on panoramic radiographs,” inInternational Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2024, pp. 110–121

  3. [3]

    T-mamba: a unified framework with long-range dependency in dual-domain for 2d & 3d tooth segmentation,

    J. Hao, Y . Zhu, L. He, M. Liu, J. K. H. Tsoi, and K. F. Hung, “T-mamba: a unified framework with long-range dependency in dual-domain for 2d & 3d tooth segmentation,”arXiv preprint arXiv:2404.01065, 2024

  4. [4]

    Char- acteristics, licensing, and ethical considerations of openly accessible oral-maxillofacial imaging datasets: a systematic review,

    J. Hao, A. Nalley, A. W. K. Yeung, R. Tanaka, Q. Y . H. Ai, W. Y . H. Lam, Z. Shan, Y . Y . Leung, A. AlHadidi, M. M. Bornsteinet al., “Char- acteristics, licensing, and ethical considerations of openly accessible oral-maxillofacial imaging datasets: a systematic review,”npj Digital Medicine, vol. 8, no. 1, p. 412, 2025

  5. [5]

    Photography-based dental plaque detection and report generation among preschool children using plaquesam,

    J. Hao, K. Guo, Z. Shan, Y . Yang, J. K.-H. Tsoi, P. P. Y . Lam, and K. F. Hung, “Photography-based dental plaque detection and report generation among preschool children using plaquesam,” 2026

  6. [6]

    Cephalometric landmark detection across ages with prototypical net- work,

    H. Wu, C. Wang, L. Mei, T. Yang, M. Zhu, D. Shen, and Z. Cui, “Cephalometric landmark detection across ages with prototypical net- work,” inInternational conference on medical image computing and computer-assisted intervention. Springer, 2024, pp. 155–165

  7. [7]

    A high magnifications histopathology image dataset for oral squamous cell carcinoma diagnosis and prognosis,

    J. Guan, J. Guo, Q. Chen, J. Chen, Y . Cai, Y . He, Z. Huang, Y . Wang, and Y . Xie, “A high magnifications histopathology image dataset for oral squamous cell carcinoma diagnosis and prognosis,”Scientific Data, 2026. AUTHORet al.: PREPARATION OF PAPERS FOR IEEE TRANSACTIONS ON MEDICAL IMAGING 13

  8. [8]

    Descriptive caption enhancement with visual specialists for multimodal perception,

    Y . Sun, J. Hao, K. Zhu, J.-J. Liu, Y . Zhao, X. Li, G. Zhang, Z. Li, and J. Wang, “Descriptive caption enhancement with visual specialists for multimodal perception,”arXiv preprint arXiv:2412.14233, 2024

  9. [9]

    Fullanno: A data engine for enhancing image comprehension of mllms,

    J. Hao, Y . Zhao, S. Chen, Y . Sun, Q. Chen, G. Zhang, K. Yao, E. Ding, and J. Wang, “Fullanno: A data engine for enhancing image comprehension of mllms,”arXiv preprint arXiv:2409.13540, 2024

  10. [10]

    Oralgpt-omni: A versatile dental multimodal large language model,

    J. Hao, Y . Liang, L. Lin, Y . Fan, W. Zhou, K. Guo, Z. Ye, Y . Sun, X. Zhang, Y . Yanget al., “Oralgpt-omni: A versatile dental multimodal large language model,”CVPR, 2026

  11. [11]

    Dentalgpt: Incentivizing multimodal complex reasoning in dentistry,

    Z. Cai, J. Zhang, J. Zhao, Z. Zeng, Y . Li, J. Liang, J. Chen, Y . Yang, J. You, S. Denget al., “Dentalgpt: Incentivizing multimodal complex reasoning in dentistry,”arXiv preprint arXiv:2512.11558, 2025

  12. [12]

    The performance of large language models in dentomax- illofacial radiology: a systematic review,

    Z. Liu, A. Nalley, J. Hao, Q. Y . H Ai, A. W. Kan Yeung, R. Tanaka, and K. F. Hung, “The performance of large language models in dentomax- illofacial radiology: a systematic review,”Dentomaxillofacial Radiology, p. twaf060, 2025

  13. [13]

    React: Synergizing reasoning and acting in language models,

    S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y . Cao, “React: Synergizing reasoning and acting in language models,” inThe eleventh international conference on learning representations, 2022

  14. [14]

    Towards better dental ai: A multimodal benchmark and instruction dataset for panoramic x-ray analysis,

    J. Hao, Y . Fan, Y . Sun, K. Guo, L. Lin, J. Yang, Q. Y . H. Ai, L. M. Wong, H. Tang, and K. F. Hung, “Towards better dental ai: A multimodal benchmark and instruction dataset for panoramic x-ray analysis,”NeurIPS 2025, 2025

  15. [15]

    A survey on rag with llms,

    M. Arslan, H. Ghanem, S. Munawar, and C. Cruz, “A survey on rag with llms,”Procedia computer science, vol. 246, pp. 3781–3790, 2024

  16. [16]

    Mdagents: An adaptive collaboration of llms for medical decision-making,

    Y . Kim, C. Park, H. Jeong, Y . S. Chan, X. Xu, D. McDuff, H. Lee, M. Ghassemi, C. Breazeal, and H. W. Park, “Mdagents: An adaptive collaboration of llms for medical decision-making,”Advances in Neural Information Processing Systems, vol. 37, pp. 79 410–79 452, 2024

  17. [17]

    Mmedagent: Learning to use medical tools with multi- modal agent,

    B. Li, T. Yan, Y . Pan, J. Luo, R. Ji, J. Ding, Z. Xu, S. Liu, H. Dong, Z. Linet al., “Mmedagent: Learning to use medical tools with multi- modal agent,” inFindings of the Association for Computational Linguis- tics: EMNLP 2024, 2024, pp. 8745–8760

  18. [18]

    Medagent-pro: Towards evidence-based multi-modal medical diagnosis via reasoning agentic workflow,

    Z. Wang, J. Wu, L. Cai, C. H. Low, X. Yang, Q. Li, and Y . Jin, “Medagent-pro: Towards evidence-based multi-modal medical diagnosis via reasoning agentic workflow,”arXiv preprint arXiv:2503.18968, 2025

  19. [19]

    Medrag: Enhancing retrieval-augmented generation with knowledge graph-elicited reasoning for healthcare copilot,

    X. Zhao, S. Liu, S.-Y . Yang, and C. Miao, “Medrag: Enhancing retrieval-augmented generation with knowledge graph-elicited reasoning for healthcare copilot,” inProceedings of the ACM on Web Conference 2025, 2025, pp. 4442–4457

  20. [20]

    Medrax: Med- ical reasoning agent for chest x-ray,

    A. Fallahpour, J. Ma, A. Munim, H. Lyu, and B. Wang, “Medrax: Med- ical reasoning agent for chest x-ray,”arXiv preprint arXiv:2502.02673, 2025

  21. [21]

    Oralgpt-plus: Learning to use visual tools via reinforcement learning for panoramic x-ray analysis,

    Y . Fan, J. Hao, H. Chen, J. Bao, Y . Shao, Y . Liang, K. F. Hung, and H. Tang, “Oralgpt-plus: Learning to use visual tools via reinforcement learning for panoramic x-ray analysis,”CVPR 2026, 2026

  22. [22]

    Dentvlm: A multimodal vision-language model for comprehensive dental diagnosis and enhanced clinical practice,

    Z. Meng, J. Hao, X. Dai, Y . Feng, J. Liu, B. Feng, H. Wu, X. Gai, H. Zhu, T. Huet al., “Dentvlm: A multimodal vision-language model for comprehensive dental diagnosis and enhanced clinical practice,”arXiv preprint arXiv:2509.23344, 2025

  23. [23]

    Opgagent: An agent for auditable dental panoramic x-ray interpretation,

    Z. Yu, L. Yang, B. Babicka, M. Hu, J. Hao, A. Huang, J. Huang, Y . Jin, J. Wu, and Z. Ge, “Opgagent: An agent for auditable dental panoramic x-ray interpretation,”arXiv preprint arXiv:2603.00462, 2026

  24. [24]

    Prompt engineering for healthcare: Methodologies and applications,

    J. Wang, E. Shi, S. Yu, Z. Wu, H. Hu, C. Ma, H. Dai, Q. Yang, Y . Kang, J. Wuet al., “Prompt engineering for healthcare: Methodologies and applications,”Meta-Radiology, p. 100190, 2025

  25. [25]

    Intentgpt: Few-shot intent discovery with large language models,

    J. A. Rodriguez, N. Botzer, D. Vazquez, C. Pal, M. Pedersoli, and I. Laradji, “Intentgpt: Few-shot intent discovery with large language models,”arXiv preprint arXiv:2411.10670, 2024

  26. [26]

    Qwen3 technical report,

    A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lvet al., “Qwen3 technical report,”arXiv preprint arXiv:2505.09388, 2025

  27. [27]

    Biomedclip: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs,

    S. Zhang, Y . Xu, N. Usuyama, H. Xu, J. Bagga, R. Tinn, S. Preston, R. Rao, M. Wei, N. Valluriet al., “Biomedclip: a multimodal biomedical foundation model pretrained from fifteen million scientific image-text pairs,”arXiv preprint arXiv:2303.00915, 2023

  28. [28]

    Sim ´eoni, H

    O. Sim ´eoni, H. V . V o, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V . Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoaet al., “Dinov3,” arXiv preprint arXiv:2508.10104, 2025

  29. [29]

    Dino: Detr with improved denoising anchor boxes for end-to- end object detection,

    H. Zhang, F. Li, S. Liu, L. Zhang, H. Su, J. Zhu, L. M. Ni, and H.-Y . Shum, “Dino: Detr with improved denoising anchor boxes for end-to- end object detection,”arXiv preprint arXiv:2203.03605, 2022

  30. [30]

    Mask dino: Towards a unified transformer-based framework for object detection and segmentation,

    F. Li, H. Zhang, H. Xu, S. Liu, L. Zhang, L. M. Ni, and H.-Y . Shum, “Mask dino: Towards a unified transformer-based framework for object detection and segmentation,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 3041– 3050

  31. [31]

    Mineru: An open-source solution for precise document content extraction,

    B. Wang, C. Xu, X. Zhao, L. Ouyang, F. Wu, Z. Zhao, R. Xu, K. Liu, Y . Qu, F. Shanget al., “Mineru: An open-source solution for precise document content extraction,”arXiv preprint arXiv:2409.18839, 2024

  32. [32]

    Qwen3 embedding: Advancing text embedding and reranking through foundation models,

    Y . Zhang, M. Li, D. Long, X. Zhang, H. Lin, B. Yang, P. Xie, A. Yang, D. Liu, J. Linet al., “Qwen3 embedding: Advancing text embedding and reranking through foundation models,”arXiv preprint arXiv:2506.05176, 2025

  33. [33]

    [Online]

    OpenAI, “Gpt-5,” 2025. [Online]. Available: https://openai.com/ zh-Hans-CN/index/introducing-gpt-5

  34. [34]

    [Online]

    ——, “o3,” 2025. [Online]. Available: https://openai.com/zh-Hans-CN/ index/introducing-o3-and-o4-mini

  35. [35]

    [Online]

    xAI, “grok4,” 2025. [Online]. Available: https://x.ai/grok

  36. [36]

    Doubao-1.5-vision-pro,

    ByteDance, “Doubao-1.5-vision-pro,” 2025. [Online]. Available: https: //www.volcengine.com/product/doubao

  37. [37]

    Gemini: a family of highly capable multimodal models,

    G. Team, R. Anil, S. Borgeaud, J.-B. Alayrac, J. Yu, R. Soricut, J. Schalkwyk, A. M. Dai, A. Hauth, K. Millicanet al., “Gemini: a family of highly capable multimodal models,”arXiv preprint arXiv:2312.11805, 2023

  38. [38]

    Glm-4.5 v and glm-4.1 v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning, 2025,

    V . Team, W. Hong, W. Yu, X. Gu, G. Wang, G. Gan, H. Tang, J. Cheng, J. Qi, J. Jiet al., “Glm-4.5 v and glm-4.1 v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning, 2025,”https://arxiv.org/abs/2507.01006, 2025

  39. [39]

    Qwen2. 5-vl technical report,

    S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tanget al., “Qwen2. 5-vl technical report,”arXiv preprint arXiv:2502.13923, 2025

  40. [40]

    Internvl3. 5: Advancing open-source multi- modal models in versatility, reasoning, and efficiency,

    W. Wang, Z. Gao, L. Gu, H. Pu, L. Cui, X. Wei, Z. Liu, L. Jing, S. Ye, J. Shaoet al., “Internvl3. 5: Advancing open-source multi- modal models in versatility, reasoning, and efficiency,”arXiv preprint arXiv:2508.18265, 2025

  41. [41]

    Llava-next: Improved reasoning, ocr, and world knowledge,

    H. Liu, C. Li, Y . Li, B. Li, Y . Zhang, S. Shen, and Y . J. Lee, “Llava-next: Improved reasoning, ocr, and world knowledge,” January 2024. [Online]. Available: https://llava-vl.github.io/blog/2024-01-30-llava-next/

  42. [42]

    Llava-onevision: Easy visual task transfer,

    B. Li, Y . Zhang, D. Guo, R. Zhang, F. Li, H. Zhang, K. Zhang, P. Zhang, Y . Li, Z. Liuet al., “Llava-onevision: Easy visual task transfer,”arXiv preprint arXiv:2408.03326, 2024

  43. [43]

    Mimo: Unlocking the reasoning potential of language model–from pretraining to posttraining,

    L. Xiaomi, B. Xia, B. Shen, D. Zhu, D. Zhang, G. Wang, H. Zhang, H. Liu, J. Xiao, J. Donget al., “Mimo: Unlocking the reasoning potential of language model–from pretraining to posttraining,”arXiv preprint arXiv:2505.07608, 2025

  44. [44]

    Phi- 4 technical report,

    M. Abdin, J. Aneja, H. Behl, S. Bubeck, R. Eldan, S. Gunasekar, M. Harrison, R. J. Hewett, M. Javaheripi, P. Kauffmannet al., “Phi- 4 technical report,”arXiv preprint arXiv:2412.08905, 2024

  45. [45]

    Mistral-small-3.1-24b-instruct-2503,

    mistralai, “Mistral-small-3.1-24b-instruct-2503,” 2025. [On- line]. Available: https://huggingface.co/mistralai/Mistral-Small-3. 1-24B-Instruct-2503

  46. [46]

    R-4b: Incen- tivizing general-purpose auto-thinking capability in mllms via bi-mode annealing and reinforce learning,

    Q. Yang, B. Ni, S. Xiang, H. Hu, H. Peng, and J. Jiang, “R-4b: Incen- tivizing general-purpose auto-thinking capability in mllms via bi-mode annealing and reinforce learning,”arXiv preprint arXiv:2508.21113, 2025

  47. [47]

    Ovis2. 5 technical report,

    S. Lu, Y . Li, Y . Xia, Y . Hu, S. Zhao, Y . Ma, Z. Wei, Y . Li, L. Duan, J. Zhaoet al., “Ovis2. 5 technical report,”arXiv preprint arXiv:2508.11737, 2025

  48. [48]

    Llava-med: Training a large language-and-vision assistant for biomedicine in one day,

    C. Li, C. Wong, S. Zhang, N. Usuyama, H. Liu, J. Yang, T. Naumann, H. Poon, and J. Gao, “Llava-med: Training a large language-and-vision assistant for biomedicine in one day,”Advances in Neural Information Processing Systems, vol. 36, pp. 28 541–28 564, 2023

  49. [49]

    Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale,

    J. Chen, C. Gui, R. Ouyang, A. Gao, S. Chen, G. H. Chen, X. Wang, R. Zhang, Z. Cai, K. Jiet al., “Huatuogpt-vision, towards injecting medical visual knowledge into multimodal llms at scale,”arXiv preprint arXiv:2406.19280, 2024

  50. [50]

    Lingshu: A generalist foundation model for unified multimodal medical understanding and reasoning,

    W. Xu, H. P. Chan, L. Li, M. Aljunied, R. Yuan, J. Wang, C. Xiao, G. Chen, C. Liu, Z. Liet al., “Lingshu: A generalist foundation model for unified multimodal medical understanding and reasoning,”arXiv preprint arXiv:2506.07044, 2025

  51. [51]

    Medvlm-r1: Incentivizing medical reasoning capability of vision-language models (vlms) via reinforcement learning,

    J. Pan, C. Liu, J. Wu, F. Liu, J. Zhu, H. B. Li, C. Chen, C. Ouyang, and D. Rueckert, “Medvlm-r1: Incentivizing medical reasoning capability of vision-language models (vlms) via reinforcement learning,” inInterna- tional Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2025, pp. 337–347

  52. [52]

    Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models,

    Y . Lai, J. Zhong, M. Li, S. Zhao, and X. Yang, “Med-r1: Reinforcement learning for generalizable medical reasoning in vision-language models,” arXiv preprint arXiv:2503.13939, 2025

  53. [53]

    H. Sun, Y . Jiang, W. Lou, Y . Zhang, W. Li, L. Wang, M. Liu, L. Liu, and X. Wang, “Chiron-o1: Igniting multimodal large language models towards generalizable medical reasoning via mentor-intern collaborative 14 IEEE TRANSACTIONS ON MEDICAL IMAGING, VOL. XX, NO. XX, XXXX 2020 search,” inThe Thirty-ninth Annual Conference on Neural Information Processing Systems

  54. [54]

    Medgemma technical report,

    A. Sellergren, S. Kazemzadeh, T. Jaroensri, A. Kiraly, M. Traverse, T. Kohlberger, S. Xu, F. Jamil, C. Hughes, C. Lauet al., “Medgemma technical report,”arXiv preprint arXiv:2507.05201, 2025

  55. [55]

    Healthgpt: A medical large vision-language model for unifying comprehension and generation via heterogeneous knowledge adaptation,

    T. Lin, W. Zhang, S. Li, Y . Yuan, B. Yu, H. Li, W. He, H. Jiang, M. Li, X. Songet al., “Healthgpt: A medical large vision-language model for unifying comprehension and generation via heterogeneous knowledge adaptation,”arXiv preprint arXiv:2502.09838, 2025

  56. [56]

    Medagents: Large language models as collaborators for zero-shot medical reasoning,

    X. Tang, A. Zou, Z. Zhang, Z. Li, Y . Zhao, X. Zhang, A. Cohan, and M. Gerstein, “Medagents: Large language models as collaborators for zero-shot medical reasoning,” inFindings of the Association for Computational Linguistics: ACL 2024, 2024, pp. 599–621

  57. [57]

    Gpt-4o system card,

    A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radfordet al., “Gpt-4o system card,”arXiv preprint arXiv:2410.21276, 2024

  58. [58]

    Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond,

    J. Bai, S. Bai, S. Yang, S. Wang, S. Tan, P. Wang, J. Lin, C. Zhou, and J. Zhou, “Qwen-vl: A versatile vision-language model for un- derstanding, localization, text reading, and beyond,”arXiv preprint arXiv:2308.12966, 2023

  59. [59]

    Deepseek-vl: towards real-world vision-language understanding,

    H. Lu, W. Liu, B. Zhang, B. Wang, K. Dong, B. Liu, J. Sun, T. Ren, Z. Li, H. Yanget al., “Deepseek-vl: towards real-world vision-language understanding,”arXiv preprint arXiv:2403.05525, 2024

  60. [60]

    Chatglm: A family of large language models from glm-130b to glm-4 all tools,

    T. GLM, A. Zeng, B. Xu, B. Wang, C. Zhang, D. Yin, D. Rojas, G. Feng, H. Zhao, H. Lai, H. Yu, H. Wang, J. Sun, J. Zhang, J. Cheng, J. Gui, J. Tang, J. Zhang, J. Li, L. Zhao, L. Wu, L. Zhong, M. Liu, M. Huang, P. Zhang, Q. Zheng, R. Lu, S. Duan, S. Zhang, S. Cao, S. Yang, W. L. Tam, W. Zhao, X. Liu, X. Xia, X. Zhang, X. Gu, X. Lv, X. Liu, X. Liu, X. Yang, ...

  61. [61]

    Meddr: Diagnosis-guided bootstrapping for large-scale medical vision-language learning,

    S. He, Y . Nie, Z. Chen, Z. Cai, H. Wang, S. Yang, and H. Chen, “Meddr: Diagnosis-guided bootstrapping for large-scale medical vision-language learning,”arXiv preprint arXiv:2404.15127, vol. 1, no. 3, p. 6, 2024

  62. [62]

    Gpt-5.4,

    OpenAI, “Gpt-5.4,” 2026. [Online]. Available: https://openai.com/ zh-Hans-CN/index/introducing-gpt-5-4/

  63. [63]

    Kimi k2. 5: Visual agentic intelligence,

    K. Team, T. Bai, Y . Bai, Y . Bao, S. Cai, Y . Cao, Y . Charles, H. Che, C. Chen, G. Chenet al., “Kimi k2. 5: Visual agentic intelligence,”arXiv preprint arXiv:2602.02276, 2026

  64. [64]

    Minimax m2.7,

    MiniMax, “Minimax m2.7,” 2026. [Online]. Available: https://www. minimaxi.com/models/text/m27

  65. [65]

    Qwen3.5,

    Q. Team, “Qwen3.5,” 2026. [Online]. Available: https://qwen.ai/blog? id=qwen3.5

  66. [66]

    Huatuogpt, towards taming language model to be a doctor,

    H. Zhang, J. Chen, F. Jiang, F. Yu, Z. Chen, G. Chen, J. Li, X. Wu, Z. Zhiyi, Q. Xiaoet al., “Huatuogpt, towards taming language model to be a doctor,” inFindings of the association for computational linguistics: EMNLP 2023, 2023, pp. 10 859–10 885

  67. [67]

    Hulu-med: A transparent generalist model towards holistic medical vision-language understanding,

    S. Jiang, Y . Wang, S. Song, T. Hu, C. Zhou, B. Pu, Y . Zhang, Z. Yang, Y . Feng, J. T. Zhouet al., “Hulu-med: A transparent generalist model towards holistic medical vision-language understanding,”arXiv preprint arXiv:2510.08668, 2025