REVIEW 2 major objections 5 minor 12 references
AI-in-the-loop: The future of biomedical visual analytics applications in the era of AI
T0 review · 2 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read This paper argues that biomedical visual analytics should put AI in the loop while keeping human experts in charge of agency and responsibility.
desk verdict A coherent, useful roadmap for biomedical VA with AI-in-the-loop, but its engine is a bet on generative AI capabilities that the paper itself admits lack training data; the comparative adoption claim is unsupported. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The paper's central organizing device is the 'AI-in-the-loop' workflow model, together with an eight-task taxonomy of AI-supported pipeline steps and an automation spectrum from fully human-generated to fully AI-generated tools. The taxonomy (Data2Data, Data2Code, Data2Image, Data2VisComponent, Task2VisConfig, VisComponent2Tool, Task2Query, Task2Interact) defines where AI enters the classic visual analytics pipeline, while the spectrum shows how these steps can combine into systems of varying automation. This machinery carries the argument by giving researchers a concrete language for describing and building future AI-assisted biomedical VA tools, and by making explicit that a fully end-to-end AI tool would require regeneration rather than reconfiguration when intent changes.
What would settle it
A prospective comparison study would settle the claim: if an end-to-end AI-generated visualization tool with no human audit matches or exceeds the reliability of an AI-in-the-loop tool with a human expert in real biomedical decision tasks, for example in tumor-board treatment planning, then the paper's insistence on human agency and visual analytics as the integrating interface would lose its empirical basis.
Extended reading notes
Core claim
The central claim is that agency and responsibility must remain with human experts in biomedical decision-making, and that visual analytics therefore should be used as a tool for integrating AI into human-centered workflows, which is the reverse of the common 'human-in-the-loop' framing that inserts humans into AI systems. On this view, AI becomes an integral part of the visual analytics loop: it assists with data processing (Data2Data), generates visualization code and components (Data2Code, Data2VisComponent), creates images directly from data (Data2Image), configures and assembles tools (Task2VisConfig, VisComponent2Tool), and mediates natural-language queries and interactions (Task2Query, Task2Interact). The paper argues that such AI-supported workflows can be mixed at various automation levels, and that even near-future generative systems, despite hallucinations and reliability limits, will make low-code biomedical VA tool development feasible while human experts audit each result. The claimed consequence is that biomedical visual analytics will remain central, and that the approach could help close the gap between approved AI radiology products and their actual 2% market penetration.
Load-bearing premise
The roadmap assumes that generative AI and large language models will become capable of correctly and reliably producing custom visualization code, direct data-to-image transformations, natural-language queries, and intent-driven interactions for complex biomedical data, despite the paper's own list of obstacles (scarce training data, proprietary VA code, hallucination, and hard-to-validate expert interactions).
Editorial extensions
If this is right
- Tool builders will be able to mix and match AI-supported stages, for instance a human-written component with an AI-generated configuration, or an AI-generated view with a human-crafted dashboard.
- For high-stakes uses, every AI-generated component must remain inspectable and editable because the human expert is ultimately responsible, and closed end-to-end generation is the least desirable end of the spectrum.
- Task2Query and Task2Interact will let users steer complex biomedical data with natural language and predicted intentions, but only when the underlying AI can reliably interpret high-level structures such as anatomy.
- VA researchers will spend less time on routine code generation and more on new algorithms, interface designs, and validation tools, since standard components can be produced by assistive AI.
- The biomedical field will adopt AI-assisted VA more slowly than recreational AI applications because training data and code are scarce and mostly proprietary.
Reading between the lines
- The taxonomy could be turned into a benchmark: each of the eight arrows names a capability that can be evaluated separately, giving the community measurable milestones for how close generative AI is to realizing the roadmap.
- The 'regeneration problem' the paper identifies, where intent-driven changes to an end-to-end tool require full regeneration, suggests a design principle that component-based, configurable AI tools will likely dominate in practice over monolithic end-to-end generation.
- The AI-in-the-loop framing may generalize beyond biomedicine to other regulated, high-stakes domains such as financial auditing or public-safety analytics, where explainability and human sovereignty are similarly mandated.
- A direct test of the roadmap would be to implement a small biomedical VA scenario, for example a multi-modal tumor-board view, and measure whether current LLMs can produce the Data2Code and Task2Query stages with acceptable reliability; the paper does not perform this test.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This Visualization Viewpoint argues that in biomedical data analytics, the increasing power of generative AI and LLMs will transform visual analytics workflows, but that agency and responsibility must remain with human experts. The authors propose an 'AI-in-the-loop' paradigm in which visual analytics serves as an interface integrating AI into human-centered workflows, in contrast to the traditional 'human-in-the-loop' framing. They present a workflow taxonomy (Data2Data, Data2Code, Data2Image, Data2VisComponent, Task2VisConfig, VisComponent2Tool, Task2Query, Task2Interact) and a spectrum from fully human-generated to end-to-end AI-generated VA tools, and discuss ethical/legal considerations and a set of research opportunities.
Significance. The central normative claim that human-centered VA should remain central in critical biomedical decision-making is coherent and consistent with current regulatory and ethical frameworks. The paper's explicit list of challenges (scarce biomedical training data, proprietary VA code, hard-to-validate expert interactions) is a useful contribution, and the proposed taxonomy may serve as a design framework. If the feasibility assumptions are met, the viewpoint provides a clear agenda for the visualization community. However, the paper does not provide evidence or a mechanism for the key capability assumptions underlying the roadmap, and one comparative prediction in the conclusion is unsupported. The contribution is a viewpoint worth publishing after revision, not a result with validated evidence.
major comments (2)
- [The future of VA workflows / Challenges in Biomedical Data Analytics] The roadmap in Figure 2 rests on the assumption that generative AI will become able to produce correct, task-appropriate visualization components and interactions for complex biomedical data. The paper asserts this repeatedly, e.g., 'As LLMs evolve and specialize, direct generation of custom visualization code adapted to the current dataset will be feasible' and 'the process of creating a tool by mixing and linking different visualization components ... will be achievable in the near future', but the same paper's Section 'Challenges in Biomedical Data Analytics', items (5)-(7), identifies that biomedical training data are scarce, that most VA code is proprietary and unavailable for training, and that expert interactions are highly variable and difficult to validate. These are precisely the resources needed to train the proposed generative capabilities. The text does not offer a mechanism, a data strategy, or benchmark evidence that this bottleneck can be overcome. As written, this is a load-bearing unsupported conjecture rather than an argued roadmap. Please either provide evidence or a concrete strategy for overcoming the data/code/interaction bottleneck, or reframe these statements as open research questions and explicitly discuss the dependence of Figure 2 on this capability assumption.
- [Conclusion] The concluding claim that 'our proposed AI-in-the-loop approach ... will more effectively close this gap than human-in-the-loop approaches' is an unsupported comparative prediction. Reference [11] reports current market penetration of radiology AI products (~2%) and discusses barriers to adoption; it does not provide evidence comparing the effectiveness of AI-in-the-loop versus human-in-the-loop paradigms. The manuscript offers no data, case study, or analytical argument that would support this comparative claim. Please either remove the comparison or reframe it explicitly as an untested hypothesis, since as stated it exceeds what the cited material and the paper's own argument can support.
minor comments (5)
- [Visual Analytics Workflows: present and future] In the 'Current VA status' subsection, 'thehuman-in-the-loop' should read 'the human-in-the-loop'.
- [Challenges in Biomedical Data Analytics] In item (5), 'the the development' contains a duplicated 'the'; it should read 'the development'.
- [The future of VA workflows] The naming of the task is inconsistent: the list uses 'Data2Viscomponent' while Figure 2 and later text use 'Data2VisComp'; please standardize.
- [Dangers, Ethical, and Legal Considerations] The phrase 'Metas' llama3' should be written as 'Meta's Llama 3' for readability and correctness.
- [Figure 3 caption] The caption uses 'Task2Config' and 'Component2Tool' while the body uses 'Task2VisConfig' and 'VisComponent2Tool'; please align the shorthand with the taxonomy terminology.
Circularity Check
No circularity: the paper is a viewpoint essay whose claims are proposals and literature-based arguments, not derivations that reduce to their own inputs.
full rationale
This is a Visualization Viewpoint essay rather than a technical derivation, so there is no chain of equations or fitted parameters whose outputs coincide with inputs. The paper's central proposal, "AI-in-the-loop," is introduced by definition and contrasted with "human-in-the-loop," but the surrounding argument depends on cited external literature (e.g., Keim et al. for the human-centered definition, Rajpurkar and Lungren for the 2% radiology AI market penetration), not on the authors' own prior results. No self-citation is load-bearing, and no uniqueness theorem or prior result by the same authors is invoked to force a choice. The taxonomy in Figure 2 (Data2Code, Data2Image, etc.) is a speculative organizational scheme for possible future AI capabilities, not a prediction that is mathematically forced by training data or by the taxonomy itself. The conclusion's expectation that AI-in-the-loop will close the radiology adoption gap more effectively than human-in-the-loop is an unsupported comparative prediction, but it is not circular: the cited 2% statistic is external evidence about current adoption, not a quantity recycled into the claim. Under the hard rules, a paper that makes no derivations or fitted predictions and is self-contained against external references should receive a score of 0, and no specific circular step can be exhibited.
Assumptions & free parameters
assumptions (4)
- domain assumption Agency and responsibility must remain with human experts in high-stakes biomedical decision making.
- ad hoc to paper Future generative and large language models will become capable of correctly generating custom visualization code and interactions for complex biomedical data.
- domain assumption Scarcity of public, labelled VA tool code and data will persist and constrain generative model training.
- domain assumption Visual Analytics will remain necessary for auditability, explainability, and uncertainty assessment of AI outputs.
invented entities (2)
-
AI-in-the-loop paradigm
-
AI companions
Cite this review
Pith. "Pith review of AI-in-the-loop: The future of biomedical visual analytics applications in the era of AI." pith.science (2026). https://pith.science/paper/RIXYN2ZE
@misc{pith2026241215876,
author = {Pith},
title = {Pith review of: AI-in-the-loop: The future of biomedical visual analytics applications in the era of AI},
year = {2026},
howpublished = {\url{https://pith.science/paper/RIXYN2ZE}},
note = {Machine review of arXiv:2412.15876}
}
read the original abstract
AI is the workhorse of modern data analytics and omnipresent across many sectors. Large Language Models and multi-modal foundation models are today capable of generating code, charts, visualizations, etc. How will these massive developments of AI in data analytics shape future data visualizations and visual analytics workflows? What is the potential of AI to reshape methodology and design of future visual analytics applications? What will be our role as visualization researchers in the future? What are opportunities, open challenges and threats in the context of an increasingly powerful AI? This Visualization Viewpoint discusses these questions in the special context of biomedical data analytics as an example of a domain in which critical decisions are taken based on complex and sensitive data, with high requirements on transparency, efficiency, and reliability. We map recent trends and developments in AI on the elements of interactive visualization and visual analytics workflows and highlight the potential of AI to transform biomedical visualization as a research field. Given that agency and responsibility have to remain with human experts, we argue that it is helpful to keep the focus on human-centered workflows, and to use visual analytics as a tool for integrating ``AI-in-the-loop''. This is in contrast to the more traditional term ``human-in-the-loop'', which focuses on incorporating human expertise into AI-based systems.
Reference graph
Works this paper leans on
- [11]
-
[1]
11em plus .33em minus .07em 4000 4000 100 4000 4000 500 `\.=1000 = #1 9 12 #1 * [1] 0pt [0pt][0pt] #1 * \| ** #1 \@IEEEauthorblockNstyle \@IEEEcompsocnotconfonly \@IEEEcompsocconfonly \@IEEEauthorblockAstyle \@IEEEcompsocnotconfonly \@IEEEcompsocconfonly \@IEEEcompsocconfonly \@IEEEauthordefaulttextstyle \@IEEEcompsocnotconfonly \@IEEEauthorblockconfadjsp...
-
[2]
A. Wu, D. Deng, M. Chen, S. Liu, D. Keim, R. Maciejewski, S. Miksch, H. Strobelt, F. Vi \'e gas, and M. Wattenberg, ``Grand challenges in visual analytics applications,'' IEEE Computer Graphics and Applications, vol. 43, no. 5, pp. 83--90, 2023
work page 2023
-
[3]
OpenAI, ``Gpt-4 technical report,'' OpenAI, Tech. Rep. gpt4-report@openai.com, 3 2023
work page 2023
- [4]
-
[5]
J. Abramson, J. Adler, J. Dunger, R. Evans, T. Green, A. Pritzel, O. Ronneberger, L. Willmore, A. J. Ballard, J. Bambrick, S. W. Bodenstein, D. A. Evans, C.-C. Hung, M. O’Neill, D. Reiman, K. Tunyasuvunakool, Z. Wu, A. Žemgulytė, E. Arvaniti, C. Beattie, O. Bertolli, A. Bridgland, A. Cherepanov, M. Congreve, A. I. Cowen-Rivers, A. Cowie, M. Figurnov, F. B...
work page 2024
-
[6]
D. Keim, G. Andrienko, J.-D. Fekete, C. G\" o rg, J. Kohlhammer, and G. Melan c on, Visual Analytics: Definition, Process, and Challenges. 1em plus 0.5em minus 0.4em Springer Berlin Heidelberg, 2007, p. 154–175
work page 2007
-
[7]
V. C. Dibia, ``Lida: A tool for automatic generation of grammar-agnostic visualizations and infographics using large language models,'' Annual Meeting of the Association for Computational Linguistics, 2023
work page 2023
Show all 12 references
-
[8]
Bubeck, V
S. Bubeck, V. Chandrasekaran, R. Eldan, J. Gehrke, E. Horvitz, E. Kamar, P. Lee, Y. T. Lee, Y. Li, S. Lundberg, H. Nori, H. Palangi, M. T. Ribeiro, and Y. Zhang, ``Sparks of artificial general intelligence: Early experiments with gpt-4,'' 2023. [Online]. Available: https://arx...
2023 arXiv
-
[9]
M. Moor, O. Banerjee, Z. S. H. Abad, H. M. Krumholz, J. Leskovec, E. J. Topol, and P. Rajpurkar, ``Foundation models for generalist medical artificial intelligence,'' Nature, vol. 616, no. 7956, pp. 259--265, Apr. 2023
2023
-
[10]
E. M. Bender, T. Gebru, A. McMillan-Major, and S. Shmitchell, ``On the dangers of stochastic parrots: Can language models be too big?'' in Proceedings of the 2021 ACM conference on fairness, accountability, and transparency, 2021, pp. 610--623
2021
-
[12]
Rajpurkar and M
P. Rajpurkar and M. P. Lungren, ``The current and future state of AI interpretation of medical images,'' New England Journal of Medicine, vol. 388, no. 21, pp. 1981--1990, 2023
1981
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.