REVIEW 3 major objections 1 minor 54 references
Understanding How Humans Inject Knowledge into Machine Learning Workflows through Visual Analytics
T0 review · 3 major / 1 minor · reviewed 2026-07-02 · grok-4.3
Pith's one-line read A survey of over 200 visualization papers identifies specific pathways for humans to inject knowledge into machine learning workflows via interactive tools.
desk verdict A structured survey of VIS4ML papers that maps knowledge-injection pathways but overstates its results as 'unequivocal evidence.' read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The coding scheme applied to four perspectives (characteristics of ML, visualization, interaction, and actions) that traces pathways of human knowledge transfer into ML workflows.
What would settle it
A search of additional papers outside the IEEE VIS corpus that documents common human knowledge injection methods into ML workflows not covered by the coding scheme would undermine the completeness of the identified pathways.
Extended reading notes
Core claim
By coding the collected papers from four perspectives, the analysis reveals different pathways that transfer human knowledge to ML workflows via interactive visualization. These pathways are explained using the conceptual model that views VA as model building and the information-theoretic cost-benefit analysis that reasons VA as for optimizing ML workflows, providing evidence of the merits of using VA in ML workflows.
Load-bearing premise
The corpus of IEEE VIS papers from the past decade together with the developed coding scheme sufficiently captures the representative ways humans inject knowledge into ML workflows through interactive visualization.
Editorial extensions
If this is right
- Distinct pathways exist for transferring human knowledge to ML workflows through interactive visualization.
- Visual analytics functions as model building in these workflows.
- Information-theoretic cost-benefit analysis positions visual analytics as a means to optimize ML workflows.
- The survey supplies evidence for the merits of visual analytics within machine learning workflows.
Reading between the lines
- Tool builders could design new visual analytics interfaces that explicitly support the identified pathways at different stages of machine learning.
- The coding scheme could be tested on papers from additional conferences to check whether the same pathways appear consistently.
- Measuring actual workflow efficiency gains when using these pathways in practice would provide a direct test of the cost-benefit reasoning.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript surveys over 200 VIS4ML papers from IEEE VIS conferences (2013–2022), develops a four-axis coding scheme covering ML characteristics, visualization, interaction, and actions, identifies pathways by which humans inject knowledge into ML workflows via interactive visualization, and interprets the results through conceptual models of VA as model building and information-theoretic cost-benefit analysis. It concludes that the survey supplies unequivocal evidence of the merits of VA in ML workflows, with the coded corpus and results released online.
Significance. If the coding scheme is shown to be reliable, the work would supply a useful taxonomy of human-in-the-loop patterns in visualization-supported ML, helping to organize an emerging subfield. The public release of the full coded dataset and analysis artifacts is a clear strength that supports reproducibility and secondary use. The significance is nevertheless tempered by the observational character of the study and the absence of validation metrics for the central methodological choices.
major comments (3)
- [Abstract and §6] Abstract and §6: The assertion that the work 'provides unequivocal evidence showing the merits of using VA in ML workflows' is not warranted by the reported methodology. The study performs qualitative coding of a venue-restricted corpus without reported validation or quantitative outcome measures; this observational design cannot supply unequivocal evidence and the claim should be moderated or supported by additional validation.
- [§3] §3 (Corpus construction): Restricting the corpus to IEEE VIS papers from a single decade introduces venue and publication bias that is not discussed. The paper should explain why other venues publishing VIS4ML work were excluded and how this choice affects the representativeness of the observed pathways.
- [§4] §4 (Coding scheme): No information is given on inter-rater reliability, pilot testing, or external validation of the four-axis coding scheme. Because the pathways, conceptual models, and final claims rest directly on the coded data, the lack of these standard methodological safeguards is load-bearing.
minor comments (1)
- [§3] The supplementary website is referenced but the main text would benefit from a compact summary table reporting basic corpus statistics (e.g., paper counts per year and per coding axis) to allow readers to assess coverage without leaving the PDF.
Simulated Author's Rebuttal
We thank the referee for their detailed and constructive comments. We address each major point below, indicating revisions where appropriate.
read point-by-point responses
-
Referee: [Abstract and §6] Abstract and §6: The assertion that the work 'provides unequivocal evidence showing the merits of using VA in ML workflows' is not warranted by the reported methodology. The study performs qualitative coding of a venue-restricted corpus without reported validation or quantitative outcome measures; this observational design cannot supply unequivocal evidence and the claim should be moderated or supported by additional validation.
Authors: We agree that 'unequivocal evidence' overstates what an observational qualitative survey can demonstrate. We will revise the abstract and §6 to moderate the language (e.g., 'provides evidence illustrating the merits' or 'offers insights into the role of VA'), while adding an explicit acknowledgment of the study's observational and qualitative character and the absence of quantitative outcome metrics. revision: yes
-
Referee: [§3] §3 (Corpus construction): Restricting the corpus to IEEE VIS papers from a single decade introduces venue and publication bias that is not discussed. The paper should explain why other venues publishing VIS4ML work were excluded and how this choice affects the representativeness of the observed pathways.
Authors: We will expand §3 with a new paragraph explaining the rationale: IEEE VIS is the flagship venue for visualization research, ensuring high-quality, peer-reviewed VIS4ML contributions within a consistent community. We will also acknowledge the resulting venue bias and limited representativeness, noting that VIS4ML work appears in other venues (e.g., EuroVis, PacificVis) and that expanding the corpus in future work would improve generalizability of the observed pathways. revision: yes
-
Referee: [§4] §4 (Coding scheme): No information is given on inter-rater reliability, pilot testing, or external validation of the four-axis coding scheme. Because the pathways, conceptual models, and final claims rest directly on the coded data, the lack of these standard methodological safeguards is load-bearing.
Authors: We will revise §4 to describe the coding process in greater detail: the scheme was developed iteratively by the author team through discussion and refinement; a pilot phase was conducted on a subset of papers to test and adjust categories; and disagreements were resolved via consensus meetings. We will note that formal inter-rater reliability statistics (e.g., Cohen's kappa) were not computed and acknowledge this as a methodological limitation, while arguing that the consensus-based approach provided sufficient consistency for the exploratory goals of the survey. revision: partial
Circularity Check
No circularity: survey of external literature with independent coding
full rationale
The paper performs a literature survey of over 200 external IEEE VIS papers (2013-2022), develops a four-axis coding scheme, and qualitatively identifies pathways. No mathematical derivations, parameter fitting, predictions, or self-referential equations exist. The central claim rests on observed patterns in the collected corpus rather than any reduction to the paper's own inputs or prior self-citations. The analysis is self-contained against external benchmarks (published papers) with no load-bearing self-citation chains or ansatzes imported from the authors' own work.
Assumptions & free parameters
assumptions (1)
- domain assumption The corpus of IEEE VIS papers from the past decade adequately represents VIS4ML research on human knowledge injection.
Cite this review
Pith. "Pith review of Understanding How Humans Inject Knowledge into Machine Learning Workflows through Visual Analytics." pith.science (2026). https://pith.science/paper/URBLI34C
@misc{pith2026260700969,
author = {Pith},
title = {Pith review of: Understanding How Humans Inject Knowledge into Machine Learning Workflows through Visual Analytics},
year = {2026},
howpublished = {\url{https://pith.science/paper/URBLI34C}},
note = {Machine review of arXiv:2607.00969}
}
read the original abstract
Visual analytics (VA) plays an increasingly important role in supporting machine learning (ML) workflows. In the field of visualization, such approaches and techniques are referred to as VIS4ML. While ML models are mostly learned automatically, the corresponding ML workflows receive a variety of human inputs, such as data labelling, feature engineering, model architecture designing, hyper-parameter tuning, and so on. In this work, we surveyed over 200 VIS4ML papers to gain an understanding of how humans inject their knowledge into ML workflows through interactive visualization. We collected a corpus of VIS4ML papers from the IEEE VIS conferences in the past decade. We developed a coding scheme to facilitate the literature research from four perspectives: characteristics of ML, visualization, interaction, and actions. The analysis of the coded dataset allows us to observe different pathways that transfer human knowledge to ML workflows via interactive visualization. Building on the analysis, we explain the phenomena of VIS4ML using the conceptual model that views VA as model building and the information-theoretic cost-benefit analysis that reasons VA as for optimizing ML workflows. This work provides unequivocal evidence showing the merits of using VA in ML workflows. The full list of surveyed papers, along with all analysis results and figures, is available at https://vis4ml4hd.github.io/ml-knowledge-inject-va/.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
S. Amershi, A. Begel, C. Bird, R. DeLine, H. Gall, E. Kamar et al. Soft- ware engineering for machine learning: A case study. In 2019 IEEE/ACM 41st International Conference on Software Engineering: Software Engi- neering in Practice (ICSE-SEIP) , pp. 291–300, 2019. 2, 18
work page 2019
-
[2]
G. Andrienko, N. Andrienko, and D. Hecker. Topic modelling for spa- tial insights: Uncovering space use from movement data. Computers & Graphics, 122:103989, 2024. 6
work page 2024
-
[3]
N. Andrienko, T. Lammarsch, G. Andrienko, G. Fuchs, D. Keim, S. Miksch et al. Viewing visual analytics as model building. Computer Graphics F orum, 37(6):275–299, 2018. 2, 8, 9
work page 2018
-
[4]
C. M. Bishop and N. M. Nasrabadi. Pattern recognition and machine learning, vol. 4. 2006. 3
work page 2006
-
[5]
C. Chai and G. Li. Human-in-the-loop techniques in machine learning. IEEE Data Eng. Bull. , 43(3):37–52, 2020. 2
work page 2020
- [6]
-
[7]
A. Chatzimparmpas, R. M. Martins, I. Jusufi, K. Kucher, F. Rossi, and A. Kerren. The state of the art in enhancing trust in machine learn- ing models with the use of visualizations. Computer Graphics F orum, 39:713–756, 6 2020. 2
work page 2020
-
[8]
A. Chatzimparmpas, R. M. Martins, and A. Kerren. Visruler: Visual analytics for extracting decision rules from bagged and boosted decision trees. Information Visualization, 22(2):115–139, 2023. 2
work page 2023
Show all 54 references
-
[9]
M. Chen. Cost-benefit analysis of data intelligence – its broader inter- pretations. In Advances in Info-Metrics: Information and Information Processing across Disciplines, pp. 433–463. 2020. 2, 9
2020
-
[10]
Chen and D
M. Chen and D. J. Edwards. ‘isms’ in visualization. In F oundations of Data Visualization, pp. 225–241. 2020. 2
2020
-
[11]
M. Chen, L. Floridi, and R. Borgo. What is visualization really for? In The Philosophy of Information Quality , vol. 358, pp. 7593. 2014. 2
2014
-
[12]
Chen and A
M. Chen and A. Golan. What may visualization processes opti- mize? IEEE Transactions on Visualization and Computer Graphics , 22(12):2619–2632, 2016. 2, 9, 10
2016
-
[13]
M. Chen, K. Mueller, and A. Ynnerman. Fusion of Visual Channels , pp. 119–127. Springer London, London, 2014. 3
2014
-
[14]
Diehl, A
A. Diehl, A. Abdul-Rahman, B. Bach, M. El-Assady, M. Kraus, R. S. Laramee et al. An analysis of the interplay and mutual benefits of grounded theory and visualization. IEEE Transactions on Visualization and Computer Graphics, 31(9):5462–5479, 2025. 3
2025
-
[15]
J. J. Dudley and P . O. Kristensson. A review of user interface design for interactive machine learning. ACM Transactions on Interactive Intelligent Systems (TiiS), 8(2):1–37, 2018. 2
2018
-
[16]
Endert, W
A. Endert, W. Ribarsky, C. Turkay, B. W. Wong, I. Nabney, I. D. Blanco et al. The state of the art in integrating machine learning into visual analytics. Computer Graphics F orum, 36:458–486, 12 2017. 2
2017
-
[17]
S. D. H. Evergreen. Effective Data Visualization: The Right Chart for the Right Data. SAGE Publications, 2016. 2
2016
-
[18]
Fernández-Delgado, E
M. Fernández-Delgado, E. Cernadas, S. Barro, and D. Amorim. Do we need hundreds of classifiers to solve real world classification problems? The journal of machine learning research , 15(1):3133–3181, 2014. 3
2014
-
[19]
Gómez-Carmona, D
O. Gómez-Carmona, D. Casado-Mansilla, D. Lopez-de Ipina, and J. García-Zubia. Human-in-the-loop machine learning: Reconceptual- izing the role of the user in interactive approaches. Internet of Things , 25:101048, 2024. 2
2024
-
[20]
Goodfellow, Y
I. Goodfellow, Y . Bengio, A. Courville, and Y . Bengio. Deep learning , vol. 1. 2016. 3
2016
-
[21]
Hohman, M
F. Hohman, M. Kahng, R. Pienta, and D. H. Chau. Visual analytics in deep learning: An interrogative survey for the next frontiers. IEEE Trans- actions on Visualization and Computer Graphics , 25:2674–2693, 8 2019. 2
2019
-
[22]
Holzinger
A. Holzinger. Interactive machine learning for health informatics: when do we need the human-in-the-loop? Brain informatics , 3(2):119–131,
-
[23]
S. B. Kotsiantis, I. Zaharakis, P . Pintelas, et al. Supervised machine learn- ing: A review of classification techniques. Emerging artificial intelli- gence applications in computer engineering , 160(1):3–24, 2007. 3
2007
-
[24]
Kumar, S
S. Kumar, S. Datta, V . Singh, D. Datta, S. Kumar Singh, and R. Sharma. Applications, challenges, and future directions of human-in-the-loop learning. IEEE Access, 12:75735–75760, 2024. 2
2024
-
[25]
S. Liu, J. Xiao, J. Liu, X. Wang, J. Wu, and J. Zhu. Visual diagnosis of tree boosting methods. IEEE Trans. Vis. Comp. Graph. , 24(1):163–173,
-
[26]
R. Marty. Applied Security Visualization. Addison-Wesley, 2009. 2
2009
-
[27]
Moussavi, G
L. Moussavi, G. Andrienko, N. Andrienko, and A. Slingsby. Visually- supported topic modeling for understanding behavioral patterns from spa- tiotemporal events. Computers & Graphics, 129:104245, 2025. 6
2025
-
[28]
O’Callaghan, D
D. O’Callaghan, D. Greene, J. Carthy, and P . Cunningham. An analysis of the coherence of descriptors in topic modeling. Expert Syst. Appl. , 42(13):56455657, Aug. 2015. 5
2015
-
[29]
H. Park, N. Das, R. Duggal, A. P . Wright, O. Shaikh, F. Hohman et al. Neurocartography: Scalable automatic visual summarization of concepts in deep neural networks. IEEE Transactions on Visualization and Com- puter Graphics, 28(1):813823, Jan. 2022. 7
2022
-
[30]
Sacha, M
D. Sacha, M. Kraus, J. Bernard, M. Behrisch, T. Schreck, Y . Asano et al. Somflow: Guided exploratory cluster analysis with self-organizing maps and analytic provenance. IEEE Transactions on Visualization and Computer Graphics, 24(1):120–130, 2018. 7
2018
-
[31]
Sacha, M
D. Sacha, M. Kraus, D. A. Keim, and M. Chen. Vis4ml: An ontology for visual analytics assisted machine learning. IEEE Transactions on Visual- ization and Computer Graphics , 25(1):385–395, 2019. 1, 2, 3
2019
-
[32]
Sculley, G
D. Sculley, G. Holt, D. Golovin, E. Davydov, T. Phillips, D. Ebner et al. Hidden technical debt in machine learning systems. Advances in neural information processing systems, 28, 2015. 2, 18
2015
-
[33]
Sevastjanova, S
R. Sevastjanova, S. V ogelbacher, A. Spitz, D. Keim, and M. El-Assady. Visual comparison of text sequences generated by large language models. In IEEE Visualization in Data Science (VDS) , pp. 11–20, 2023. 2
2023
-
[34]
Q. Shen, Y . Wu, Y . Jiang, W. Zeng, K. Alexis, A. Vianova et al. Visual in- terpretation of recurrent neural network on multi-dimensional time-series forecast. In 2020 IEEE Pacific visualization symposium (PacificVis) , pp. 61–70, 2020. 2
2020
-
[35]
Sperrle, M
F. Sperrle, M. ElÄêAssady, G. Guo, R. Borgo, D. H. Chau, A. Endert et al. A Survey of Human-Centered Evaluations in Human-Centered Machine Learning. Computer Graphics F orum, 40:543–568, 6 2021. 2
2021
-
[36]
J. T. Stasko. V alue-driven evaluation of visualizations. In Proc. 5th Work- shop on Beyond Time and Errors: Novel Evaluation Methods for Visual- ization (BELIV). 2014. 2
2014
-
[37]
Streeb, M
D. Streeb, M. El-Assady, D. A. Keim, and M. Chen. Visualize? argu- ments for visual support in decision making. IEEE Computer Graphics and Applications, 41(2):17–22, 2021. 2
2021
-
[38]
Streeb, M
D. Streeb, M. El-Assady, D. A. Keim, and M. Chen. Why visualize? untangling a large network of arguments. IEEE Transactions on Visual- ization and Computer Graphics , 27(3):2220–2236, 2021. 2
2021
-
[39]
Strobelt, A
H. Strobelt, A. Webson, V . Sanh, B. Hoover, J. Beyer, H. Pfister et al. In- teractive and visual prompt engineering for ad-hoc task adaptation with large language models. IEEE Transactions on Visualization and Com- puter Graphics, pp. 1–11, 2022. 2
2022
-
[40]
A. Suh, A. Mosca, E. Wu, and R. Chang. A grammar of hypotheses for visualization, data, and analysis, 2023. 2
2023
-
[41]
G. K. Tam, V . Kothari, and M. Chen. An analysis of machine-and human- analytics in classification. IEEE Trans. Vis. Comp. Graph. , 23:71–80,
-
[42]
E. R. Tufte. The Visual Display of Quantitative Information . 2nd ed.,
-
[43]
J. J. van Wijk. The value of visualization. In Proc. IEEE Visualization, pp. 79–86. 2005. 2
2005
-
[44]
J. Wang, S. Hazarika, C. Li, and H.-W. Shen. Visualization and visual analysis of ensemble data: A survey. IEEE Transactions on Visualization and Computer Graphics, 25(9):2853–2872, 2019. 1
2019
-
[45]
J. Wang, S. Liu, and W. Zhang. Visual analytics for machine learning: A data perspective survey. IEEE Transactions on Visualization and Com- puter Graphics, 30(12):7637–7656, 2024. 1, 2
2024
-
[46]
J. Wang, W. Zhang, H. Y ang, C.-C. M. Y eh, and L. Wang. Visual ana- lytics for rnn-based deep reinforcement learning. IEEE Trans. Vis. Comp. Graph., 28(12):4141–4155, 2021. 2
2021
-
[47]
Q. Wang, W. Alexander, J. Pegg, H. Qu, and M. Chen. Hypoml: Vi- sual analysis for hypothesis-based evaluation of machine learning models. IEEE Trans. Vis. Comp. Graph., 27(2):1417–1426, 2020. 2
2020
-
[48]
C. Ware. Information Visualization: Perception for Design. 2nd ed., 2004. 2
2004
-
[49]
Wexler, M
J. Wexler, M. Pushkarna, T. Bolukbasi, M. Wattenberg, F. Viegas, and J. Wilson. The what-if tool: Interactive probing of machine learning mod- els. IEEE Transactions on Visualization and Computer Graphics, pp. 1–1,
-
[50]
X. Wu, L. Xiao, Y . Sun, J. Zhang, T. Ma, and L. He. A survey of human- in-the-loop for machine learning. Future Generation Computer Systems , 135:364–381, 2022. 2
2022
-
[51]
D. Xin, L. Ma, J. Liu, S. Macke, S. Song, and A. Parameswaran. Acceler- ating human-in-the-loop machine learning: Challenges and opportunities. In Proceedings of the second workshop on data management for end-to- end machine learning, pp. 1–4, 2018. 2
2018
-
[52]
Y uan, C
J. Y uan, C. Chen, W. Y ang, M. Liu, J. Xia, and S. Liu. A survey of visual analytics techniques for machine learning. Computational Visual Media, 7:3–36, 3 2021. 1, 2 APPENDICES Understanding How Humans Inject Knowledge into Machine Learning Workflows through Visual Analytics ...
2021
-
[53]
Logical/Symbolic Models; 3
Linear Models; 2. Logical/Symbolic Models; 3. Statistical/Graphical Models; 4. Instance-Based Models; 5. Kernel Methods; 6. Neural Models; 7. Sequence Models; 8. Decision Tree-Based Models
-
[54]
steering
Autoencoder-Based Models; 10. Perceptron-Based Models; 11. Ensemble Methods; 12. Evolutionary Methods; 13. Others *1.5 ML_TASK: Target problems addressed by the model. 1.Classification; 2.Detection/Recognition; 3.Regression; 4.Clustering; 5.Association; 6.Feature Selec- tion; 7...
2016
Reviewed July 2, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.