Pith. sign in

REVIEW 3 major objections 1 minor 54 references

Understanding How Humans Inject Knowledge into Machine Learning Workflows through Visual Analytics

T0 review · 3 major / 1 minor · reviewed 2026-07-02 · grok-4.3

Pith's one-line read A survey of over 200 visualization papers identifies specific pathways for humans to inject knowledge into machine learning workflows via interactive tools.

desk verdict A structured survey of VIS4ML papers that maps knowledge-injection pathways but overstates its results as 'unequivocal evidence.' read the letter →

arxiv 2607.00969 v1 pith:URBLI34C submitted 2026-07-01 cs.HC cs.LG

classification cs.HCcs.LG
keywords visualanalyticsmachinelearningworkflowshumanknowledgeinjectionVIS4MLinteractivevisualizationliteraturesurveycodingschemetransferpathways
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper surveys more than 200 VIS4ML papers from IEEE VIS conferences over the past decade to map how humans add knowledge to machine learning through visual analytics. It builds a coding scheme across machine learning traits, visualization, interaction, and actions to trace distinct pathways of knowledge transfer. The results are explained through a model that treats visual analytics as model building and through cost-benefit reasoning that positions it as a way to optimize machine learning workflows. A reader would care because the work shows concrete ways visualization makes machine learning more responsive to human insight rather than fully automatic.

What carries the argument

The coding scheme applied to four perspectives (characteristics of ML, visualization, interaction, and actions) that traces pathways of human knowledge transfer into ML workflows.

What would settle it

A search of additional papers outside the IEEE VIS corpus that documents common human knowledge injection methods into ML workflows not covered by the coding scheme would undermine the completeness of the identified pathways.

Watch

Extended reading notes

Core claim

By coding the collected papers from four perspectives, the analysis reveals different pathways that transfer human knowledge to ML workflows via interactive visualization. These pathways are explained using the conceptual model that views VA as model building and the information-theoretic cost-benefit analysis that reasons VA as for optimizing ML workflows, providing evidence of the merits of using VA in ML workflows.

Load-bearing premise

The corpus of IEEE VIS papers from the past decade together with the developed coding scheme sufficiently captures the representative ways humans inject knowledge into ML workflows through interactive visualization.

Editorial extensions

If this is right

  • Distinct pathways exist for transferring human knowledge to ML workflows through interactive visualization.
  • Visual analytics functions as model building in these workflows.
  • Information-theoretic cost-benefit analysis positions visual analytics as a means to optimize ML workflows.
  • The survey supplies evidence for the merits of visual analytics within machine learning workflows.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Tool builders could design new visual analytics interfaces that explicitly support the identified pathways at different stages of machine learning.
  • The coding scheme could be tested on papers from additional conferences to check whether the same pathways appear consistently.
  • Measuring actual workflow efficiency gains when using these pathways in practice would provide a direct test of the cost-benefit reasoning.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

3 major / 1 minor

Summary. The manuscript surveys over 200 VIS4ML papers from IEEE VIS conferences (2013–2022), develops a four-axis coding scheme covering ML characteristics, visualization, interaction, and actions, identifies pathways by which humans inject knowledge into ML workflows via interactive visualization, and interprets the results through conceptual models of VA as model building and information-theoretic cost-benefit analysis. It concludes that the survey supplies unequivocal evidence of the merits of VA in ML workflows, with the coded corpus and results released online.

Significance. If the coding scheme is shown to be reliable, the work would supply a useful taxonomy of human-in-the-loop patterns in visualization-supported ML, helping to organize an emerging subfield. The public release of the full coded dataset and analysis artifacts is a clear strength that supports reproducibility and secondary use. The significance is nevertheless tempered by the observational character of the study and the absence of validation metrics for the central methodological choices.

major comments (3)
  1. [Abstract and §6] Abstract and §6: The assertion that the work 'provides unequivocal evidence showing the merits of using VA in ML workflows' is not warranted by the reported methodology. The study performs qualitative coding of a venue-restricted corpus without reported validation or quantitative outcome measures; this observational design cannot supply unequivocal evidence and the claim should be moderated or supported by additional validation.
  2. [§3] §3 (Corpus construction): Restricting the corpus to IEEE VIS papers from a single decade introduces venue and publication bias that is not discussed. The paper should explain why other venues publishing VIS4ML work were excluded and how this choice affects the representativeness of the observed pathways.
  3. [§4] §4 (Coding scheme): No information is given on inter-rater reliability, pilot testing, or external validation of the four-axis coding scheme. Because the pathways, conceptual models, and final claims rest directly on the coded data, the lack of these standard methodological safeguards is load-bearing.
minor comments (1)
  1. [§3] The supplementary website is referenced but the main text would benefit from a compact summary table reporting basic corpus statistics (e.g., paper counts per year and per coding axis) to allow readers to assess coverage without leaving the PDF.

Simulated Author's Rebuttal

3 responses · 0 unresolved

We thank the referee for their detailed and constructive comments. We address each major point below, indicating revisions where appropriate.

read point-by-point responses
  1. Referee: [Abstract and §6] Abstract and §6: The assertion that the work 'provides unequivocal evidence showing the merits of using VA in ML workflows' is not warranted by the reported methodology. The study performs qualitative coding of a venue-restricted corpus without reported validation or quantitative outcome measures; this observational design cannot supply unequivocal evidence and the claim should be moderated or supported by additional validation.

    Authors: We agree that 'unequivocal evidence' overstates what an observational qualitative survey can demonstrate. We will revise the abstract and §6 to moderate the language (e.g., 'provides evidence illustrating the merits' or 'offers insights into the role of VA'), while adding an explicit acknowledgment of the study's observational and qualitative character and the absence of quantitative outcome metrics. revision: yes

  2. Referee: [§3] §3 (Corpus construction): Restricting the corpus to IEEE VIS papers from a single decade introduces venue and publication bias that is not discussed. The paper should explain why other venues publishing VIS4ML work were excluded and how this choice affects the representativeness of the observed pathways.

    Authors: We will expand §3 with a new paragraph explaining the rationale: IEEE VIS is the flagship venue for visualization research, ensuring high-quality, peer-reviewed VIS4ML contributions within a consistent community. We will also acknowledge the resulting venue bias and limited representativeness, noting that VIS4ML work appears in other venues (e.g., EuroVis, PacificVis) and that expanding the corpus in future work would improve generalizability of the observed pathways. revision: yes

  3. Referee: [§4] §4 (Coding scheme): No information is given on inter-rater reliability, pilot testing, or external validation of the four-axis coding scheme. Because the pathways, conceptual models, and final claims rest directly on the coded data, the lack of these standard methodological safeguards is load-bearing.

    Authors: We will revise §4 to describe the coding process in greater detail: the scheme was developed iteratively by the author team through discussion and refinement; a pilot phase was conducted on a subset of papers to test and adjust categories; and disagreements were resolved via consensus meetings. We will note that formal inter-rater reliability statistics (e.g., Cohen's kappa) were not computed and acknowledge this as a methodological limitation, while arguing that the consensus-based approach provided sufficient consistency for the exploratory goals of the survey. revision: partial

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: survey of external literature with independent coding

full rationale

The paper performs a literature survey of over 200 external IEEE VIS papers (2013-2022), develops a four-axis coding scheme, and qualitatively identifies pathways. No mathematical derivations, parameter fitting, predictions, or self-referential equations exist. The central claim rests on observed patterns in the collected corpus rather than any reduction to the paper's own inputs or prior self-citations. The analysis is self-contained against external benchmarks (published papers) with no load-bearing self-citation chains or ansatzes imported from the authors' own work.

Assumptions & free parameters 0 free parameters · 1 assumptions · 0 invented entities

Survey paper; no free parameters or invented entities. Relies on domain assumption that the selected conference corpus represents the field.

assumptions (1)
  • domain assumption The corpus of IEEE VIS papers from the past decade adequately represents VIS4ML research on human knowledge injection.
    Paper states it collected papers from IEEE VIS conferences in the past decade as the basis for the survey.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Understanding How Humans Inject Knowledge into Machine Learning Workflows through Visual Analytics." pith.science (2026). https://pith.science/paper/URBLI34C

@misc{pith2026260700969,
  author       = {Pith},
  title        = {Pith review of: Understanding How Humans Inject Knowledge into Machine Learning Workflows through Visual Analytics},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/URBLI34C}},
  note         = {Machine review of arXiv:2607.00969}
}
read the original abstract

Visual analytics (VA) plays an increasingly important role in supporting machine learning (ML) workflows. In the field of visualization, such approaches and techniques are referred to as VIS4ML. While ML models are mostly learned automatically, the corresponding ML workflows receive a variety of human inputs, such as data labelling, feature engineering, model architecture designing, hyper-parameter tuning, and so on. In this work, we surveyed over 200 VIS4ML papers to gain an understanding of how humans inject their knowledge into ML workflows through interactive visualization. We collected a corpus of VIS4ML papers from the IEEE VIS conferences in the past decade. We developed a coding scheme to facilitate the literature research from four perspectives: characteristics of ML, visualization, interaction, and actions. The analysis of the coded dataset allows us to observe different pathways that transfer human knowledge to ML workflows via interactive visualization. Building on the analysis, we explain the phenomena of VIS4ML using the conceptual model that views VA as model building and the information-theoretic cost-benefit analysis that reasons VA as for optimizing ML workflows. This work provides unequivocal evidence showing the merits of using VA in ML workflows. The full list of surveyed papers, along with all analysis results and figures, is available at https://vis4ml4hd.github.io/ml-knowledge-inject-va/.

Figures

Figures reproduced from arXiv: 2607.00969 by the authors.

Figure 1
Figure 1. Overview of the survey process: (a) Paper screening; (b) Coding [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Selected pairwise Sankey diagrams (part 1 of 2). [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Selected pairwise Sankey diagrams (part 2 of 2). [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Word clouds of important terms in four selected topics support [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: The survey allows us to answer questions about knowledge may be injected into ML workflows. (a) possible injection pathways, and (b)-(f) [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

54 extracted references · 54 canonical work pages

  1. [1]

    Amershi, A

    S. Amershi, A. Begel, C. Bird, R. DeLine, H. Gall, E. Kamar et al. Soft- ware engineering for machine learning: A case study. In 2019 IEEE/ACM 41st International Conference on Software Engineering: Software Engi- neering in Practice (ICSE-SEIP) , pp. 291–300, 2019. 2, 18

  2. [2]

    Andrienko, N

    G. Andrienko, N. Andrienko, and D. Hecker. Topic modelling for spa- tial insights: Uncovering space use from movement data. Computers & Graphics, 122:103989, 2024. 6

  3. [3]

    Andrienko, T

    N. Andrienko, T. Lammarsch, G. Andrienko, G. Fuchs, D. Keim, S. Miksch et al. Viewing visual analytics as model building. Computer Graphics F orum, 37(6):275–299, 2018. 2, 8, 9

  4. [4]

    C. M. Bishop and N. M. Nasrabadi. Pattern recognition and machine learning, vol. 4. 2006. 3

  5. [5]

    Chai and G

    C. Chai and G. Li. Human-in-the-loop techniques in machine learning. IEEE Data Eng. Bull. , 43(3):37–52, 2020. 2

  6. [6]

    Chang, C

    R. Chang, C. Ziemkiewicz, T. M. Green, and W. Ribarsky. Defining insight for visual analytics. IEEE Computer Graphics and Applications , 29(2):14–17, 2009. 2

  7. [7]

    Chatzimparmpas, R

    A. Chatzimparmpas, R. M. Martins, I. Jusufi, K. Kucher, F. Rossi, and A. Kerren. The state of the art in enhancing trust in machine learn- ing models with the use of visualizations. Computer Graphics F orum, 39:713–756, 6 2020. 2

  8. [8]

    Chatzimparmpas, R

    A. Chatzimparmpas, R. M. Martins, and A. Kerren. Visruler: Visual analytics for extracting decision rules from bagged and boosted decision trees. Information Visualization, 22(2):115–139, 2023. 2

Show all 54 references
  1. [9]

    M. Chen. Cost-benefit analysis of data intelligence – its broader inter- pretations. In Advances in Info-Metrics: Information and Information Processing across Disciplines, pp. 433–463. 2020. 2, 9

  2. [10]

    Chen and D

    M. Chen and D. J. Edwards. ‘isms’ in visualization. In F oundations of Data Visualization, pp. 225–241. 2020. 2

  3. [11]

    M. Chen, L. Floridi, and R. Borgo. What is visualization really for? In The Philosophy of Information Quality , vol. 358, pp. 7593. 2014. 2

  4. [12]

    Chen and A

    M. Chen and A. Golan. What may visualization processes opti- mize? IEEE Transactions on Visualization and Computer Graphics , 22(12):2619–2632, 2016. 2, 9, 10

  5. [13]

    M. Chen, K. Mueller, and A. Ynnerman. Fusion of Visual Channels , pp. 119–127. Springer London, London, 2014. 3

  6. [14]

    Diehl, A

    A. Diehl, A. Abdul-Rahman, B. Bach, M. El-Assady, M. Kraus, R. S. Laramee et al. An analysis of the interplay and mutual benefits of grounded theory and visualization. IEEE Transactions on Visualization and Computer Graphics, 31(9):5462–5479, 2025. 3

  7. [15]

    J. J. Dudley and P . O. Kristensson. A review of user interface design for interactive machine learning. ACM Transactions on Interactive Intelligent Systems (TiiS), 8(2):1–37, 2018. 2

  8. [16]

    Endert, W

    A. Endert, W. Ribarsky, C. Turkay, B. W. Wong, I. Nabney, I. D. Blanco et al. The state of the art in integrating machine learning into visual analytics. Computer Graphics F orum, 36:458–486, 12 2017. 2

  9. [17]

    S. D. H. Evergreen. Effective Data Visualization: The Right Chart for the Right Data. SAGE Publications, 2016. 2

  10. [18]

    Fernández-Delgado, E

    M. Fernández-Delgado, E. Cernadas, S. Barro, and D. Amorim. Do we need hundreds of classifiers to solve real world classification problems? The journal of machine learning research , 15(1):3133–3181, 2014. 3

  11. [19]

    Gómez-Carmona, D

    O. Gómez-Carmona, D. Casado-Mansilla, D. Lopez-de Ipina, and J. García-Zubia. Human-in-the-loop machine learning: Reconceptual- izing the role of the user in interactive approaches. Internet of Things , 25:101048, 2024. 2

  12. [20]

    Goodfellow, Y

    I. Goodfellow, Y . Bengio, A. Courville, and Y . Bengio. Deep learning , vol. 1. 2016. 3

  13. [21]

    Hohman, M

    F. Hohman, M. Kahng, R. Pienta, and D. H. Chau. Visual analytics in deep learning: An interrogative survey for the next frontiers. IEEE Trans- actions on Visualization and Computer Graphics , 25:2674–2693, 8 2019. 2

  14. [22]

    Holzinger

    A. Holzinger. Interactive machine learning for health informatics: when do we need the human-in-the-loop? Brain informatics , 3(2):119–131,

  15. [23]

    S. B. Kotsiantis, I. Zaharakis, P . Pintelas, et al. Supervised machine learn- ing: A review of classification techniques. Emerging artificial intelli- gence applications in computer engineering , 160(1):3–24, 2007. 3

  16. [24]

    Kumar, S

    S. Kumar, S. Datta, V . Singh, D. Datta, S. Kumar Singh, and R. Sharma. Applications, challenges, and future directions of human-in-the-loop learning. IEEE Access, 12:75735–75760, 2024. 2

  17. [25]

    S. Liu, J. Xiao, J. Liu, X. Wang, J. Wu, and J. Zhu. Visual diagnosis of tree boosting methods. IEEE Trans. Vis. Comp. Graph. , 24(1):163–173,

  18. [26]

    R. Marty. Applied Security Visualization. Addison-Wesley, 2009. 2

  19. [27]

    Moussavi, G

    L. Moussavi, G. Andrienko, N. Andrienko, and A. Slingsby. Visually- supported topic modeling for understanding behavioral patterns from spa- tiotemporal events. Computers & Graphics, 129:104245, 2025. 6

  20. [28]

    O’Callaghan, D

    D. O’Callaghan, D. Greene, J. Carthy, and P . Cunningham. An analysis of the coherence of descriptors in topic modeling. Expert Syst. Appl. , 42(13):56455657, Aug. 2015. 5

  21. [29]

    H. Park, N. Das, R. Duggal, A. P . Wright, O. Shaikh, F. Hohman et al. Neurocartography: Scalable automatic visual summarization of concepts in deep neural networks. IEEE Transactions on Visualization and Com- puter Graphics, 28(1):813823, Jan. 2022. 7

  22. [30]

    Sacha, M

    D. Sacha, M. Kraus, J. Bernard, M. Behrisch, T. Schreck, Y . Asano et al. Somflow: Guided exploratory cluster analysis with self-organizing maps and analytic provenance. IEEE Transactions on Visualization and Computer Graphics, 24(1):120–130, 2018. 7

  23. [31]

    Sacha, M

    D. Sacha, M. Kraus, D. A. Keim, and M. Chen. Vis4ml: An ontology for visual analytics assisted machine learning. IEEE Transactions on Visual- ization and Computer Graphics , 25(1):385–395, 2019. 1, 2, 3

  24. [32]

    Sculley, G

    D. Sculley, G. Holt, D. Golovin, E. Davydov, T. Phillips, D. Ebner et al. Hidden technical debt in machine learning systems. Advances in neural information processing systems, 28, 2015. 2, 18

  25. [33]

    Sevastjanova, S

    R. Sevastjanova, S. V ogelbacher, A. Spitz, D. Keim, and M. El-Assady. Visual comparison of text sequences generated by large language models. In IEEE Visualization in Data Science (VDS) , pp. 11–20, 2023. 2

  26. [34]

    Q. Shen, Y . Wu, Y . Jiang, W. Zeng, K. Alexis, A. Vianova et al. Visual in- terpretation of recurrent neural network on multi-dimensional time-series forecast. In 2020 IEEE Pacific visualization symposium (PacificVis) , pp. 61–70, 2020. 2

  27. [35]

    Sperrle, M

    F. Sperrle, M. ElÄêAssady, G. Guo, R. Borgo, D. H. Chau, A. Endert et al. A Survey of Human-Centered Evaluations in Human-Centered Machine Learning. Computer Graphics F orum, 40:543–568, 6 2021. 2

  28. [36]

    J. T. Stasko. V alue-driven evaluation of visualizations. In Proc. 5th Work- shop on Beyond Time and Errors: Novel Evaluation Methods for Visual- ization (BELIV). 2014. 2

  29. [37]

    Streeb, M

    D. Streeb, M. El-Assady, D. A. Keim, and M. Chen. Visualize? argu- ments for visual support in decision making. IEEE Computer Graphics and Applications, 41(2):17–22, 2021. 2

  30. [38]

    Streeb, M

    D. Streeb, M. El-Assady, D. A. Keim, and M. Chen. Why visualize? untangling a large network of arguments. IEEE Transactions on Visual- ization and Computer Graphics , 27(3):2220–2236, 2021. 2

  31. [39]

    Strobelt, A

    H. Strobelt, A. Webson, V . Sanh, B. Hoover, J. Beyer, H. Pfister et al. In- teractive and visual prompt engineering for ad-hoc task adaptation with large language models. IEEE Transactions on Visualization and Com- puter Graphics, pp. 1–11, 2022. 2

  32. [40]

    A. Suh, A. Mosca, E. Wu, and R. Chang. A grammar of hypotheses for visualization, data, and analysis, 2023. 2

  33. [41]

    G. K. Tam, V . Kothari, and M. Chen. An analysis of machine-and human- analytics in classification. IEEE Trans. Vis. Comp. Graph. , 23:71–80,

  34. [42]

    E. R. Tufte. The Visual Display of Quantitative Information . 2nd ed.,

  35. [43]

    J. J. van Wijk. The value of visualization. In Proc. IEEE Visualization, pp. 79–86. 2005. 2

  36. [44]

    J. Wang, S. Hazarika, C. Li, and H.-W. Shen. Visualization and visual analysis of ensemble data: A survey. IEEE Transactions on Visualization and Computer Graphics, 25(9):2853–2872, 2019. 1

  37. [45]

    J. Wang, S. Liu, and W. Zhang. Visual analytics for machine learning: A data perspective survey. IEEE Transactions on Visualization and Com- puter Graphics, 30(12):7637–7656, 2024. 1, 2

  38. [46]

    J. Wang, W. Zhang, H. Y ang, C.-C. M. Y eh, and L. Wang. Visual ana- lytics for rnn-based deep reinforcement learning. IEEE Trans. Vis. Comp. Graph., 28(12):4141–4155, 2021. 2

  39. [47]

    Q. Wang, W. Alexander, J. Pegg, H. Qu, and M. Chen. Hypoml: Vi- sual analysis for hypothesis-based evaluation of machine learning models. IEEE Trans. Vis. Comp. Graph., 27(2):1417–1426, 2020. 2

  40. [48]

    C. Ware. Information Visualization: Perception for Design. 2nd ed., 2004. 2

  41. [49]

    Wexler, M

    J. Wexler, M. Pushkarna, T. Bolukbasi, M. Wattenberg, F. Viegas, and J. Wilson. The what-if tool: Interactive probing of machine learning mod- els. IEEE Transactions on Visualization and Computer Graphics, pp. 1–1,

  42. [50]

    X. Wu, L. Xiao, Y . Sun, J. Zhang, T. Ma, and L. He. A survey of human- in-the-loop for machine learning. Future Generation Computer Systems , 135:364–381, 2022. 2

  43. [51]

    D. Xin, L. Ma, J. Liu, S. Macke, S. Song, and A. Parameswaran. Acceler- ating human-in-the-loop machine learning: Challenges and opportunities. In Proceedings of the second workshop on data management for end-to- end machine learning, pp. 1–4, 2018. 2

  44. [52]

    Y uan, C

    J. Y uan, C. Chen, W. Y ang, M. Liu, J. Xia, and S. Liu. A survey of visual analytics techniques for machine learning. Computational Visual Media, 7:3–36, 3 2021. 1, 2 APPENDICES Understanding How Humans Inject Knowledge into Machine Learning Workflows through Visual Analytics ...

  45. [53]

    Logical/Symbolic Models; 3

    Linear Models; 2. Logical/Symbolic Models; 3. Statistical/Graphical Models; 4. Instance-Based Models; 5. Kernel Methods; 6. Neural Models; 7. Sequence Models; 8. Decision Tree-Based Models

  46. [54]

    steering

    Autoencoder-Based Models; 10. Perceptron-Based Models; 11. Ensemble Methods; 12. Evolutionary Methods; 13. Others *1.5 ML_TASK: Target problems addressed by the model. 1.Classification; 2.Detection/Recognition; 3.Regression; 4.Clustering; 5.Association; 6.Feature Selec- tion; 7...

Pith tools

Reviewed July 2, 2026 · model on record in the stance chip above.