REVIEW 3 major objections 5 minor 46 references
Visualization in the preprocessing phase: an interview study with enterprise professionals
T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read Ten insights define how visualization should support data preprocessing
desk verdict Useful consolidation of preprocessing-focused visualization needs, with an honest small-sample caveat; worth refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the ten-insight list itself, each insight being a single recommendation for what visualization should offer during preprocessing, built through an iterative, incremental coding method. The list began from the first participant's wishlist; every later interview updated it, and after all interviews the combined entries were reviewed, labelled, ordered by frequency, and filtered so that items mentioned by fewer than two participants were dropped. A comparison with three prior interview studies merged their design implications with the list, adding the tenth insight. The list does the argument's work: it converts a heterogeneous set of interview responses into a compact, ordered set of requirements that future tool designs can be checked against.
What would settle it
Run the same semi-structured protocol with a larger and more diverse sample, such as fifty analysts across government, healthcare, retail, and finance in more than one country; if most of the ten insights no longer reach the two-participant mention threshold, or several new high-frequency insights appear, the claim that this list is a consolidated requirements baseline for enterprise preprocessing would not hold.
Extended reading notes
Core claim
The central claim is a consolidated list of ten insights stating how visualization should support preprocessing activities from the data analyst's perspective. The insights are: keep visualization simple; keep context by remaining compatible with the programming environments analysts already use; save the analyst's time; support large-scale or Big Data scenarios; allow interaction beyond static reports; recognize tables as a legitimate visual format; pay attention to under-served work scopes such as feature creation and deep-learning interpretation; treat preprocessing as part of a back-and-forth cycle rather than a linear first phase; allow comparison of data before and after transformations; and capture the metadata or logic behind automatic transformations. The authors report that the first nine emerged from their interviews, the tenth came from the related work, and three of the insights were not present in any of the three comparison studies. Readers are asked to accept this list as a requirements baseline for new visualization efforts in visual data exploration.
Load-bearing premise
The study assumes that thirteen self-selected data analysts, mostly male and mostly working in one country's IT industry, can reveal enough about enterprise preprocessing practice that the ten insights generalize beyond that group.
Editorial extensions
If this is right
- Visualization tools for data mining should run inside Python and R workflows instead of forcing analysts to switch environments and lose context.
- Scalability is a precondition: analysts working with very large datasets need aggregation and density-based rendering before they can benefit from novel chart types.
- Table views are a legitimate visualization; design effort should go into enhancing tables with interaction and pixel-oriented techniques rather than replacing them.
- Preprocessing tooling should let analysts compare data before and after transformations and should record the logic behind automatic cleaning steps.
- The ten insights can serve as a checklist for planning or evaluating new visualization solutions aimed at initial exploratory analysis.
Reading between the lines
- The authors do not spell it out, but the ten insights also function as an evaluation rubric: a tool that violates several of them, for example by hiding transformation logic or requiring context switches, should predictably struggle to gain adoption, and that prediction could be tested in a comparative study.
- Insight 8, which reframes preprocessing as a back-and-forth cycle, points toward a concrete design direction: show partial cleaning results as they become available and let the analyst keep interacting while the tool works.
- Because the sample is mostly male, mostly IT-industry, and drawn from one country, the frequency ordering of the insights is the most fragile part; a testable extension is to rerun the same protocol with analysts in government, healthcare, and retail and compare which insights reach the mention threshold.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript reports a semi-structured interview study with thirteen enterprise data analysts, focusing on how visualization can support the preprocessing phase of data mining workflows. The authors describe the participants' profiles, their data analysis processes, prevalent preprocessing activities, data quality concerns, and their current use of visualization. From an iterative, incremental coding of the interview responses, the authors derive a consolidated list of ten insights for visualization tools, compare this list with three related works (RW1, RW2, RW3), and position the list as a requirements baseline for future visualization research. The paper concludes with acknowledged limitations, including the sample's narrow demographic and geographic representativeness.
Significance. If the ten-insight list is robust, it would provide a concise, actionable checklist for visualization researchers and tool builders targeting the preprocessing phase, an area that the authors argue is underserved relative to later-stage visualization. The study's strengths include its transparent description of the coding procedure, the explicit triangulation with three related interview studies, and an honest section on limitations. The paper also offers useful descriptive detail about practitioner workflows and data quality issues. However, the central contribution rests on a small, self-selected, single-country sample, and the manuscript's framing as a 'consolidated set of requirements' exceeds what the evidence can support. The insights are best regarded as candidate requirements or hypotheses for further validation, not as a settled baseline.
major comments (3)
- [Section III-C and Section IV, Step 4] There is an inconsistency in the inclusion threshold that determines which items become insights. Section III-C states 'we considered the items reported by more than two participants,' which means at least three participants. In contrast, Section IV, Step 4 states 'the items that were not mentioned by at least two participants were not included,' which means at least two participants. With n=13, the difference between 2 and 3 mentions is material (15% versus 23% of the sample) and directly affects the composition of the central ten-insight list. The authors should reconcile these statements, report the exact number of participants supporting each insight, and justify the chosen threshold.
- [Abstract, Section I, and Section V] The abstract and introduction describe the ten insights as 'a consolidated set of requirements' for future visualization research, but Section V explicitly concedes that 'the data collected and its analysis cannot be considered a representation of all data analysts.' Given the sample of thirteen participants, mostly male and mostly working in the IT industry in one country, the requirements framing overstates the evidence. I recommend reframing the contribution as candidate insights or requirements hypotheses that need further validation, and adjusting the abstract and Section I accordingly so that the claims match the scope of the study.
- [Section IV and Section V] The consolidated list treats all ten insights uniformly, but their provenance differs: insights 7-9 are unique to this small interview sample, while insight 10 is taken solely from related works, as the authors acknowledge. This mixing is not itself an error, but the final presentation should make the provenance of each insight explicit in the main text (not only in Figure 4) and should discuss the evidentiary status of each category. In particular, the three insights that are not corroborated by any related work should be flagged as tentative, and the manuscript should suggest what additional evidence would be needed to solidify them.
minor comments (5)
- [Section III-B] The sentence 'The same environment configuration was used for all participants, face-to-face or online conversations' is ambiguous; clarify whether the same physical setup or the same set of instructions was used across the two modes of interviewing.
- [Figure 5 caption] There is a typo in the caption: 'insigths' should be 'insights.'
- [Section III-A] The manuscript states that participants were located in three cities in the same country but does not name the country or cities. Naming the country (or at least the region) would help readers calibrate the generalizability of the findings, especially because the study's own Section V emphasizes the sample's limitations.
- [Section III-B] The paper notes that 'parts of the sessions were recorded' but does not specify which parts or how the audio was used beyond reviewing notes. A brief clarification of the recording and transcription procedure would improve methodological transparency.
- [Section IV, Step 1] The coding description says the list 'started based on the inputs received from participant one' and was then iteratively revised. This is a reasonable approach, but the paper should state whether any of the final insights were added or removed during this iterative process and how disagreements or ambiguous responses were handled.
Circularity Check
No circularity: the ten insights are an empirical synthesis of interview responses and related-work triangulation, with each item's provenance explicitly disclosed.
full rationale
The paper's central claim is a qualitative research finding: a consolidated list of ten insights about how visualization can support preprocessing, derived from semi-structured interviews with thirteen data analysts and then compared with three related interview studies. The derivation chain is transparent and empirical: participant responses were coded iteratively (Section IV), items mentioned by fewer than two participants were excluded, and the final list was ordered by frequency, with insights unique to this study and one insight imported solely from related work clearly labeled. No parameter is fitted to data and then renamed as a prediction; there are no equations, no formal model, and no target quantity being reproduced from its own inputs. The authors do not rely on self-citations for any load-bearing premise, and the related works are used only for triangulation, with the paper explicitly stating which insights came from interviews and which came from the literature. The acknowledged limitation in Section V, that the convenience sample 'cannot be considered a representation of all data analysts,' is a generalizability concern rather than evidence of circularity, because it does not make the resulting list equivalent to its inputs by construction. The paper is therefore self-contained as an interview study and shows no significant circularity.
Assumptions & free parameters
assumptions (4)
- domain assumption Participant self-reports are accurate about their data analysis practices.
- ad hoc to paper The threshold of at least two participants for an insight to be included is a valid way to identify recurring needs.
- domain assumption The three related works selected (RW1, RW2, RW3) are representative of prior interview studies in this space.
- domain assumption Qualitative thematic coding without inter-rater reliability is a sound basis for the insights.
Cite this review
Pith. "Pith review of Visualization in the preprocessing phase: an interview study with enterprise professionals." pith.science (2026). https://pith.science/paper/47YOLHLQ
@misc{pith2026190807894,
author = {Pith},
title = {Pith review of: Visualization in the preprocessing phase: an interview study with enterprise professionals},
year = {2026},
howpublished = {\url{https://pith.science/paper/47YOLHLQ}},
note = {Machine review of arXiv:1908.07894}
}
read the original abstract
The current information age has increasingly required organizations to become data-driven. However, analyzing and managing raw data is still a challenging part of the data mining process. Even though we can find interview studies proposing design implications or recommendations for future visualization solutions in the data mining scope, they cover the entire workflow and do not fully focus on the challenges during the preprocessing phase and on how visualization can support it. Moreover, they do not organize a final list of insights consolidating the findings of other related studies. Hence, to better understand the current practice of enterprise professionals in data mining workflows, in particular during the preprocessing phase, and how visualization supports this process, we conducted semi-structured interviews with thirteen data analysts. The discussion about the challenges and opportunities based on the responses of the interviewees resulted in a list of ten insights. This list was compared with the closest related works, improving the reliability of our findings and providing background, as a consolidated set of requirements, for future visualization research papers applied to visual data exploration in data mining. Furthermore, we provide greater details on the profile of the data analysts, the main challenges they face, and the opportunities that arise while they are engaged in data mining projects in diverse organizational areas.
Figures
Reference graph
Works this paper leans on
-
[1]
Exploratory Data Mining and Data Cleaning
Dasu T and Johnson T. Exploratory Data Mining and Data Cleaning . 1 ed. New Y ork, NY , USA: John Wiley & Sons, Inc., 2003. ISBN 0471268518
work page 2003
-
[2]
Data Mining: Concepts and Techniques
Han J, Kamber M and Pei J. Data Mining: Concepts and Techniques . 3rd ed. San Francisco, CA, USA: Morgan Kaufmann Publishers I nc.,
-
[3]
Knowledge Discovery in Databases
Piateski G and Frawley W. Knowledge Discovery in Databases . Cambridge, MA, USA: MIT Press, 1991. ISBN 0262660709
work page 1991
-
[4]
The crisp-dm model: the new blueprint for data mining
Shearer C. The crisp-dm model: the new blueprint for data mining. Journal of data warehousing 2000; 5(4): 13–22
work page 2000
-
[5]
Wrangler: Inter active visual specification of data transformation scripts
Kandel S, Paepcke A, Hellerstein J et al. Wrangler: Inter active visual specification of data transformation scripts. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems . CHI ’11, New Y ork, NY , USA: ACM. ISBN 978-1-4503-0228-9, pp. 3363–33 72. DOI:10.1145/1978942.1979444
-
[6]
Quantitative data cleaning for large da tabases, 2008
Hellerstein JM. Quantitative data cleaning for large da tabases, 2008
work page 2008
-
[7]
Jugulum R. Competing with High Quality Data: Concepts, Tools, and Techniques for Building a Successful Approach to Data Quali ty. Wiley,
-
[8]
Introduction to Data Mining, (First Edition)
Tan PN, Steinbach M and Kumar V . Introduction to Data Mining, (First Edition). Boston, MA, USA: Addison-Wesley Longman Publishing Co., Inc., 2005. ISBN 0321321367
work page 2005
Show all 46 references
-
[9]
From visual data exploration to visual data mining: A survey
Ferreira de Oliveira MC and Levkowitz H. From visual data exploration to visual data mining: A survey. IEEE Transactions on Visualization and Computer Graphics 2003; 9(3): 378–394. DOI:10.1109/TVCG.2003. 1207445
2003 doi
-
[10]
Interactive Data Visualization: F oundations, Techniques, and Applications, Second Edition - 360 Degree Business
Ward MO, Grinstein G and Keim D. Interactive Data Visualization: F oundations, Techniques, and Applications, Second Edition - 360 Degree Business. 2nd ed. Natick, MA, USA: A. K. Peters, Ltd., 2015. ISBN 1482257378, 9781482257373
2015
-
[11]
The interactive visualization g ap in initial exploratory data analysis
Batch A and Elmqvist N. The interactive visualization g ap in initial exploratory data analysis. IEEE Transactions on Visualization and Computer Graphics 2018; 24(1): 278–287. DOI:10.1109/TVCG.2017. 2743990
2018 doi
-
[12]
Enterprise da ta analysis and visualization: An interview study
Kandel S, Paepcke A, Hellerstein JM et al. Enterprise da ta analysis and visualization: An interview study. IEEE Transactions on Visualization and Computer Graphics 2012; 18(12): 2917–2926. DOI:10.1109/TVCG. 2012.219
2012 doi
-
[13]
Futzing and moseying: I nter- views with professional data analysts on exploration pract ices
Alspaugh S, Zokaei N, Liu A et al. Futzing and moseying: I nter- views with professional data analysts on exploration pract ices. IEEE Transactions on Visualization and Computer Graphics 2018; : 1–1DOI: 10.1109/TVCG.2018.2865040
2018
-
[14]
Guidelines for snowballing in systematic lit erature studies and a replication in software engineering
Wohlin C. Guidelines for snowballing in systematic lit erature studies and a replication in software engineering. In Proceedings of the 18th International Conference on Evaluation and Assessment in S oftware Engineering. EASE ’14, New Y ork, NY , USA: ACM. ISBN 978-1- 4503-247...
-
[15]
Empirical studies in i nformation visualization: Seven scenarios
Lam H, Bertini E, Isenberg P et al. Empirical studies in i nformation visualization: Seven scenarios. IEEE Transactions on Visualization and Computer Graphics 2012; 18(9): 1520–1536. DOI:10.1109/TVCG.2011. 279
2012 doi
-
[16]
Research Methods in Human- Computer Interaction, 2nd Edition
Lazar J, Feng JH and Hochheiser H. Research Methods in Human- Computer Interaction, 2nd Edition . Elsevier Science, 2017. ISBN 9780128093436
2017
-
[17]
What is big data? a cons ensual definition and a review of key research topics
De Mauro A, Greco M and Grimaldi M. What is big data? a cons ensual definition and a review of key research topics. In AIP conference proceedings, volume 1644. AIP , pp. 97–104
-
[18]
Python. Python. https://www.python.org/, 2018. [Online; accessed 12- December-2018]
2018
-
[19]
The r project for statistical computing
R. The r project for statistical computing. https://www.r-project.org/,
-
[20]
Databricks: Making big data simple
Databricks. Databricks: Making big data simple. https://databricks.com/,
-
[21]
Knime - the konstanz i nformation miner: V ersion 2.0 and beyond
Berthold MR, Cebron N, Dill F et al. Knime - the konstanz i nformation miner: V ersion 2.0 and beyond. SIGKDD Explor Newsl 2009; 11(1): 26–31. DOI:10.1145/1656274.1656280
2009
-
[22]
Knime: Open for innovation
KNIME. Knime: Open for innovation. https://www.knime.com/, 2018. [Online; accessed 12-December-2018]
2018
-
[24]
The open graph viz platform
Gephi. The open graph viz platform. https://gephi.org/, 2018. [Online; accessed 12-December-2018]
2018
-
[25]
Orange: Data mining to olbox in python
Demˇ sar J, Curk T, Erjavec A et al. Orange: Data mining to olbox in python. Journal of Machine Learning Research 2013; 14: 2349–2353
2013
-
[26]
Gephi: an open source s oftware for exploring and manipulating networks
Bastian M, Heymann S and Jacomy M. Gephi: an open source s oftware for exploring and manipulating networks. In Third international AAAI conference on weblogs and social media . pp. 361–362
-
[27]
Spark: Cluste r computing with working sets
Zaharia M, Chowdhury M, Franklin MJ et al. Spark: Cluste r computing with working sets. In Proceedings of the 2Nd USENIX Conference on Hot Topics in Cloud Computing . HotCloud’10, Berkeley, CA, USA: USENIX Association, pp. 10–10
-
[28]
Apache spark: Unied analytics engine for big da ta
Spark A. Apache spark: Unied analytics engine for big da ta. https://spark.apache.org/, 2018. [Online; accessed 12-December-2018]
2018
-
[29]
Orange: Data mining fruitful and fun
Orange. Orange: Data mining fruitful and fun. https://orange.biolab.si/,
-
[30]
[Online; accessed 12-December-2018]
2018
-
[31]
ggplot2 - tidyverse
ggplot2. ggplot2 - tidyverse. https://ggplot2.tidyverse.org/, 2018. [Online; accessed 12-December-2018]
2018
-
[32]
Apache hadoop
Hadoop A. Apache hadoop. https://hadoop.apache.org/, 2018. [Online; accessed 12-December-2018]
2018
-
[33]
Matplotlib: Python plotting - matplotlib 3.0.2 documentation
Matplotlib. Matplotlib: Python plotting - matplotlib 3.0.2 documentation. https://matplotlib.org/, 2018. [Online; accessed 12-December-2018]
2018
-
[34]
seaborn: statistical data visualization - 0
Seaborn. seaborn: statistical data visualization - 0. 9.0 documenta- tion. https://seaborn.pydata.org/, 2018. [Online; accessed 12-December- 2018]
2018
-
[35]
Vim: Visualization and imputation of missing values
Templ M, Alfons A, Kowarik A et al. Vim: Visualization and imputation of missing values. 11 https://cran.r-project.org/web/packages/VIM/index.html, 2018. [Online; accessed 12-December-2018]
2018
-
[36]
Tableau. Tableau. http://www.tableau.com/, 2018. [Online; accessed 12-December-2018]
2018
-
[37]
Sas analytics
SAS. Sas analytics. https://www.sas.com/, 2018. [Online; accessed 12-December-2018]
2018
-
[38]
Imputation with the R package VIM
Kowarik A and Templ M. Imputation with the R package VIM. Journal of Statistical Software 2016; 74(7): 1–16. DOI:10.18637/jss.v074.i07
2016 doi
-
[39]
The eyes have it: A task by data type taxon omy for information visualizations
Shneiderman B. The eyes have it: A task by data type taxon omy for information visualizations. In Proceedings of the 1996 IEEE Symposium on Visual Languages . VL ’96, Washington, DC, USA: IEEE Computer Society. ISBN 0-8186-7508-X, pp. 336–
1996
-
[40]
The table lens: Merging graphical and s ymbolic representations in an interactive focus + context visualiz ation for tabular information
Rao R and Card SK. The table lens: Merging graphical and s ymbolic representations in an interactive focus + context visualiz ation for tabular information. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems . CHI ’94, New Y ork, NY , USA: ACM. ISBN ...
-
[41]
Qlik: Data analytics for modern business intelli gence
Qlik. Qlik: Data analytics for modern business intelli gence. https://www.qlik.com, 2018. [Online; accessed 12-December-2018]
2018
-
[42]
Seeing, sensing, and scrutinizing
Rensink RA. Seeing, sensing, and scrutinizing. Vision Research 2000; 40(10): 1469 – 1487. DOI:https://doi.org/10.1016/S0042- 6989(00) 00003-1
2000 doi
-
[43]
Progressive visual analy tics: User- driven visual exploration of in-progress analytics
Stolper CD, Perer A and Gotz D. Progressive visual analy tics: User- driven visual exploration of in-progress analytics. IEEE Transactions on Visualization and Computer Graphics 2014; 20(12): 1653–1662. DOI: 10.1109/TVCG.2014.2346574
2014
-
[44]
List_Insights_horizontal.JPG
Kindlmann G and Scheidegger C. An algebraic process for visualization design. IEEE Transactions on Visualization and Computer Graphics 2014; 20(12): 2181–2190. DOI:10.1109/TVCG.2014.2346325 . This figure "List_Insights_horizontal.JPG" is available in "JPG" format from: http://...
2014
-
[45]
Designing pixel-oriented visualization techn iques: theory and applications
Keim D. Designing pixel-oriented visualization techn iques: theory and applications. IEEE Transactions on Visualization and Computer Graphics 2000; 6(1): 59–78. DOI:10.1109/2945.841121
-
[46]
Mastering the Information Age: Solving Problems with Visual Analytics, Eurographics Association
Keim D, Kohlhammer J and Ellis G. Mastering the Information Age: Solving Problems with Visual Analytics, Eurographics Association. Germany: Eurographics Association, 2010. ISBN 978-3-9056 73-77-7
2010
-
[2011]
ISBN 0123814790, 9780123814791
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.