Pith. sign in

REVIEW 3 major objections 5 minor 46 references

Visualization in the preprocessing phase: an interview study with enterprise professionals

T0 review · 3 major / 5 minor · reviewed 2026-08-14 · deepseek-v4-flash

Pith's one-line read Ten insights define how visualization should support data preprocessing

desk verdict Useful consolidation of preprocessing-focused visualization needs, with an honest small-sample caveat; worth refereeing. read the letter →

arxiv 1908.07894 v1 pith:47YOLHLQ submitted 2019-08-21 cs.HC

classification cs.HC
keywords visualizationpreprocessingvisualdataexplorationmininginterviewstudyenterpriseprofessionalsqualitydesigninsights
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that visualization can support the preprocessing phase of enterprise data mining in ten specific, reusable ways. Drawing on semi-structured interviews with thirteen data analysts, the authors consolidate recurring challenges and wishes into a list of ten insights, including 'keep it simple,' 'tables are OK,' 'allow comparison,' and 'capture metadata.' They then compare this list against three earlier interview studies, which ratifies most items and contributes one insight the interviews had not raised. If the list holds, it gives future visualization tool builders a straightforward requirements baseline for helping analysts during data cleaning, standardization, and feature selection rather than only when presenting final results.

What carries the argument

The central object is the ten-insight list itself, each insight being a single recommendation for what visualization should offer during preprocessing, built through an iterative, incremental coding method. The list began from the first participant's wishlist; every later interview updated it, and after all interviews the combined entries were reviewed, labelled, ordered by frequency, and filtered so that items mentioned by fewer than two participants were dropped. A comparison with three prior interview studies merged their design implications with the list, adding the tenth insight. The list does the argument's work: it converts a heterogeneous set of interview responses into a compact, ordered set of requirements that future tool designs can be checked against.

What would settle it

Run the same semi-structured protocol with a larger and more diverse sample, such as fifty analysts across government, healthcare, retail, and finance in more than one country; if most of the ten insights no longer reach the two-participant mention threshold, or several new high-frequency insights appear, the claim that this list is a consolidated requirements baseline for enterprise preprocessing would not hold.

Watch

Extended reading notes

Core claim

The central claim is a consolidated list of ten insights stating how visualization should support preprocessing activities from the data analyst's perspective. The insights are: keep visualization simple; keep context by remaining compatible with the programming environments analysts already use; save the analyst's time; support large-scale or Big Data scenarios; allow interaction beyond static reports; recognize tables as a legitimate visual format; pay attention to under-served work scopes such as feature creation and deep-learning interpretation; treat preprocessing as part of a back-and-forth cycle rather than a linear first phase; allow comparison of data before and after transformations; and capture the metadata or logic behind automatic transformations. The authors report that the first nine emerged from their interviews, the tenth came from the related work, and three of the insights were not present in any of the three comparison studies. Readers are asked to accept this list as a requirements baseline for new visualization efforts in visual data exploration.

Load-bearing premise

The study assumes that thirteen self-selected data analysts, mostly male and mostly working in one country's IT industry, can reveal enough about enterprise preprocessing practice that the ten insights generalize beyond that group.

Editorial extensions

If this is right

  • Visualization tools for data mining should run inside Python and R workflows instead of forcing analysts to switch environments and lose context.
  • Scalability is a precondition: analysts working with very large datasets need aggregation and density-based rendering before they can benefit from novel chart types.
  • Table views are a legitimate visualization; design effort should go into enhancing tables with interaction and pixel-oriented techniques rather than replacing them.
  • Preprocessing tooling should let analysts compare data before and after transformations and should record the logic behind automatic cleaning steps.
  • The ten insights can serve as a checklist for planning or evaluating new visualization solutions aimed at initial exploratory analysis.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The authors do not spell it out, but the ten insights also function as an evaluation rubric: a tool that violates several of them, for example by hiding transformation logic or requiring context switches, should predictably struggle to gain adoption, and that prediction could be tested in a comparative study.
  • Insight 8, which reframes preprocessing as a back-and-forth cycle, points toward a concrete design direction: show partial cleaning results as they become available and let the analyst keep interacting while the tool works.
  • Because the sample is mostly male, mostly IT-industry, and drawn from one country, the frequency ordering of the insights is the most fragile part; a testable extension is to rerun the same protocol with analysts in government, healthcare, and retail and compare which insights reach the mention threshold.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This manuscript reports a semi-structured interview study with thirteen enterprise data analysts, focusing on how visualization can support the preprocessing phase of data mining workflows. The authors describe the participants' profiles, their data analysis processes, prevalent preprocessing activities, data quality concerns, and their current use of visualization. From an iterative, incremental coding of the interview responses, the authors derive a consolidated list of ten insights for visualization tools, compare this list with three related works (RW1, RW2, RW3), and position the list as a requirements baseline for future visualization research. The paper concludes with acknowledged limitations, including the sample's narrow demographic and geographic representativeness.

Significance. If the ten-insight list is robust, it would provide a concise, actionable checklist for visualization researchers and tool builders targeting the preprocessing phase, an area that the authors argue is underserved relative to later-stage visualization. The study's strengths include its transparent description of the coding procedure, the explicit triangulation with three related interview studies, and an honest section on limitations. The paper also offers useful descriptive detail about practitioner workflows and data quality issues. However, the central contribution rests on a small, self-selected, single-country sample, and the manuscript's framing as a 'consolidated set of requirements' exceeds what the evidence can support. The insights are best regarded as candidate requirements or hypotheses for further validation, not as a settled baseline.

major comments (3)
  1. [Section III-C and Section IV, Step 4] There is an inconsistency in the inclusion threshold that determines which items become insights. Section III-C states 'we considered the items reported by more than two participants,' which means at least three participants. In contrast, Section IV, Step 4 states 'the items that were not mentioned by at least two participants were not included,' which means at least two participants. With n=13, the difference between 2 and 3 mentions is material (15% versus 23% of the sample) and directly affects the composition of the central ten-insight list. The authors should reconcile these statements, report the exact number of participants supporting each insight, and justify the chosen threshold.
  2. [Abstract, Section I, and Section V] The abstract and introduction describe the ten insights as 'a consolidated set of requirements' for future visualization research, but Section V explicitly concedes that 'the data collected and its analysis cannot be considered a representation of all data analysts.' Given the sample of thirteen participants, mostly male and mostly working in the IT industry in one country, the requirements framing overstates the evidence. I recommend reframing the contribution as candidate insights or requirements hypotheses that need further validation, and adjusting the abstract and Section I accordingly so that the claims match the scope of the study.
  3. [Section IV and Section V] The consolidated list treats all ten insights uniformly, but their provenance differs: insights 7-9 are unique to this small interview sample, while insight 10 is taken solely from related works, as the authors acknowledge. This mixing is not itself an error, but the final presentation should make the provenance of each insight explicit in the main text (not only in Figure 4) and should discuss the evidentiary status of each category. In particular, the three insights that are not corroborated by any related work should be flagged as tentative, and the manuscript should suggest what additional evidence would be needed to solidify them.
minor comments (5)
  1. [Section III-B] The sentence 'The same environment configuration was used for all participants, face-to-face or online conversations' is ambiguous; clarify whether the same physical setup or the same set of instructions was used across the two modes of interviewing.
  2. [Figure 5 caption] There is a typo in the caption: 'insigths' should be 'insights.'
  3. [Section III-A] The manuscript states that participants were located in three cities in the same country but does not name the country or cities. Naming the country (or at least the region) would help readers calibrate the generalizability of the findings, especially because the study's own Section V emphasizes the sample's limitations.
  4. [Section III-B] The paper notes that 'parts of the sessions were recorded' but does not specify which parts or how the audio was used beyond reviewing notes. A brief clarification of the recording and transcription procedure would improve methodological transparency.
  5. [Section IV, Step 1] The coding description says the list 'started based on the inputs received from participant one' and was then iteratively revised. This is a reasonable approach, but the paper should state whether any of the final insights were added or removed during this iterative process and how disagreements or ambiguous responses were handled.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the ten insights are an empirical synthesis of interview responses and related-work triangulation, with each item's provenance explicitly disclosed.

full rationale

The paper's central claim is a qualitative research finding: a consolidated list of ten insights about how visualization can support preprocessing, derived from semi-structured interviews with thirteen data analysts and then compared with three related interview studies. The derivation chain is transparent and empirical: participant responses were coded iteratively (Section IV), items mentioned by fewer than two participants were excluded, and the final list was ordered by frequency, with insights unique to this study and one insight imported solely from related work clearly labeled. No parameter is fitted to data and then renamed as a prediction; there are no equations, no formal model, and no target quantity being reproduced from its own inputs. The authors do not rely on self-citations for any load-bearing premise, and the related works are used only for triangulation, with the paper explicitly stating which insights came from interviews and which came from the literature. The acknowledged limitation in Section V, that the convenience sample 'cannot be considered a representation of all data analysts,' is a generalizability concern rather than evidence of circularity, because it does not make the resulting list equivalent to its inputs by construction. The paper is therefore self-contained as an interview study and shows no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The central claim rests on methodological assumptions about the accuracy of self-reports, the validity of the inclusion threshold, the representativeness of the related works, and the reliability of the coding. These are normal for qualitative research but are not independently verified.

assumptions (4)
  • domain assumption Participant self-reports are accurate about their data analysis practices.
    The study relies entirely on what participants said during interviews, with no direct observation of their work; only notes and partial recordings were taken (Section III.B).
  • ad hoc to paper The threshold of at least two participants for an insight to be included is a valid way to identify recurring needs.
    Section III.B states this rule without justification; it is a post hoc choice made by the authors rather than a standard qualitative threshold.
  • domain assumption The three related works selected (RW1, RW2, RW3) are representative of prior interview studies in this space.
    Section II describes selection via a systematic review and snowballing, but the final list of three studies is small and all are from the same research community.
  • domain assumption Qualitative thematic coding without inter-rater reliability is a sound basis for the insights.
    Section III.B describes a three-level coding process but does not report multiple coders or reliability checks, so the coding process is potentially subjective.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Visualization in the preprocessing phase: an interview study with enterprise professionals." pith.science (2026). https://pith.science/paper/47YOLHLQ

@misc{pith2026190807894,
  author       = {Pith},
  title        = {Pith review of: Visualization in the preprocessing phase: an interview study with enterprise professionals},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/47YOLHLQ}},
  note         = {Machine review of arXiv:1908.07894}
}
read the original abstract

The current information age has increasingly required organizations to become data-driven. However, analyzing and managing raw data is still a challenging part of the data mining process. Even though we can find interview studies proposing design implications or recommendations for future visualization solutions in the data mining scope, they cover the entire workflow and do not fully focus on the challenges during the preprocessing phase and on how visualization can support it. Moreover, they do not organize a final list of insights consolidating the findings of other related studies. Hence, to better understand the current practice of enterprise professionals in data mining workflows, in particular during the preprocessing phase, and how visualization supports this process, we conducted semi-structured interviews with thirteen data analysts. The discussion about the challenges and opportunities based on the responses of the interviewees resulted in a list of ten insights. This list was compared with the closest related works, improving the reliability of our findings and providing background, as a consolidated set of requirements, for future visualization research papers applied to visual data exploration in data mining. Furthermore, we provide greater details on the profile of the data analysts, the main challenges they face, and the opportunities that arise while they are engaged in data mining projects in diverse organizational areas.

Figures

Figures reproduced from arXiv: 1908.07894 by the authors.

Figure 1
Figure 1. The inclusion criteria for each analyzed study were p [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Additional information on the profile of the thirteen [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Example of workflows used during data analysis. The st [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Complete list of the insights. (Top of figure, dark blu [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Consolidated list of insigths for new visualization [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

46 extracted references · 35 canonical work pages

  1. [1]

    Exploratory Data Mining and Data Cleaning

    Dasu T and Johnson T. Exploratory Data Mining and Data Cleaning . 1 ed. New Y ork, NY , USA: John Wiley & Sons, Inc., 2003. ISBN 0471268518

  2. [2]

    Data Mining: Concepts and Techniques

    Han J, Kamber M and Pei J. Data Mining: Concepts and Techniques . 3rd ed. San Francisco, CA, USA: Morgan Kaufmann Publishers I nc.,

  3. [3]

    Knowledge Discovery in Databases

    Piateski G and Frawley W. Knowledge Discovery in Databases . Cambridge, MA, USA: MIT Press, 1991. ISBN 0262660709

  4. [4]

    The crisp-dm model: the new blueprint for data mining

    Shearer C. The crisp-dm model: the new blueprint for data mining. Journal of data warehousing 2000; 5(4): 13–22

  5. [5]

    Wrangler: Inter active visual specification of data transformation scripts

    Kandel S, Paepcke A, Hellerstein J et al. Wrangler: Inter active visual specification of data transformation scripts. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems . CHI ’11, New Y ork, NY , USA: ACM. ISBN 978-1-4503-0228-9, pp. 3363–33 72. DOI:10.1145/1978942.1979444

  6. [6]

    Quantitative data cleaning for large da tabases, 2008

    Hellerstein JM. Quantitative data cleaning for large da tabases, 2008

  7. [7]

    Competing with High Quality Data: Concepts, Tools, and Techniques for Building a Successful Approach to Data Quali ty

    Jugulum R. Competing with High Quality Data: Concepts, Tools, and Techniques for Building a Successful Approach to Data Quali ty. Wiley,

  8. [8]

    Introduction to Data Mining, (First Edition)

    Tan PN, Steinbach M and Kumar V . Introduction to Data Mining, (First Edition). Boston, MA, USA: Addison-Wesley Longman Publishing Co., Inc., 2005. ISBN 0321321367

Show all 46 references
  1. [9]

    From visual data exploration to visual data mining: A survey

    Ferreira de Oliveira MC and Levkowitz H. From visual data exploration to visual data mining: A survey. IEEE Transactions on Visualization and Computer Graphics 2003; 9(3): 378–394. DOI:10.1109/TVCG.2003. 1207445

  2. [10]

    Interactive Data Visualization: F oundations, Techniques, and Applications, Second Edition - 360 Degree Business

    Ward MO, Grinstein G and Keim D. Interactive Data Visualization: F oundations, Techniques, and Applications, Second Edition - 360 Degree Business. 2nd ed. Natick, MA, USA: A. K. Peters, Ltd., 2015. ISBN 1482257378, 9781482257373

  3. [11]

    The interactive visualization g ap in initial exploratory data analysis

    Batch A and Elmqvist N. The interactive visualization g ap in initial exploratory data analysis. IEEE Transactions on Visualization and Computer Graphics 2018; 24(1): 278–287. DOI:10.1109/TVCG.2017. 2743990

  4. [12]

    Enterprise da ta analysis and visualization: An interview study

    Kandel S, Paepcke A, Hellerstein JM et al. Enterprise da ta analysis and visualization: An interview study. IEEE Transactions on Visualization and Computer Graphics 2012; 18(12): 2917–2926. DOI:10.1109/TVCG. 2012.219

  5. [13]

    Futzing and moseying: I nter- views with professional data analysts on exploration pract ices

    Alspaugh S, Zokaei N, Liu A et al. Futzing and moseying: I nter- views with professional data analysts on exploration pract ices. IEEE Transactions on Visualization and Computer Graphics 2018; : 1–1DOI: 10.1109/TVCG.2018.2865040

  6. [14]

    Guidelines for snowballing in systematic lit erature studies and a replication in software engineering

    Wohlin C. Guidelines for snowballing in systematic lit erature studies and a replication in software engineering. In Proceedings of the 18th International Conference on Evaluation and Assessment in S oftware Engineering. EASE ’14, New Y ork, NY , USA: ACM. ISBN 978-1- 4503-247...

  7. [15]

    Empirical studies in i nformation visualization: Seven scenarios

    Lam H, Bertini E, Isenberg P et al. Empirical studies in i nformation visualization: Seven scenarios. IEEE Transactions on Visualization and Computer Graphics 2012; 18(9): 1520–1536. DOI:10.1109/TVCG.2011. 279

  8. [16]

    Research Methods in Human- Computer Interaction, 2nd Edition

    Lazar J, Feng JH and Hochheiser H. Research Methods in Human- Computer Interaction, 2nd Edition . Elsevier Science, 2017. ISBN 9780128093436

  9. [17]

    What is big data? a cons ensual definition and a review of key research topics

    De Mauro A, Greco M and Grimaldi M. What is big data? a cons ensual definition and a review of key research topics. In AIP conference proceedings, volume 1644. AIP , pp. 97–104

  10. [18]

    Python. Python. https://www.python.org/, 2018. [Online; accessed 12- December-2018]

  11. [19]

    The r project for statistical computing

    R. The r project for statistical computing. https://www.r-project.org/,

  12. [20]

    Databricks: Making big data simple

    Databricks. Databricks: Making big data simple. https://databricks.com/,

  13. [21]

    Knime - the konstanz i nformation miner: V ersion 2.0 and beyond

    Berthold MR, Cebron N, Dill F et al. Knime - the konstanz i nformation miner: V ersion 2.0 and beyond. SIGKDD Explor Newsl 2009; 11(1): 26–31. DOI:10.1145/1656274.1656280

  14. [22]

    Knime: Open for innovation

    KNIME. Knime: Open for innovation. https://www.knime.com/, 2018. [Online; accessed 12-December-2018]

  15. [24]

    The open graph viz platform

    Gephi. The open graph viz platform. https://gephi.org/, 2018. [Online; accessed 12-December-2018]

  16. [25]

    Orange: Data mining to olbox in python

    Demˇ sar J, Curk T, Erjavec A et al. Orange: Data mining to olbox in python. Journal of Machine Learning Research 2013; 14: 2349–2353

  17. [26]

    Gephi: an open source s oftware for exploring and manipulating networks

    Bastian M, Heymann S and Jacomy M. Gephi: an open source s oftware for exploring and manipulating networks. In Third international AAAI conference on weblogs and social media . pp. 361–362

  18. [27]

    Spark: Cluste r computing with working sets

    Zaharia M, Chowdhury M, Franklin MJ et al. Spark: Cluste r computing with working sets. In Proceedings of the 2Nd USENIX Conference on Hot Topics in Cloud Computing . HotCloud’10, Berkeley, CA, USA: USENIX Association, pp. 10–10

  19. [28]

    Apache spark: Unied analytics engine for big da ta

    Spark A. Apache spark: Unied analytics engine for big da ta. https://spark.apache.org/, 2018. [Online; accessed 12-December-2018]

  20. [29]

    Orange: Data mining fruitful and fun

    Orange. Orange: Data mining fruitful and fun. https://orange.biolab.si/,

  21. [30]

    [Online; accessed 12-December-2018]

  22. [31]

    ggplot2 - tidyverse

    ggplot2. ggplot2 - tidyverse. https://ggplot2.tidyverse.org/, 2018. [Online; accessed 12-December-2018]

  23. [32]

    Apache hadoop

    Hadoop A. Apache hadoop. https://hadoop.apache.org/, 2018. [Online; accessed 12-December-2018]

  24. [33]

    Matplotlib: Python plotting - matplotlib 3.0.2 documentation

    Matplotlib. Matplotlib: Python plotting - matplotlib 3.0.2 documentation. https://matplotlib.org/, 2018. [Online; accessed 12-December-2018]

  25. [34]

    seaborn: statistical data visualization - 0

    Seaborn. seaborn: statistical data visualization - 0. 9.0 documenta- tion. https://seaborn.pydata.org/, 2018. [Online; accessed 12-December- 2018]

  26. [35]

    Vim: Visualization and imputation of missing values

    Templ M, Alfons A, Kowarik A et al. Vim: Visualization and imputation of missing values. 11 https://cran.r-project.org/web/packages/VIM/index.html, 2018. [Online; accessed 12-December-2018]

  27. [36]

    Tableau. Tableau. http://www.tableau.com/, 2018. [Online; accessed 12-December-2018]

  28. [37]

    Sas analytics

    SAS. Sas analytics. https://www.sas.com/, 2018. [Online; accessed 12-December-2018]

  29. [38]

    Imputation with the R package VIM

    Kowarik A and Templ M. Imputation with the R package VIM. Journal of Statistical Software 2016; 74(7): 1–16. DOI:10.18637/jss.v074.i07

  30. [39]

    The eyes have it: A task by data type taxon omy for information visualizations

    Shneiderman B. The eyes have it: A task by data type taxon omy for information visualizations. In Proceedings of the 1996 IEEE Symposium on Visual Languages . VL ’96, Washington, DC, USA: IEEE Computer Society. ISBN 0-8186-7508-X, pp. 336–

  31. [40]

    The table lens: Merging graphical and s ymbolic representations in an interactive focus + context visualiz ation for tabular information

    Rao R and Card SK. The table lens: Merging graphical and s ymbolic representations in an interactive focus + context visualiz ation for tabular information. In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems . CHI ’94, New Y ork, NY , USA: ACM. ISBN ...

  32. [41]

    Qlik: Data analytics for modern business intelli gence

    Qlik. Qlik: Data analytics for modern business intelli gence. https://www.qlik.com, 2018. [Online; accessed 12-December-2018]

  33. [42]

    Seeing, sensing, and scrutinizing

    Rensink RA. Seeing, sensing, and scrutinizing. Vision Research 2000; 40(10): 1469 – 1487. DOI:https://doi.org/10.1016/S0042- 6989(00) 00003-1

  34. [43]

    Progressive visual analy tics: User- driven visual exploration of in-progress analytics

    Stolper CD, Perer A and Gotz D. Progressive visual analy tics: User- driven visual exploration of in-progress analytics. IEEE Transactions on Visualization and Computer Graphics 2014; 20(12): 1653–1662. DOI: 10.1109/TVCG.2014.2346574

  35. [44]

    List_Insights_horizontal.JPG

    Kindlmann G and Scheidegger C. An algebraic process for visualization design. IEEE Transactions on Visualization and Computer Graphics 2014; 20(12): 2181–2190. DOI:10.1109/TVCG.2014.2346325 . This figure "List_Insights_horizontal.JPG" is available in "JPG" format from: http://...

  36. [45]

    Designing pixel-oriented visualization techn iques: theory and applications

    Keim D. Designing pixel-oriented visualization techn iques: theory and applications. IEEE Transactions on Visualization and Computer Graphics 2000; 6(1): 59–78. DOI:10.1109/2945.841121

  37. [46]

    Mastering the Information Age: Solving Problems with Visual Analytics, Eurographics Association

    Keim D, Kohlhammer J and Ellis G. Mastering the Information Age: Solving Problems with Visual Analytics, Eurographics Association. Germany: Eurographics Association, 2010. ISBN 978-3-9056 73-77-7

  38. [2011]

    ISBN 0123814790, 9780123814791

Pith tools

Reviewed August 14, 2026 · model on record in the stance chip above.