{"id":"2af7a80f-793c-4763-811d-3e554702349f","arxiv_id":"2501.09862","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Ten experts on the same complex data library held diverse, incomplete mental models that diverged from the data model they themselves had built and used.","lead":"The paper interviews ten researchers who build and use a complex data library, asking them to describe and draw the data model the library stores. It finds that team members, including the people who designed the model, hold different and incomplete mental pictures of that data, which can stall analysis.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The control question added to distinguish 'forgetting' from 'not understanding' is never reported; without its results, omitted components cannot be read as mental-model gaps.","rationale":"The reader identified the shared-context omission risk as the weakest assumption; the missing Q7 results are a more specific, internally designed check on exactly that risk. The control question was added precisely to separate 'forgetting the hierarchy' from 'not understanding it,' so failing to report its outcomes leaves the central claim's key inference untested. This does not overturn the paper's contribution: the qualitative themes, the diversity of drawings, and the two suggested hazards remain useful, and the authors are transparent about their positionality and method. But the missing control results should be a condition of acceptance, because they are cheap to report and directly bear on whether the findings describe mental-model content or merely in-the-moment communicative choices. I retain the reader's CONDITIONAL verdict, adding this specific condition to the existing requests for supplements and protocol-revision reporting. My concern is a partial agreement with the reader rather than full agreement because the reader framed the issue as an interpretive assumption, whereas I see an unreported internal validation step that could settle it empirically.","tokens_in":18997,"tokens_out":4088,"duration_ms":44428,"concrete_test":"Extract the Q7 drawings and transcripts for the nine participants who received the control question (P2–P10) and code each for (a) presence of a tree with A calling both B and C and (b) presence of associated duration data. If all nine pass, the central finding should be reframed as spontaneous recall differences under inferred shared context, with hazards described as retrieval or expression costs rather than representational absence; if any fail, the current forgetting framing survives for those participants. Report per-participant outcomes in the paper or supplement.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central inference is that omitted parts of the reified model reveal divergent mental models. Section 4.2 introduces Q7 as a control 'to verify that they could abstract and describe a small call tree and its associated data, to separate \"forgetting\" the hierarchy versus not understanding it,' and notes it was added after the first interview. Yet Section 5 never reports what P2–P10 drew for Q7. This is the exact experiment that would validate or invalidate the forgetting interpretation: if all nine drew A calls B and C correctly, then the interview omissions are better explained by shared-context conversational omission or retrieval failure, not by absence of the hierarchy in the mental model. The 'parallel hazards' claim would then be overstated, because participants could reason correctly when the task required it. If several failed Q7, the finding is strongly supported. Because the paper does not disclose these outcomes, the load-bearing distinction between 'does not know' and 'did not say' is unresolved. The Section 4.5 positionality note acknowledges the shared-familiarity risk but does not supply the control results that would mitigate it.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative study of ten data workers who develop and use a shared domain-specific data analysis library (EnsembleAPI). Through semi-structured interviews, sketching tasks, and analysis-task questions, the authors elicited participants' mental data models (MDMs) and compared them with the reified data model (RDM) embodied in the library's data structures. The central finding is that participants' mental models were diverse and frequently diverged from the RDM, even for the designers of that model. The authors identify two 'parallel hazards': (1) an inaccurate mental model can prevent a worker from planning an analysis, and (2) an incomplete understanding of the reified model can prevent a worker from expressing an analysis in code. They also report themes about data sources, analysis workflows, and the particular haziness of metadata, and they propose design implications for data analysis tools and data-model design.","tokens_in":19061,"tokens_out":3957,"duration_ms":41748,"significance":"If the findings hold, this is a valuable empirical contribution to HCI and visualization research on mental models of data: it extends prior work such as Williams et al. [53] to a real, complex, heterogeneous data model used by experts over an extended period, and it documents divergence even among the model's designers. The study has notable strengths: it recruited nearly the entire relevant population (10 of 11 team members), combined multiple elicitation methods (description, ranking, sketching, task questions), grounds its themes in direct quotes and participant drawings, and includes an unusually candid positionality statement and limitations section. The parallel-hazards framing and the metadata findings are likely to be useful to designers of data science tools. However, the paper's load-bearing inference from omissions to forgetting rests on a control question (Q7) whose results are never reported, and several cross-participant claims are affected by protocol changes made during data collection. These issues are fixable but currently leave the central claim less strongly supported than the abstract suggests.","major_comments":[{"comment":"Section 4.2 introduces Q7 as a control 'to verify that they could abstract and describe a small call tree and its associated data, to separate \"forgetting\" the hierarchy versus not understanding it,' and notes it was added after the first interview. Section 5 never reports what P2-P10 drew for Q7. This is the experiment that would validate or invalidate the interpretation that omitted components reflect gaps in mental models. If all nine drew A calls B and C correctly, the interview omissions would be better explained by shared-context conversational omission or retrieval failure, and the 'parallel hazards' claim would be substantially weakened. If several failed, the finding would be strongly supported. Please report the control outcomes for all participants who received Q7, or reframe the RQ1 and T1.1 claims as being about what participants spontaneously expressed rather than what they remembered.","section":"Section 4.2 and Section 5.1 (T1.1, RQ1)"},{"comment":"The protocol was refined during the study: Q5b was changed and Q7 was added after the first interview. The central comparative evidence in Figure 4 treats recall across all ten participants, but P1 did not receive the same protocol as P2-P10. This is a potential confound for cross-participant comparisons such as the statement in Section 5.3 that P1 had 'the most sparse description and representation.' Please provide per-participant protocol-version information and a sensitivity analysis excluding P1, or hedge the cross-participant comparisons accordingly.","section":"Section 4.2, Section 5, and Figure 4"},{"comment":"Section 5.1 states that 'The statistics table was most forgotten' but then excludes it from further analysis 'because it exists for derived data and is not populated on collection.' Including this component in a forgetting finding is misleading, since participants were not expected to mention a table that is not part of initial data collection. Relatedly, Section 6.2.1's first hazard claims that 'sufficiently inaccurate mental models resulted in errors which would prevent practical analysis,' but the main supporting evidence is two accidental dimension-drops (P6 in their drawing and P7 in task 5a) that participants corrected during the interview. The evidence supports difficulty, slowdown, and the need for re-scaffolding, but it does not clearly support 'prevent.' Please align the hazard wording with the strength of the evidence.","section":"Section 5.1 and Section 6.2.1"}],"minor_comments":[{"comment":"The single-coder inductive phase is justified by the reflexive thematic analysis approach, but the supplementary materials are referenced without stating how they can be accessed. Including the codebook and an excerpt of the theme-development process would strengthen the audit trail and help readers assess the trustworthiness of the inductive themes.","section":"Section 4.4"},{"comment":"Figure 4's timeline would be easier to interpret if it distinguished explicit mention, implicit allusion, and inclusion in a drawing; the text describes all three modes (e.g., P6 and P10 alluding to performance data rather than naming it), but the figure appears to collapse them.","section":"Section 5.1 and Figure 4"},{"comment":"The placement of participants on the utilitarian-to-structural and seniority axes is described only as 'qualitatively assessed'; please state the criteria or coding used to place each participant so that the figure is reproducible and not read as arbitrary.","section":"Figure 8"},{"comment":"There are several typographical errors: the abstract and introduction contain 'a a different mental model', RQ1 in Section 1 reads 'an known form', and Section 6.2.3 contains 'thier mental data model'. These should be corrected in revision.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a good fit for the venue and the empirical contribution is valuable. The main blocker is the unreported Q7 control, which is directly load-bearing for the 'forgetting' interpretation; however, this is an addressable revision rather than a fundamental flaw. I see no derivational circularity: the comparison with Williams et al. [53] is external and appropriate, and the paper does not reduce its central claim to any cited result. The single-coder inductive analysis is a defensible methodological choice under reflexive TA, provided the supplementary materials are made available. I would not reject on the current evidence, but the authors should either supply the control results or adjust the central claims to what the reported data actually support."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: this is a genuine empirical contribution, not a repackaging. Ten people who all work with the same reified data model—most of them helped design it—show markedly different mental models of it, and the divergence has practical consequences. The \"two parallel hazards\" framing (can't reason about data because your mental model is wrong; can't express a correct analysis because you don't understand the reified model) is useful and well-supported by quotes and sketches.\n\nThe study extends Williams et al. from non-experts with small unfamiliar datasets to experts with a complex, heterogeneous model they've used for over a year. That's a real gap and this paper fills it. The method is appropriate for the question: semi-structured interviews, sketching, task scenarios. The timeline of when components were recalled (Figure 4) is a nice artifact, and the metadata-as-grab-bag theme rings true and has design implications.\n\nNow the soft spot. Section 4.2 introduces Q7 as a control added after the first interview, explicitly to separate \"forgetting the hierarchy\" from \"not understanding it.\" Section 5 never reports what P2–P10 drew. That matters because the paper's central inference is that omitted components reveal mental-model gaps. If everyone drew A calls B and C correctly, then the omissions are better read as shared-context ellipsis or retrieval failure, not absence from the model. The authors flag the shared-familiarity risk in 4.5, but they built a control for it and then don't report it. That's a real gap, and it's easily fixable.\n\nSmaller issues: the inductive coding used a single coder (they justify this via Braun and Clarke, and the deductive phase had three), and the interview protocol was refined mid-study—both are minor and disclosed. The n=10 is small but it's the entire team minus one, which is defensible. The citation pattern is clean; coauthored prior work is used as an external comparison, not as an input to the finding.\n\nBottom line: this deserves a serious referee, not a desk reject. A reviewer should ask for the Q7 results and a data-availability statement. If the control came out clean, the forgetting interpretation needs softening; if it didn't, the central claim gets stronger. Either way the paper is worth engaging.","headline":"Solid qualitative study of expert mental models of a shared reified data model, but the unreported control question leaves the central 'forgetting' claim partly unverified.","tokens_in":19707,"tokens_out":2283,"would_cite":true,"duration_ms":23378,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A team of ten expert data workers who all use the same code-encoded data model carried diverse mental pictures of it, and the mismatches blocked both planning and coding analyses.","keywords":["mental models","data models","heterogeneous data","thematic analysis","exploratory data analysis","data workers","high performance computing","interviews"],"falsifier":"A replication using an interviewer with no knowledge of EnsembleAPI, who never prompts for specific components, would settle the point: if most participants then spontaneously describe the full reference model—call tree, metadata, performance and statistics tables, and their links—the reported forgetting and divergence would be shown to be partly a shared-context artifact.","tokens_in":18691,"feed_emoji":"🧠","tokens_out":4090,"duration_ms":38042,"temperature":0.7,"pith_summary":"This paper sets out to show that even expert data workers who design and maintain a shared, code-encoded data structure do not hold a single mental picture of it. Studying ten members of one team that built and used a heterogeneous data model—a call tree joined to performance, metadata, and statistics tables—the authors find that participants' mental data models were diverse and often omitted core components, including among the model's designers. The authors argue this divergence creates two parallel hazards: an incomplete mental model can block deciding what analysis to run, and an incomplete grasp of the stored model can block expressing that analysis in code. The stakes are practical: if the team that built a data model cannot reliably recall it, the wider audience of data workers using similar libraries likely needs interfaces that surface and bridge these mismatches.","feed_headline":"Even the designers of a complex data model hold conflicting views of it","feed_subtitle":"Interview study of ten experts shows their mental pictures of the data diverge from the code, blocking analysis.","key_machinery":"The paper's central objects are the Mental Data Model (MDM), the internal and subjective representation of data a worker carries, and the Reified Data Model (RDM), the external, shared representation encoded in code or documentation. The argument runs on the measured gap between the two, elicited through semi-structured interviews, freehand sketching, task-based recall questions, and a control question that separates genuine forgetting from inability to abstract a simple hierarchy; reflexive thematic analysis organizes the transcripts into themes. The two 'parallel hazards'—one rooted in the mental model, one in the reified model—are the mechanism that converts the observed mismatch into concrete analysis failures.","core_discovery":"The central claim is that participants had diverse mental data models that differed from the reified data model, even among team members who had designed the model. In interviews and sketches, the performance table dominated most mental models, while the call tree and the metadata table were hazy, implicit, or missing; some participants thought hierarchically and others tabularly, and senior team members often described their data through the final charts of an analysis rather than the stored structure. The authors interpret this as evidence of two parallel hazards: a data worker with an inaccurate mental model may be unable to do the analyses they want regardless of the reified model, and a worker who knows what they want but does not understand the reified model may be unable to express it in code. From these observations they recommend design interventions—data-model reminders, visual recommendation, graphical scripting bridges, structured metadata—rather than trying to force everyone onto one canonical model.","pith_inferences":["An extension the authors leave implicit: if this divergence appears within a team of ten experts who designed the model, the mismatch is likely at least as large for casual users of similar data libraries, which argues for default-on data-model views in analysis tools.","A testable extension would be an intervention study: give analysts a lightweight always-visible schematic of the data model and measure whether task completion time and code errors drop relative to a control group.","The findings suggest data-model recall could be improved by design, such as naming and visually distinguishing components consistently across documentation, API, and plots; the paper gestures at this but does not test it."],"forward_implications":["Data analysis tools should surface the stored data model, for example through lightweight embedded overviews, so workers do not have to recall it from memory.","Metadata cannot be treated as a grab bag: giving metadata machine-readable semantic types would let tools guide grouping, filtering, and display.","Interfaces should embrace multiple coexisting mental models, including both table-centric and hierarchy-centric views, rather than forcing one canonical representation.","Probing the intended users' mental models early, before a data model is fixed in code, may reduce later engineering debt and analysis failures.","Visual recommender and graphical-scripting approaches that show how a chart is derived from the stored model could bridge chart-based mental models to the reified model."],"supporting_citations":[{"why":"Closest prior work; it showed non-experts hold diverse data representations for small unfamiliar datasets, the baseline this study extends to expert long-term users.","marker":"[53]"},{"why":"Supplies the reflexive thematic analysis methodology used to develop themes from interviews and sketches.","marker":"[16]"},{"why":"Defines the data abstraction types (tables, networks, geometry) used to characterize the reference model as heterogeneous.","marker":"[40]"},{"why":"Defines the 'data workers' population that the paper's implications target.","marker":"[32]"},{"why":"Supplies the pitfall that task elicitation from tool developers differs from front-line analysts, used to interpret role-based differences.","marker":"[46]"},{"why":"Argues for more human-legible metadata handling, which the paper's metadata recommendation builds on.","marker":"[4]"},{"why":"Provides the embedded continuous data profiling idea that the paper suggests as a model-reminder intervention.","marker":"[22]"}],"fun_headline_variants":["Even data model designers see their data differently","Mental models of data clash with the code—even for creators","Ten experts, one data model, many conflicting mental pictures","Designers of a data model can't agree on what it means","Data workers' internal models diverge from the reified code"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that a component absent from a participant's description or drawing is genuinely missing from their mental model, rather than something they assumed the interviewer already knew.","fun_headline_variants_meta":{"raw":{"variants":["Even data model designers see their data differently","Mental models of data clash with the code—even for creators","Ten experts, one data model, many conflicting mental pictures","Designers of a data model can't agree on what it means","Data workers' internal models diverge from the reified code"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1478,"prompt_tokens":870,"completion_tokens":608,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":486,"completion_tokens_details":{"reasoning_tokens":526}},"tokens_in":486,"tokens_out":608,"duration_ms":6600,"temperature":1.0,"reasoning_tokens":526,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T19:36:43.882990+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A replication using an interviewer with no knowledge of EnsembleAPI, who never prompts for specific components, would settle the point: if most participants then spontaneously describe the full reference model—call tree, metadata, performance and statistics tables, and their links—the reported forgetting and divergence would be shown to be partly a shared-context artifact.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Closest prior work; it showed non-experts hold diverse data representations for small unfamiliar datasets, the baseline this study extends to expert long-term users."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the reflexive thematic analysis methodology used to develop themes from interviews and sketches."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the data abstraction types (tables, networks, geometry) used to characterize the reference model as heterogeneous."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the 'data workers' population that the paper's implications target."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the pitfall that task elicitation from tool developers differs from front-line analysts, used to interpret role-based differences."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Argues for more human-legible metadata handling, which the paper's metadata recommendation builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the embedded continuous data profiling idea that the paper suggests as a model-reminder intervention."}],"review_version":1}