{"id":"61e05be4-e80a-4dd9-92a4-eeb2c4b4d083","arxiv_id":"2504.12495","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A qualitative interview study finds that expert document research is iterative, personal, and socially contextual, and argues NLP tools should model documents as social objects, not just text.","lead":"The authors interview 16 experts in materials science and law or policy about how they find, read, and evaluate documents. They argue that current text-focused NLP tools miss the social context, iterative mental modeling, and personalization these experts rely on.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Interview guide may prime the exact findings: Appendix A Q7 seeds 'iterative mental model' and Q10-12 seed 'social context,' so 'rely extensively' may reflect prompted endorsement rather than unprompted priorities.","rationale":"The reader identified self-report accuracy as the weakest assumption. I sharpen this to a construct-validity problem: several interview questions are leading and directly name the constructs claimed as findings. This matters because the abstract's first clause ('participants' processes... rely extensively on social context') is the empirical foundation for the paper's normative call; if the effect is partly instrument-driven, the comparison to document-centric tools loses its evidentiary base. The proposed re-coding test would separate spontaneous from prompted articulations. This does not move the verdict: the paper remains a useful qualitative needs assessment, but the strong claim needs the caveat the conditional verdict already imposes.","tokens_in":15800,"tokens_out":4200,"duration_ms":44951,"concrete_test":"Re-analyze the 16 transcripts (authors have the data) coding each occurrence of the three headline themes (iteration/mental model, idiosyncrasy, social-context reliance) as 'spontaneous' (raised by participant in response to a non-directive question before the topic is named) or 'prompted' (first named in the interviewer's question, e.g., Q7 or Q10-12). Report the proportion of participants whose social-context reliance is only articulated after Q10-12. If it is above, say, 50%, the central claim about 'rely extensively' should be downgraded to 'endorse when prompted.' If transcripts are unavailable, a re-interview study with open-ended questions (e.g., 'Tell me how you decide whether to trust a document') before any mention of metadata would settle it.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that expert document research is idiosyncratic, iterative, and heavily reliant on social context, and therefore that document-centric NLP better reflects user priorities—rests on interview self-reports. The weakness is not merely that self-reports can diverge from behavior; the interview instrument appears to prime the exact conclusions. Appendix A, Q7 asks: 'To what degree is finding documents or facts an iterative process? Is there a mental model that you have of the space of possible documents that you update as you find new documents?' This states the hypothesized 'iterative mental model' and asks for a degree, inviting endorsement. Q10-12 ask 'How much of that depends on content...', 'How much depends on knowledge of the field...', 'How much depends on metadata, like citations, author affiliations, venue, etc.?', which explicitly segments evaluation into content, field knowledge, and metadata and suggests social-context factors are expected. Quotes such as 'at best...40 to 45%' (40) are direct responses to that segmentation. The grounded-theory coding then extracts themes that were partly supplied by the questions. This does not make the findings false, but it weakens the inference that social context is a spontaneously central priority of participants rather than an interviewer-prompted category.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a qualitative interview study of 16 expert document researchers, 10 in materials science and 6 in law/policy. Using semi-structured interviews and grounded-theory coding, the authors identify recurring task types (local context tasks, global context tasks, and corpus construction) and broader themes including iterative mental-model building, reliance on metadata and social context, and highly personal research processes. They argue that current text-centric NLP systems treat documents as containers for information, whereas existing document-centric approaches in HCI and citation-aware NLP better reflect participants' priorities, and they close with recommendations for accessible, personalizable, iterative, and socially aware tools. The full interview guide is included in Appendix A.","tokens_in":16100,"tokens_out":6345,"duration_ms":68036,"significance":"If the main findings hold, the paper provides a useful, quote-anchored characterization of expert document research and connects it to STS document theory and to existing HCI/NLP tool design. The study's strengths include transparent reporting of the sampling and coding process, inclusion of the complete interview guide, use of verbatim quotes, and an explicit limitations section. The design recommendations are concrete and actionable. However, the central claims about the relative importance of social context and the superiority of document-centric tools are weakened by interview-guide priming and by the absence of any direct participant comparison between document-centric and text-centric systems; the manuscript should either supply additional evidence or qualify these claims.","major_comments":[{"comment":"The main empirical claim that participants 'rely extensively' on social context and that their processes are iterative is at risk of being partly an artifact of the interview instrument. Appendix A, Q7 asks: 'To what degree is finding documents or facts an iterative process? Is there a mental model that you have of the space of possible documents that you update as you find new documents?' This presupposes both iterativeness and an updatable mental model. Q10-12 then ask participants to apportion evaluation among 'content of the document itself,' 'knowledge of the field... not explicitly in the document,' and 'metadata, like citations, author affiliations, venue, etc.', supplying the exact decomposition used in Section 5.6. The quote that content accounts for 'at best... 40 to 45%' (participant 40) is a direct response to that prompted decomposition, so it does not by itself establish that social context was a spontaneously central priority. I recommend that the authors report how often the iterative and social-context themes arose in answer to unprompted questions (e.g., Q3, Q4, Q9, Q13) or, failing that, soften the claim to something like 'when asked, participants described relying on...'.","section":"Section 3 / Appendix A, Q7 and Q10-12"},{"comment":"The conclusion that document-centric NLP tools 'tend to better reflect our participants' priorities' is presented as a direct outcome of the interviews, but participants were not exposed to or asked about such tools. Section 4.3 notes that very few participants had access to advanced NLP document tools, and Section 5 discusses document-aware systems that are 'less accessible outside their research communities.' The mapping from interview descriptions of tasks to the capabilities of document-centric systems is therefore the authors' interpretive synthesis rather than a participant-stated comparison. This is legitimate in a qualitative study, but it should be explicitly labeled as an inference, and the analytic chain linking specific participant statements to specific tool affordances should be laid out in Section 5. Without that, the abstract's comparative claim overstates the empirical support.","section":"Section 5.6 / Section 6"},{"comment":"The paper repeatedly uses quantitative-sounding characterizations: processes are 'consistently informed' by social context (Section 1), personalization is 'a consistent theme' (Section 4.3), and social/metadata signals are 'Perhaps the most commonly used signal' (Section 5.6). However, no coding frequencies, counts of participants per theme, inter-coder agreement, or saturation analysis are reported. For a sample of 16 interviews, these quantifiers are not self-evident. The authors should either report the distribution of codes across participants (for example, how many participants spontaneously mentioned citation chaining, author positionality, or provenance checks) or rephrase these claims as qualitative observations about recurring themes rather than as statements of relative prevalence.","section":"Section 4.3 / Section 5.6"}],"minor_comments":[{"comment":"The abstract contains two typos that should be corrected: 'our participants processes' should be 'our participants' processes' and 'in addition its content' should be 'in addition to its content'.","section":"Abstract"},{"comment":"In the sentence 'These two sources illustrate the a possible origin for the elision of the document,' the phrase 'the a' should be corrected to 'a.' Also, 'a document... serve as a mode of social coordination' should be 'serves as a mode.'","section":"Section 2"},{"comment":"The sentence 'This is different to standard benchmark datasets' should read 'This is different from standard benchmark datasets.'","section":"Section 5.1"},{"comment":"The phrase 'which additional can help address issues' should be rewritten, for example as 'which can additionally help address issues.'","section":"Section 5.3"},{"comment":"The text contains 'original or indended audience'; 'indended' should be 'intended.' In addition, because participants are referenced by number throughout, a small table summarizing each participant's domain, role, and age bracket would improve traceability.","section":"Section 5.6"}],"recommendation":"major_revision","confidential_remarks":"This is a well-written qualitative study with a clear appendix and honest limitations, but the interview-guide priming issue is central to the paper's headline claim and should be addressed directly in revision. The paper also cites several of the authors' own prior works on evaluation critique and identity in NLP; this is natural for a research agenda, but the cluster is noticeable and the framing could benefit from a slightly broader set of voices from the qualitative HCI and information science communities."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nHere's the short version: this is a competent, readable interview study that deserves a real referee, but the central claim should be read as partly shaped by the questions they asked.\n\nWhat's new: 16 interviews with materials scientists and law/policy experts, analyzed with grounded theory into a three-part task taxonomy (local context, global context, corpus construction). The quotes are vivid, and the comparison to current NLP capabilities (IE, OCR, iterative search, terminology, corpus construction, social context) is a useful mapping. The authors are transparent about their convenience sample and expert-only focus in the Limitations section. I agree with the reader that this is a 5-ish significance, 6 novelty paper.\n\nThe soft spots are real, though not fatal. The biggest one is not addressed in the paper's own limitations: the interview guide primes the headline finding. Appendix A Q7 explicitly asks about the 'iterative mental model' and Q10-12 separately seed content, field knowledge, and metadata as evaluation dimensions. So when the paper says participants 'rely extensively' on social context, much of that reliance is an answer to a question that named those categories. The quote about 40-45% content comes straight from that segmentation. This doesn't make the finding false—participants could have pushed back—but it weakens the inference that social context is a spontaneously central priority.\n\nOther nits: no inter-rater reliability numbers, codebook not shared, and the comparison to 'existing NLP approaches' is a bit cherry-picked—they don't systematically survey what text-based LLM tools could do, just assert a mismatch. The small convenience sample is acknowledged, so I won't lean on it.\n\nStill, the taxonomy and the expressed needs feel like real signal. The paper would be a useful anchor for anyone designing document-aware reading support, and the methods section is honest enough to let a careful reader see the scaffolding. I'd send it out for review, with a request that the authors address the priming concern directly—perhaps by reporting which themes were volunteered vs. prompted.\n\nBring it to reading group? Maybe—it's a good discussion piece about how interview guides can shape qualitative findings.","headline":"A useful qualitative needs assessment whose headline claim is partly an artifact of the interview guide; still worth peer review.","tokens_in":16550,"tokens_out":2267,"would_cite":true,"duration_ms":24681,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Domain experts treat documents as social objects, not text containers, so NLP tools that chunk and decontextualize text miss how materials scientists and law/policy researchers actually work.","keywords":["document research","domain experts","qualitative interviews","grounded theory","social context of documents","NLP tool design","information extraction","mental models"],"falsifier":"A behavioral study that records ten to fifteen experts' actual document research sessions (screen capture, queries, reading order, notes) and checks whether workflows are iterative and draw on authorship, provenance, and version differences. If logged sessions show mostly linear, single-query, content-only searches, the self-reported centrality of social context and iteration would be contradicted.","tokens_in":15554,"feed_emoji":"📄","tokens_out":5222,"duration_ms":46748,"temperature":0.7,"pith_summary":"The paper reports interviews with sixteen domain experts in materials science, law, and policy about how they actually perform document research. It argues that these experts' processes are idiosyncratic, iterative, and heavily dependent on social context—authorship, provenance, community terminology, and document versioning—rather than on document content alone. The authors conclude that current text-centric NLP systems, which treat documents as containers of extractable facts, are misaligned with expert practice, and that document-aware tools better reflect how experts work. A sympathetic reader would take the central claim as: useful document-research tools must preserve the document as a unit and support personalization, iteration, and social awareness.","feed_headline":"Expert document work is social, iterative — not text mining","feed_subtitle":"Sixteen interviews show why NLP tools should keep provenance, authorship, and iteration in view.","key_machinery":"The central object is the distinction between the document as an object and the document as a container of text: for the interviewed experts, the document is the unit that carries authorship, provenance, version history, and social meaning. The methodological machinery is a grounded-theory analysis of sixteen semi-structured interviews, coded first openly and then through a closed coding frame, which produced three task categories: local context tasks (within-document extraction), global context tasks (corpus-level mental models built through iteration), and corpus construction (assembling and verifying collections of documents). This taxonomy is what lets the paper compare expert needs against existing NLP capabilities.","core_discovery":"On the paper's own terms, the discovery is that the expert document researchers interviewed describe their work as building and refining mental models of a corpus through repeated, iterative searches, and as relying on social signals—who wrote a document, why, for whom, and how it differs from related versions—to judge relevance and trustworthiness. Participants reported that content alone accounted for at most 40–45 percent of a document's trustworthiness. The paper contrasts this with the dominant NLP framing of documents as containers of information that can be chunked, decontextualized, and summarized, and argues that document-centric approaches from adjacent fields better match expert priorities, even though they are less accessible. The conclusion is a call for NLP systems that are accessible, personalizable, iterative, and socially aware.","pith_inferences":["Beyond the paper: a testable design extension would be a document research assistant that records a user's validation moves (which metadata they check, what they flag) and turns them into personalizable filters, then measures whether such filters reduce time spent on irrelevant documents.","Beyond the paper: the social-context finding likely generalizes to other expert document work—journalism, medicine, archives—where provenance and versioning matter; a replication with screen-capture logging would test whether the interview-reported processes match observed behavior.","Beyond the paper: the 40–45 percent trustworthiness figure, if reliable, quantifies how much of document evaluation is non-textual; a follow-up rating study could measure the same ratio with a larger sample and ground tool design in it."],"forward_implications":["If the claim is right, retrieval and question-answering systems that split documents into semantic chunks and discard the source document will misalign with expert workflows, because experts use the document-level unit to evaluate trust.","Document-aware tools for scientific literature—citation-based exploration, enriched PDF readers—should be treated as better models for expert support than generic text-only assistants.","Evaluation of NLP tools for document research should include whether they support iterative exploration and verification, not just one-shot accuracy.","Tools for law, policy, and other non-scientific domains need the same document-aware features (provenance, versioning, authorship) that scientific tools have, rather than treating those features as science-only.","Systems that adapt to an individual researcher's personal heuristics for relevance and quality would match expert practice better than one-size-fits-all summarizers."],"supporting_citations":[{"why":"Provides the closest prior interview/think-aloud study of how data scientists review literature, giving a comparison point for expert reading processes.","marker":"Mysore et al. (2023)"},{"why":"Supplies the finding that physicists use purpose-laden schemas combining content and authorship, a pattern echoed by the materials scientists in this study.","marker":"Bazerman (1985)"},{"why":"Supplies the theory that documents coordinate social practices rather than merely communicate, underpinning the paper's social-context framing.","marker":"Brown and Duguid (1996)"},{"why":"Documents the historical shift from documents to disembodied information, which the paper identifies as an origin of the text-container view in NLP.","marker":"Lund (2009)"},{"why":"Argues for the contingent, social role of scientific publication, supporting the claim that documents are traces of social processes.","marker":"Frohmann (2004)"},{"why":"Exemplifies document-aware reading support with an enriched PDF reader, used as evidence that document-centric tools exist and work.","marker":"Lo et al. (2024)"},{"why":"Exemplifies citation-aware visual exploration of publications, a document-aware approach closer to expert priorities.","marker":"He et al. (2019)"},{"why":"Provides the translational NLP call for application-focused research, which the paper extends to document research tools.","marker":"Newman-Griffis et al. (2021)"}],"fun_headline_variants":["Document research isn't just text — it's social context","Why NLP tools need to see documents, not just text","Expert document trust: >50% from context, not content","Social signals drive document trust more than text content","NLP's blind spot: documents are social artifacts, not just text"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The study assumes that what the sixteen experts said in semi-structured interviews about their document research accurately describes what they actually do, since no direct observation or behavior logging was used.","fun_headline_variants_meta":{"raw":{"variants":["Document research isn't just text — it's social context","Why NLP tools need to see documents, not just text","Expert document trust: >50% from context, not content","Social signals drive document trust more than text content","NLP's blind spot: documents are social artifacts, not just text"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000974,"raw_usage":{"total_tokens":4103,"prompt_tokens":873,"completion_tokens":3230,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":489,"completion_tokens_details":{"reasoning_tokens":3147}},"tokens_in":489,"tokens_out":3230,"duration_ms":21489,"temperature":1.0,"reasoning_tokens":3147,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T12:29:40.412890+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A behavioral study that records ten to fifteen experts' actual document research sessions (screen capture, queries, reading order, notes) and checks whether workflows are iterative and draw on authorship, provenance, and version differences. If logged sessions show mostly linear, single-query, content-only searches, the self-reported centrality of social context and iteration would be contradicted.","supporting_citations":[],"review_version":1}