{"id":"e99991f3-c6b6-4663-a588-d0478f8d7920","arxiv_id":"2411.15430","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Interactive IR researchers mostly discover reusable data through personal connections and publications, rarely through repositories, and system-oriented researchers reuse data more often than user-oriented ones.","lead":"Researchers interviewed 21 interactive information retrieval (IIR) scientists about how they reuse data collected by others. The study maps why they reuse data, how they find it, and what makes a dataset seem reusable, which could inform data-sharing infrastructure.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The system-/user-oriented grouping that drives the central claim is never validated; if the orientation labels are unreliable, the main comparative findings collapse.","rationale":"I identified the most load-bearing concern in good faith. The reader flagged sample representativeness and self-report as the weakest assumption. However, the paper's central contribution is the orientation-based pattern. If the independent variable (orientation) is not reliably measured, the core findings cannot be interpreted, regardless of sample size. The authors treat the two-community distinction as self-evident and do not test it. This is more fundamental than the acknowledged sampling limitations: a non-representative but correctly classified sample could still support a qualitative claim about divergent practices, but a misclassified sample cannot. The proposed check is minimal and feasible from the paper's own Table 1; it directly tests whether the grouping that drives the narrative can be independently recovered. The verdict remains CONDITIONAL (unchanged) because the paper is otherwise transparent and the concern is testable, but the condition should include orientation-label validation.","tokens_in":21373,"tokens_out":5780,"duration_ms":51332,"concrete_test":"Have two independent coders, blinded to the study findings, classify the 21 participants as system-oriented, user-oriented, or mixed using only the research-interest descriptions in Table 1 and a one-paragraph definition of the two orientations taken from Section 4.2. Compute Cohen's kappa between the two coders and between each coder and the authors' labels. If kappa is below 0.6, or if any participant is assigned mixed by either coder, the orientation-based central claim is not robust and would need to be reframed or re-analyzed with a validated grouping.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim is a comparative one: system-oriented IIR researchers reuse external data frequently for ground-truthing, while user-oriented researchers reuse less, within trusted networks, for exploration. This entire contrast rests on assigning each of the 21 participants to one of two camps (12 user-oriented, 9 system-oriented, Section 4.2). Yet the paper provides no explicit classification criteria, no rubric, and no reliability check for this assignment. Table 1 lists research interests, but several entries are ambiguous or mixed (e.g., P1 'recommender systems, metrics for IR system evaluation'; P4 'task-oriented conversational systems'; P17 'user behavior analysis, retrieval models'). If the authors' categorization is not reproducible, the observed differences in motivations (5.1), discovery (5.2), reusability assessment (5.3), and concerns (5.4) could be an artifact of how the sample was split, rather than evidence about the IIR community. The transcript-coding reliability check (Section 4.3) applies to themes, not to the orientation construct. This is a load-bearing internal-validity threat that is not acknowledged in Section 6.1.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative interview study of data reuse practices among 21 Interactive Information Retrieval (IIR) researchers. Through semi-structured interviews and two-round thematic coding, the authors identify motivations for reusing data (exploration and ground-truthing), describe how researchers discover and access data (largely through personal connections and publications rather than repositories), and characterize the criteria used to assess reusability (understandability, trustworthiness, previous usage, and data collection methods). The central comparative claim is that system-oriented researchers reuse external data frequently for ground-truthing and comparability, while user-oriented researchers reuse data less often, mostly within trusted networks, and primarily for exploration. The paper concludes with implications for data-sharing infrastructures and community-level standards in IIR.","tokens_in":21556,"tokens_out":3453,"duration_ms":30737,"significance":"If the findings are accepted, the paper offers a useful empirical map of data reuse in a field where reuse has been advocated but little studied. Its strengths include a transparent methodology: the interview guide and codebook are provided in the supplementary materials, the coding procedure is described in Section 4.3, and inter-coder agreement was checked on a sample of transcripts. The authors also explicitly acknowledge limitations such as self-report and small sample size in Section 6.1. However, the central comparative claim about orientation differences rests on an unvalidated grouping of participants, which is a load-bearing internal-validity threat that the limitations section does not address.","major_comments":[{"comment":"The assignment of the 21 participants to 'user-oriented' and 'system-oriented' groups is central to the paper's contribution, yet the manuscript gives no operational definition of these categories, no coding procedure for how the labels were applied, and no reliability check for this grouping. Table 1 shows several profiles that are plausibly mixed (e.g., P1 'recommender systems, metrics for IR system evaluation'; P4 'task-oriented conversational systems'; P17 'user behavior analysis, retrieval models'). Because Sections 5.1–5.4 and the Discussion contrast these two groups at every turn, the comparative findings would be an artifact of an unreliable split. The authors should either provide a transparent rubric and demonstrate inter-coder agreement for the orientation labels, or re-frame the findings as within-sample patterns without the group contrast.","section":"4.2, Table 1"},{"comment":"The interview guide in Section 9.1 asks participants to consider 'proposed aspects' including data quality, metadata completeness, source credibility, licensing, format, recency, and documentation when evaluating reusability. This priming is likely to shape the responses that underlie the reusability assessment findings in Section 5.3. The paper does not acknowledge this as a methodological limitation, and Section 6.1's limitations list does not mention it. The authors should discuss the potential priming effect and temper the strength of the claims in 5.3.","section":"9.1 / 5.3"},{"comment":"The results report orientation-based differences without indicating how many of the 9 system-oriented and 12 user-oriented participants expressed each pattern. For example, Section 5.1 states that user-oriented participants primarily reused data for exploration and rarely for ground-truthing, but no counts or per-group tallies are provided. With N=21, providing the distribution of responses per theme across the two groups would materially strengthen the credibility of the comparative claims and allow readers to assess the degree of overlap.","section":"5.1–5.4"}],"minor_comments":[{"comment":"The phrase 'It's worth nothing' should read 'It's worth noting' (Section 5.1, discussing the two main purposes of reuse).","section":"5.1"},{"comment":"The word 'regrading' should be 'regarding' (Section 5.4.1, first sentence: 'Challenges in understanding other's data were a major concern for participants regrading data reuse').","section":"5.4.1"},{"comment":"The sentence 'Researchers have also studied resource reuse in IIR and attempted and madetheoretical contributions and provide suggestions future practices' is garbled; it should be rephrased, e.g., 'Researchers have also studied resource reuse in IIR and attempted to make theoretical contributions and provide suggestions for future practices.'","section":"2.3"},{"comment":"The citation 'Bishop (Bishop, 2009)' redundantly repeats the author name inside the parenthetical; it should be 'Bishop (2009)' or '(Bishop, 2009)'.","section":"2.1"},{"comment":"Cohen's kappa values above 0.6 are described as 'substantial agreement,' but 0.61–0.80 is typically labeled substantial; the authors may want to report the exact kappa value and interpret it accordingly rather than using 'exceeded 0.6.'","section":"4.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a well-described qualitative study that fits the journal's scope. The main risk is that the central system-/user-oriented contrast is not methodologically secured; this is fixable within revision by adding a clear classification rubric and either validating it or softening the comparative claims. I recommend major revision rather than rejection because the descriptive findings are valuable and the methodology is otherwise transparent."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the first interview-based study of data reuse practices specifically in IIR, and it's a genuinely useful descriptive map. The findings are plausible: researchers reuse data for exploration and ground-truthing, discover it through personal connections and publications rather than repositories, assess reusability via understandability, trustworthiness, prior usage, and collection methods, and worry about context loss and community acceptance. These align with data-reuse research in other fields, which adds confidence.\n\nThe distinctive contribution is the orientation contrast: system-oriented researchers reuse external data frequently for comparability and ground-truthing, while user-oriented researchers reuse less, within trusted networks, mostly for exploration. The paper is transparent about method—interview guide and codebook in the supplement, two-round coding, and independent coding of three transcripts with kappa above 0.6. Limitations are honestly listed.\n\nThe soft spot is the one the stress-test flagged, and I think it's real: the system/user grouping is never operationalized. Section 4.2 says they recognized two communities and recruited for balance, but there is no rubric, criteria, or reliability check for assigning the 21 participants. Table 1 lists research interests, and several entries are genuinely ambiguous—P1 mixes recommender systems and evaluation metrics, P17 mixes user behavior and retrieval models. Since the main comparative findings in Sections 5 and 6 lean entirely on that split, this is a genuine internal-validity threat. The kappa check covers theme coding, not orientation labels, and the limitation section does not mention this.\n\nIt's not fatal. The descriptive findings stand on their own, and the orientation differences are plausible. But the comparative claims should be treated as exploratory until the classification is validated or at least made explicit and checked. Also, as the authors admit, the sample is small and snowball-recruited, mostly senior PhDs and professors, which limits generalization.\n\nWho should read this: anyone working on data-sharing infrastructure, IIR evaluation methodology, or open science in IR. It deserves a serious referee—it fills a real gap and is honest about its limits. I'd send it to review, with a request to clarify and ideally validate the orientation classification before acceptance.","headline":"First empirical map of IIR data reuse, but the system/user orientation split that drives its main contrast rests on an unvalidated classification; worth reviewing with that caveat.","tokens_in":22076,"tokens_out":2714,"would_cite":true,"duration_ms":24065,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Data reuse in interactive information retrieval splits by research orientation: system-oriented researchers reuse external data routinely, user-oriented researchers reuse mostly within trusted circles.","keywords":["interactive information retrieval","data reuse","data reusability assessment","research data discovery","qualitative interviews","research data infrastructure","system-oriented research","user-oriented research"],"falsifier":"A large-scale survey of IIR researchers, or a log-based study of data repository and lab-website downloads matched against later publication reuse, could test the orientation split: if system-oriented and user-oriented researchers show similar reuse rates and discovery channels when behaviors are observed rather than self-reported, the central contrast would not survive. Alternatively, an ethnographic observation study following researchers through a real data-seeking episode could check whether discovery is as passive and connection-driven as claimed.","tokens_in":1705,"feed_emoji":"🔁","tokens_out":2808,"duration_ms":94525,"temperature":0.7,"pith_summary":"The paper sets out to map how researchers in interactive information retrieval (IIR) actually reuse other people's data — why they do it, where they find it, and how they decide it is good enough to use. Based on 21 semi-structured interviews, it argues that data reuse in IIR splits along methodological orientation: system-oriented researchers routinely reuse external datasets for ground-truthing, generalizing findings, and comparability with prior work, while user-oriented researchers reuse data less often, mostly from labmates and close collaborators, and mostly to explore new questions. It also finds that discovery is driven by personal connections and the academic literature rather than by data repositories, and that reusability is judged through four lenses: understandability, trustworthiness, previous usage, and data collection methods. If these patterns hold, efforts to build data-sharing infrastructure for IIR should focus less on generic repositories and more on surfacing data through the channels researchers already trust.","feed_headline":"Data reuse in interactive search research runs on people, not portals","feed_subtitle":"System- and user-oriented labs find, trust, and reuse data differently — infrastructure should follow that split.","key_machinery":"The analytic engine of the study is the distinction between system-oriented and user-oriented IIR researchers, drawn from the community's own methodological split and used to organize every finding about motivation, discovery, assessment, and concern. The second load-bearing component is the four-aspect account of reusability assessment — understandability (can I make sense of the variables and context?), trustworthiness (is the data reliable and valid?), previous usage (has it been used before, and for what?), and data collection methods (how exactly was it gathered?) — which the authors derive from the interviews and present as the criteria reusers apply. These two components together carry the argument that data reuse in IIR is an interpersonal, context-dependent practice rather than a repository-driven one.","core_discovery":"The central claim is that data reuse in IIR is not a single practice but two different practices shaped by research orientation. System-oriented IIR researchers reuse data often, comfortably drawing on datasets produced by strangers, because reuse gives them ground truth for evaluation, a way to test whether findings generalize across datasets, and comparability with earlier results; for them trust in a dataset can be inherited from the reputation of its producers and from how widely it has already been used. User-oriented researchers reuse data rarely, and almost always from people they know, because their data is deeply tied to specific study designs and contexts; they reuse mainly for exploration and worry more about lost contextual information, ethical issues, and community acceptance. The paper further claims that data discovery is largely passive — through papers, advisors, workshops, and personal contacts — with repositories playing a minor role, and that researchers assess reusability by trying to understand the data, judging its trustworthiness, tracing its previous usages, and scrutinizing how it was collected. The study positions this as an initial empirical map of IIR researchers' data reuse behavior, intended to guide infrastructure and standards for sharing and reusing IIR research data.","pith_inferences":["Editorial extension: because discovery is reported as passive and connection-driven, a data recommendation layer embedded in the literature — recommending datasets related to the paper being read — could plausibly outperform standalone repository search tools.","The trust-in-reputation pattern described in the paper implies a self-reinforcing visibility loop: datasets already used in many papers look more trustworthy, get reused more, and can crowd out equally strong datasets from less-known groups.","A quantitative survey or repository-log analysis across the IIR community could test whether the orientation split generalizes beyond this interview sample, since the paper's own limitations section concedes that self-reports may differ from actual practice.","The user-oriented half's dependence on tacit context suggests that documentation templates alone may not solve reuse; reusable user-study data may require ongoing involvement of the data creators, such as shared protocol ownership or collaboration norms."],"forward_implications":["Data-sharing infrastructure for IIR should prioritize helping reusers find data through publications and trusted contacts, for example by embedding dataset links and provenance records in papers, rather than assuming researchers will actively search general-purpose repositories.","Documentation standards that capture study design, variable definitions, and collection procedures would directly lower the main barrier user-oriented researchers report: loss of contextual information.","Because system-oriented researchers already inherit trust from prior usage, systematic records of dataset provenance and usage histories would make datasets more reusable across the board.","Community-level agreement on data collection and documentation conventions, like the shared confidence TREC datasets enjoy, is a plausible route to expanding reuse beyond personal networks.","Efforts to promote a reuse culture in IIR must address perceived innovation and reliability, not only technical access, since reusers worry about how peers and reviewers will judge studies based on other people's data."],"supporting_citations":[{"why":"Supplies the definition of data reuse — using data collected by others for a purpose different from the original one — that the study adopts throughout.","marker":"Zimmerman, 2008"},{"why":"Provides the reference point for how researchers in another field assess the reusability of colleagues' data, which the interview analysis extends to IIR.","marker":"I. M. Faniel & Jacobsen, 2010"},{"why":"Supplies the data creators' advantage and context-loss framing that organizes the study's concerns about understanding and trusting others' data.","marker":"Pasquetto et al., 2019"},{"why":"Contributes the six-factor model of social scientists' data reuse that informed the preliminary codebook for the interviews.","marker":"R. G. Curty, 2016"},{"why":"Documents IIR-specific challenges in documenting and archiving research resources and proposes principles for improving reuse that this study builds on.","marker":"Gäde et al., 2021"},{"why":"Provides an IIR reusability assessment framework that the paper references when interpreting how researchers evaluate validity and reusability.","marker":"Liu, 2022"},{"why":"Supplies the six dimensions of distance between data creators and reusers that support the paper's analysis of information loss in IIR data.","marker":"Borgman & Groth, 2024"},{"why":"Documents broad scientific attitudes toward and frequency of data reuse, motivating the need for discipline-specific empirical work like this study.","marker":"Tenopir et al., 2020"},{"why":"Provides the semi-structured interview procedure that the study follows for data collection.","marker":"Pickard, 2013"}],"fun_headline_variants":["IIR data reuse splits by research orientation: system vs user","Repositories rarely drive data reuse in interactive search studies","For user-centric IR, data reuse demands familiarity and context","System-oriented researchers reuse data strangers; user-oriented stick to contacts","Data reuse in IIR: passive discovery, trust in provenance, not portals"],"cache_read_input_tokens":24320,"weakest_assumption_plain":"The load-bearing premise is that what 21 purposefully and snowball-recruited researchers said about their own behavior — 17 of whom had reused data, mostly from people they knew — accurately reflects how the wider IIR community goes about data reuse.","fun_headline_variants_meta":{"raw":{"variants":["IIR data reuse splits by research orientation: system vs user","Repositories rarely drive data reuse in interactive search studies","For user-centric IR, data reuse demands familiarity and context","System-oriented researchers reuse data strangers; user-oriented stick to contacts","Data reuse in IIR: passive discovery, trust in provenance, not portals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000214,"raw_usage":{"total_tokens":1473,"prompt_tokens":1041,"completion_tokens":432,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":657,"completion_tokens_details":{"reasoning_tokens":346}},"tokens_in":657,"tokens_out":432,"duration_ms":3937,"temperature":1.0,"reasoning_tokens":346,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:17:36.265207+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A large-scale survey of IIR researchers, or a log-based study of data repository and lab-website downloads matched against later publication reuse, could test the orientation split: if system-oriented and user-oriented researchers show similar reuse rates and discovery channels when behaviors are observed rather than self-reported, the central contrast would not survive. Alternatively, an ethnographic observation study following researchers through a real data-seeking episode could check whether discovery is as passive and connection-driven as claimed.","supporting_citations":[],"review_version":1}