REVIEW 3 major objections 5 minor 16 references
The Landscape of Data Reuse in Interactive Information Retrieval: Motivations, Sources, and Evaluation of Reusability
T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash
Pith's one-line read Data reuse in interactive information retrieval splits by research orientation: system-oriented researchers reuse external data routinely, user-oriented researchers reuse mostly within trusted circles.
desk verdict First empirical map of IIR data reuse, but the system/user orientation split that drives its main contrast rests on an unvalidated classification; worth reviewing with that caveat. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The analytic engine of the study is the distinction between system-oriented and user-oriented IIR researchers, drawn from the community's own methodological split and used to organize every finding about motivation, discovery, assessment, and concern. The second load-bearing component is the four-aspect account of reusability assessment — understandability (can I make sense of the variables and context?), trustworthiness (is the data reliable and valid?), previous usage (has it been used before, and for what?), and data collection methods (how exactly was it gathered?) — which the authors derive from the interviews and present as the criteria reusers apply. These two components together carry the argument that data reuse in IIR is an interpersonal, context-dependent practice rather than a repository-driven one.
What would settle it
A large-scale survey of IIR researchers, or a log-based study of data repository and lab-website downloads matched against later publication reuse, could test the orientation split: if system-oriented and user-oriented researchers show similar reuse rates and discovery channels when behaviors are observed rather than self-reported, the central contrast would not survive. Alternatively, an ethnographic observation study following researchers through a real data-seeking episode could check whether discovery is as passive and connection-driven as claimed.
Extended reading notes
Core claim
The central claim is that data reuse in IIR is not a single practice but two different practices shaped by research orientation. System-oriented IIR researchers reuse data often, comfortably drawing on datasets produced by strangers, because reuse gives them ground truth for evaluation, a way to test whether findings generalize across datasets, and comparability with earlier results; for them trust in a dataset can be inherited from the reputation of its producers and from how widely it has already been used. User-oriented researchers reuse data rarely, and almost always from people they know, because their data is deeply tied to specific study designs and contexts; they reuse mainly for exploration and worry more about lost contextual information, ethical issues, and community acceptance. The paper further claims that data discovery is largely passive — through papers, advisors, workshops, and personal contacts — with repositories playing a minor role, and that researchers assess reusability by trying to understand the data, judging its trustworthiness, tracing its previous usages, and scrutinizing how it was collected. The study positions this as an initial empirical map of IIR researchers' data reuse behavior, intended to guide infrastructure and standards for sharing and reusing IIR research data.
Load-bearing premise
The load-bearing premise is that what 21 purposefully and snowball-recruited researchers said about their own behavior — 17 of whom had reused data, mostly from people they knew — accurately reflects how the wider IIR community goes about data reuse.
Editorial extensions
If this is right
- Data-sharing infrastructure for IIR should prioritize helping reusers find data through publications and trusted contacts, for example by embedding dataset links and provenance records in papers, rather than assuming researchers will actively search general-purpose repositories.
- Documentation standards that capture study design, variable definitions, and collection procedures would directly lower the main barrier user-oriented researchers report: loss of contextual information.
- Because system-oriented researchers already inherit trust from prior usage, systematic records of dataset provenance and usage histories would make datasets more reusable across the board.
- Community-level agreement on data collection and documentation conventions, like the shared confidence TREC datasets enjoy, is a plausible route to expanding reuse beyond personal networks.
- Efforts to promote a reuse culture in IIR must address perceived innovation and reliability, not only technical access, since reusers worry about how peers and reviewers will judge studies based on other people's data.
Reading between the lines
- Editorial extension: because discovery is reported as passive and connection-driven, a data recommendation layer embedded in the literature — recommending datasets related to the paper being read — could plausibly outperform standalone repository search tools.
- The trust-in-reputation pattern described in the paper implies a self-reinforcing visibility loop: datasets already used in many papers look more trustworthy, get reused more, and can crowd out equally strong datasets from less-known groups.
- A quantitative survey or repository-log analysis across the IIR community could test whether the orientation split generalizes beyond this interview sample, since the paper's own limitations section concedes that self-reports may differ from actual practice.
- The user-oriented half's dependence on tacit context suggests that documentation templates alone may not solve reuse; reusable user-study data may require ongoing involvement of the data creators, such as shared protocol ownership or collaboration norms.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper reports a qualitative interview study of data reuse practices among 21 Interactive Information Retrieval (IIR) researchers. Through semi-structured interviews and two-round thematic coding, the authors identify motivations for reusing data (exploration and ground-truthing), describe how researchers discover and access data (largely through personal connections and publications rather than repositories), and characterize the criteria used to assess reusability (understandability, trustworthiness, previous usage, and data collection methods). The central comparative claim is that system-oriented researchers reuse external data frequently for ground-truthing and comparability, while user-oriented researchers reuse data less often, mostly within trusted networks, and primarily for exploration. The paper concludes with implications for data-sharing infrastructures and community-level standards in IIR.
Significance. If the findings are accepted, the paper offers a useful empirical map of data reuse in a field where reuse has been advocated but little studied. Its strengths include a transparent methodology: the interview guide and codebook are provided in the supplementary materials, the coding procedure is described in Section 4.3, and inter-coder agreement was checked on a sample of transcripts. The authors also explicitly acknowledge limitations such as self-report and small sample size in Section 6.1. However, the central comparative claim about orientation differences rests on an unvalidated grouping of participants, which is a load-bearing internal-validity threat that the limitations section does not address.
major comments (3)
- [4.2, Table 1] The assignment of the 21 participants to 'user-oriented' and 'system-oriented' groups is central to the paper's contribution, yet the manuscript gives no operational definition of these categories, no coding procedure for how the labels were applied, and no reliability check for this grouping. Table 1 shows several profiles that are plausibly mixed (e.g., P1 'recommender systems, metrics for IR system evaluation'; P4 'task-oriented conversational systems'; P17 'user behavior analysis, retrieval models'). Because Sections 5.1–5.4 and the Discussion contrast these two groups at every turn, the comparative findings would be an artifact of an unreliable split. The authors should either provide a transparent rubric and demonstrate inter-coder agreement for the orientation labels, or re-frame the findings as within-sample patterns without the group contrast.
- [9.1 / 5.3] The interview guide in Section 9.1 asks participants to consider 'proposed aspects' including data quality, metadata completeness, source credibility, licensing, format, recency, and documentation when evaluating reusability. This priming is likely to shape the responses that underlie the reusability assessment findings in Section 5.3. The paper does not acknowledge this as a methodological limitation, and Section 6.1's limitations list does not mention it. The authors should discuss the potential priming effect and temper the strength of the claims in 5.3.
- [5.1–5.4] The results report orientation-based differences without indicating how many of the 9 system-oriented and 12 user-oriented participants expressed each pattern. For example, Section 5.1 states that user-oriented participants primarily reused data for exploration and rarely for ground-truthing, but no counts or per-group tallies are provided. With N=21, providing the distribution of responses per theme across the two groups would materially strengthen the credibility of the comparative claims and allow readers to assess the degree of overlap.
minor comments (5)
- [5.1] The phrase 'It's worth nothing' should read 'It's worth noting' (Section 5.1, discussing the two main purposes of reuse).
- [5.4.1] The word 'regrading' should be 'regarding' (Section 5.4.1, first sentence: 'Challenges in understanding other's data were a major concern for participants regrading data reuse').
- [2.3] The sentence 'Researchers have also studied resource reuse in IIR and attempted and madetheoretical contributions and provide suggestions future practices' is garbled; it should be rephrased, e.g., 'Researchers have also studied resource reuse in IIR and attempted to make theoretical contributions and provide suggestions for future practices.'
- [2.1] The citation 'Bishop (Bishop, 2009)' redundantly repeats the author name inside the parenthetical; it should be 'Bishop (2009)' or '(Bishop, 2009)'.
- [4.3] Cohen's kappa values above 0.6 are described as 'substantial agreement,' but 0.61–0.80 is typically labeled substantial; the authors may want to report the exact kappa value and interpret it accordingly rather than using 'exceeded 0.6.'
Circularity Check
No significant circularity: the interview findings are grounded in participant data, with only minor non-load-bearing self-citation.
full rationale
This paper is an interview-based qualitative study; it contains no equations, fitted parameters, or quantitative derivation chain whose outputs could reduce to its inputs. The central comparative claim—that system-oriented IIR researchers reuse external data frequently for ground-truthing/comparability while user-oriented researchers reuse less, mostly within trusted networks and for exploration—is presented as an empirical result of coding 21 semi-structured interviews, not as a logical consequence of the orientation labels used in recruitment. Participants in Section 4.2 were recruited by research orientation (system-oriented vs. user-oriented), but the reuse behaviors, motivations, discovery channels, reusability criteria, and concerns reported in Sections 5.1–5.4 were elicited and coded from transcripts, so the outcome categories are not definitionally entailed by the input categories. The preliminary codebook in Section 4.3 was informed by prior frameworks including Liu (2022), a self-citation by the corresponding author; however, the paper explicitly states that new themes and sub-themes were allowed to emerge and that inter-coder agreement was checked, and the findings are not derived from that framework by construction. The one Results passage citing Liu (2022) for user-oriented validity assessment is corroborative rather than load-bearing, and the same section quotes participant reasoning. The skeptic's concern that the orientation split is not validated is a legitimate internal-validity limitation, but it is not a circularity: nothing in the definition of 'system-oriented' or 'user-oriented' analytically forces the observed differences in reuse frequency, motivation, or concern. Overall, the derivation chain is self-contained and the reported patterns are grounded in the interview data, so no circular step is present; the only mild caveat is the presence of self-citations, which are not load-bearing here.
Assumptions & free parameters
assumptions (3)
- domain assumption Participants' self-reported accounts reflect their actual data reuse practices.
- domain assumption The purposeful plus snowball sample represents the diversity of IIR researchers.
- domain assumption Cohen's kappa above 0.6 indicates reliable coding.
Cite this review
Pith. "Pith review of The Landscape of Data Reuse in Interactive Information Retrieval: Motivations, Sources, and Evaluation of Reusability." pith.science (2026). https://pith.science/paper/V7PDPDM6
@misc{pith2026241115430,
author = {Pith},
title = {Pith review of: The Landscape of Data Reuse in Interactive Information Retrieval: Motivations, Sources, and Evaluation of Reusability},
year = {2026},
howpublished = {\url{https://pith.science/paper/V7PDPDM6}},
note = {Machine review of arXiv:2411.15430}
}
read the original abstract
Sharing and reusing research data can effectively reduce redundant efforts in data collection and curation, especially for small labs and research teams conducting human-centered system research, and enhance the replicability of evaluation experiments. Building a sustainable data reuse process and culture relies on frameworks that encompass policies, standards, roles, and responsibilities, all of which must address the diverse needs of data providers, curators, and reusers. To advance the knowledge and accumulate empirical understandings on data reuse, this study investigated the data reuse practices of experienced researchers from the area of Interactive Information Retrieval (IIR) studies, where data reuse has been strongly advocated but still remains a challenge. To enhance the knowledge on data reuse behavior and reusability assessment strategies within IIR community, we conducted 21 semi-structured in-depth interviews with IIR researchers from varying demographic backgrounds, institutions, and stages of careers on their motivations, experiences, and concerns over data reuse. We uncovered the reasons, strategies of reusability assessments, and challenges faced by data reusers within the field of IIR as they attempt to reuse researcher data in their studies. The empirical finding improves our understanding of researchers' motivations for reusing data, their approaches to discovering reusable research data, as well as their concerns and criteria for assessing data reusability, and also enriches the on-going discussions on evaluating user-generated data and research resources and promoting community-level data reuse culture and standards.
Reference graph
Works this paper leans on
-
[1]
What are the main intents/motivations for researchers to reuse data?
-
[2]
How do IIR researchers discover and obtain reusable research data in their research practices?
-
[3]
How do IIR researchers assess the reusability of research data shared by others?
-
[4]
What are the main concerns that demotivate IIR researchers from reusing others’ data? 4 METHODS Survey and interview represent the two most commonly used methods in examining data sharing and reuse. Survey research enables researchers to identify common practices, behaviors, and perceptions related to data sharing and reuse within a targe t population. It...
-
[5]
Data format and compatibility with your research tools
-
[6]
Data recency and relevance to your research objectives
-
[7]
Documentation and annotations accompanying the dataset 20 JIANG T. ET AL
-
[8]
Imagine you are browsing a data repository with the intention of finding data that could enhanc your research, for what purpose would you try to do the data? 9.2 Codebook TABLE 4 The Codebook Theme Sub-theme Code Purposes and motivations Purposes for reusing others’ data Validate their own data or research results Test applicability of their findings Answ...
Show all 16 references
-
[9]
Data quality and reliability
-
[10]
Metadata completeness and accuracy
-
[11]
Data source credibility and origin
-
[12]
Licensing and permissions for data usage
-
[227]
data thrifting
https://doi.org/10.1007/s00799-015-0157-z Borgman, C. L., & Groth, P. T. (2024). From data creator to data reuser: Distance matters. https://doi.org/10.48550/ARXIV. 2402.07926 Borgman, C. L., Scharnhorst, A., & Golshan, M. S. (2019). Digital data archives as knowledge infrastr...
-
[304]
thick description: Toward an interpretive theory of culture
https://doi.org/10.1145/2467696.2467712 Faniel, I. M., & Jacobsen, T. E. (2010). Reusing scientific data: How earthquake engineering researchers assess the reusability of colleagues’ data. Computer Supported Cooperative Work (CSCW), 19(3), 355–375. https://doi.org/10.1007/s106...
2010
-
[381]
V., Borgman, C
https://doi.org/10.1016/S0306-4573(00)00053-4 Pasquetto, I. V., Borgman, C. L., & Wofford, M. F. (2019). Uses and reuses of scientific data: The data creators’ advantage. Harvard Data Science Review, 1(2). https://doi.org/10.1162/99608f92.fc14bf2d 18 JIANG T. ET AL Pasquetto, ...
2019
-
[2016]
These efforts provided valuable insights and firsthand experiences on the challenges and opportunities associated with sharing and reusing IIR research materials
Tracks, the INEX Interactive Track (iTrack) (Pharo et al., 2010) , Cultural Heritage in CLEF (CHiC) Interactive Task Petras et al., 2013, the Repository of Assigned Search Tasks (RepAST) (Freund & Wildemuth, 2014), and the INEX 2014 Social Book Search Track (Bellot et al., 201...
2019
Reviewed August 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.