Pith. sign in

REVIEW 3 major objections 5 minor 16 references

The Landscape of Data Reuse in Interactive Information Retrieval: Motivations, Sources, and Evaluation of Reusability

T0 review · 3 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Data reuse in interactive information retrieval splits by research orientation: system-oriented researchers reuse external data routinely, user-oriented researchers reuse mostly within trusted circles.

desk verdict First empirical map of IIR data reuse, but the system/user orientation split that drives its main contrast rests on an unvalidated classification; worth reviewing with that caveat. read the letter →

arxiv 2411.15430 v1 pith:V7PDPDM6 submitted 2024-11-23 cs.IR cs.DL

classification cs.IRcs.DL
keywords interactiveinformationretrievaldatareusereusabilityassessmentresearchdiscoveryqualitativeinterviewsinfrastructuresystem-orienteduser-oriented
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to map how researchers in interactive information retrieval (IIR) actually reuse other people's data — why they do it, where they find it, and how they decide it is good enough to use. Based on 21 semi-structured interviews, it argues that data reuse in IIR splits along methodological orientation: system-oriented researchers routinely reuse external datasets for ground-truthing, generalizing findings, and comparability with prior work, while user-oriented researchers reuse data less often, mostly from labmates and close collaborators, and mostly to explore new questions. It also finds that discovery is driven by personal connections and the academic literature rather than by data repositories, and that reusability is judged through four lenses: understandability, trustworthiness, previous usage, and data collection methods. If these patterns hold, efforts to build data-sharing infrastructure for IIR should focus less on generic repositories and more on surfacing data through the channels researchers already trust.

What carries the argument

The analytic engine of the study is the distinction between system-oriented and user-oriented IIR researchers, drawn from the community's own methodological split and used to organize every finding about motivation, discovery, assessment, and concern. The second load-bearing component is the four-aspect account of reusability assessment — understandability (can I make sense of the variables and context?), trustworthiness (is the data reliable and valid?), previous usage (has it been used before, and for what?), and data collection methods (how exactly was it gathered?) — which the authors derive from the interviews and present as the criteria reusers apply. These two components together carry the argument that data reuse in IIR is an interpersonal, context-dependent practice rather than a repository-driven one.

What would settle it

A large-scale survey of IIR researchers, or a log-based study of data repository and lab-website downloads matched against later publication reuse, could test the orientation split: if system-oriented and user-oriented researchers show similar reuse rates and discovery channels when behaviors are observed rather than self-reported, the central contrast would not survive. Alternatively, an ethnographic observation study following researchers through a real data-seeking episode could check whether discovery is as passive and connection-driven as claimed.

Watch

Extended reading notes

Core claim

The central claim is that data reuse in IIR is not a single practice but two different practices shaped by research orientation. System-oriented IIR researchers reuse data often, comfortably drawing on datasets produced by strangers, because reuse gives them ground truth for evaluation, a way to test whether findings generalize across datasets, and comparability with earlier results; for them trust in a dataset can be inherited from the reputation of its producers and from how widely it has already been used. User-oriented researchers reuse data rarely, and almost always from people they know, because their data is deeply tied to specific study designs and contexts; they reuse mainly for exploration and worry more about lost contextual information, ethical issues, and community acceptance. The paper further claims that data discovery is largely passive — through papers, advisors, workshops, and personal contacts — with repositories playing a minor role, and that researchers assess reusability by trying to understand the data, judging its trustworthiness, tracing its previous usages, and scrutinizing how it was collected. The study positions this as an initial empirical map of IIR researchers' data reuse behavior, intended to guide infrastructure and standards for sharing and reusing IIR research data.

Load-bearing premise

The load-bearing premise is that what 21 purposefully and snowball-recruited researchers said about their own behavior — 17 of whom had reused data, mostly from people they knew — accurately reflects how the wider IIR community goes about data reuse.

Editorial extensions

If this is right

  • Data-sharing infrastructure for IIR should prioritize helping reusers find data through publications and trusted contacts, for example by embedding dataset links and provenance records in papers, rather than assuming researchers will actively search general-purpose repositories.
  • Documentation standards that capture study design, variable definitions, and collection procedures would directly lower the main barrier user-oriented researchers report: loss of contextual information.
  • Because system-oriented researchers already inherit trust from prior usage, systematic records of dataset provenance and usage histories would make datasets more reusable across the board.
  • Community-level agreement on data collection and documentation conventions, like the shared confidence TREC datasets enjoy, is a plausible route to expanding reuse beyond personal networks.
  • Efforts to promote a reuse culture in IIR must address perceived innovation and reliability, not only technical access, since reusers worry about how peers and reviewers will judge studies based on other people's data.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial extension: because discovery is reported as passive and connection-driven, a data recommendation layer embedded in the literature — recommending datasets related to the paper being read — could plausibly outperform standalone repository search tools.
  • The trust-in-reputation pattern described in the paper implies a self-reinforcing visibility loop: datasets already used in many papers look more trustworthy, get reused more, and can crowd out equally strong datasets from less-known groups.
  • A quantitative survey or repository-log analysis across the IIR community could test whether the orientation split generalizes beyond this interview sample, since the paper's own limitations section concedes that self-reports may differ from actual practice.
  • The user-oriented half's dependence on tacit context suggests that documentation templates alone may not solve reuse; reusable user-study data may require ongoing involvement of the data creators, such as shared protocol ownership or collaboration norms.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper reports a qualitative interview study of data reuse practices among 21 Interactive Information Retrieval (IIR) researchers. Through semi-structured interviews and two-round thematic coding, the authors identify motivations for reusing data (exploration and ground-truthing), describe how researchers discover and access data (largely through personal connections and publications rather than repositories), and characterize the criteria used to assess reusability (understandability, trustworthiness, previous usage, and data collection methods). The central comparative claim is that system-oriented researchers reuse external data frequently for ground-truthing and comparability, while user-oriented researchers reuse data less often, mostly within trusted networks, and primarily for exploration. The paper concludes with implications for data-sharing infrastructures and community-level standards in IIR.

Significance. If the findings are accepted, the paper offers a useful empirical map of data reuse in a field where reuse has been advocated but little studied. Its strengths include a transparent methodology: the interview guide and codebook are provided in the supplementary materials, the coding procedure is described in Section 4.3, and inter-coder agreement was checked on a sample of transcripts. The authors also explicitly acknowledge limitations such as self-report and small sample size in Section 6.1. However, the central comparative claim about orientation differences rests on an unvalidated grouping of participants, which is a load-bearing internal-validity threat that the limitations section does not address.

major comments (3)
  1. [4.2, Table 1] The assignment of the 21 participants to 'user-oriented' and 'system-oriented' groups is central to the paper's contribution, yet the manuscript gives no operational definition of these categories, no coding procedure for how the labels were applied, and no reliability check for this grouping. Table 1 shows several profiles that are plausibly mixed (e.g., P1 'recommender systems, metrics for IR system evaluation'; P4 'task-oriented conversational systems'; P17 'user behavior analysis, retrieval models'). Because Sections 5.1–5.4 and the Discussion contrast these two groups at every turn, the comparative findings would be an artifact of an unreliable split. The authors should either provide a transparent rubric and demonstrate inter-coder agreement for the orientation labels, or re-frame the findings as within-sample patterns without the group contrast.
  2. [9.1 / 5.3] The interview guide in Section 9.1 asks participants to consider 'proposed aspects' including data quality, metadata completeness, source credibility, licensing, format, recency, and documentation when evaluating reusability. This priming is likely to shape the responses that underlie the reusability assessment findings in Section 5.3. The paper does not acknowledge this as a methodological limitation, and Section 6.1's limitations list does not mention it. The authors should discuss the potential priming effect and temper the strength of the claims in 5.3.
  3. [5.1–5.4] The results report orientation-based differences without indicating how many of the 9 system-oriented and 12 user-oriented participants expressed each pattern. For example, Section 5.1 states that user-oriented participants primarily reused data for exploration and rarely for ground-truthing, but no counts or per-group tallies are provided. With N=21, providing the distribution of responses per theme across the two groups would materially strengthen the credibility of the comparative claims and allow readers to assess the degree of overlap.
minor comments (5)
  1. [5.1] The phrase 'It's worth nothing' should read 'It's worth noting' (Section 5.1, discussing the two main purposes of reuse).
  2. [5.4.1] The word 'regrading' should be 'regarding' (Section 5.4.1, first sentence: 'Challenges in understanding other's data were a major concern for participants regrading data reuse').
  3. [2.3] The sentence 'Researchers have also studied resource reuse in IIR and attempted and madetheoretical contributions and provide suggestions future practices' is garbled; it should be rephrased, e.g., 'Researchers have also studied resource reuse in IIR and attempted to make theoretical contributions and provide suggestions for future practices.'
  4. [2.1] The citation 'Bishop (Bishop, 2009)' redundantly repeats the author name inside the parenthetical; it should be 'Bishop (2009)' or '(Bishop, 2009)'.
  5. [4.3] Cohen's kappa values above 0.6 are described as 'substantial agreement,' but 0.61–0.80 is typically labeled substantial; the authors may want to report the exact kappa value and interpret it accordingly rather than using 'exceeded 0.6.'

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: the interview findings are grounded in participant data, with only minor non-load-bearing self-citation.

full rationale

This paper is an interview-based qualitative study; it contains no equations, fitted parameters, or quantitative derivation chain whose outputs could reduce to its inputs. The central comparative claim—that system-oriented IIR researchers reuse external data frequently for ground-truthing/comparability while user-oriented researchers reuse less, mostly within trusted networks and for exploration—is presented as an empirical result of coding 21 semi-structured interviews, not as a logical consequence of the orientation labels used in recruitment. Participants in Section 4.2 were recruited by research orientation (system-oriented vs. user-oriented), but the reuse behaviors, motivations, discovery channels, reusability criteria, and concerns reported in Sections 5.1–5.4 were elicited and coded from transcripts, so the outcome categories are not definitionally entailed by the input categories. The preliminary codebook in Section 4.3 was informed by prior frameworks including Liu (2022), a self-citation by the corresponding author; however, the paper explicitly states that new themes and sub-themes were allowed to emerge and that inter-coder agreement was checked, and the findings are not derived from that framework by construction. The one Results passage citing Liu (2022) for user-oriented validity assessment is corroborative rather than load-bearing, and the same section quotes participant reasoning. The skeptic's concern that the orientation split is not validated is a legitimate internal-validity limitation, but it is not a circularity: nothing in the definition of 'system-oriented' or 'user-oriented' analytically forces the observed differences in reuse frequency, motivation, or concern. Overall, the derivation chain is self-contained and the reported patterns are grounded in the interview data, so no circular step is present; the only mild caveat is the presence of self-citations, which are not load-bearing here.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters or invented entities are present in this qualitative study. The central claims rest on domain assumptions about self-report accuracy, sample representativeness, and coding reliability, all of which the authors partially acknowledge in the Limitations section.

assumptions (3)
  • domain assumption Participants' self-reported accounts reflect their actual data reuse practices.
    The study relies on interview self-reports; authors acknowledge in Limitations (Section 6.1) that reports may differ from actual practices.
  • domain assumption The purposeful plus snowball sample represents the diversity of IIR researchers.
    Recruitment through professional network and snowball sampling (Section 4.2) may over-represent connected researchers; the paper claims theoretical saturation but provides no saturation evidence.
  • domain assumption Cohen's kappa above 0.6 indicates reliable coding.
    Section 4.3 reports kappa exceeded 0.6 on a randomly selected sample of 3 transcripts, interpreted as substantial agreement; this is a conventional threshold but only checked on a small subset.

how reviews work

0 comments
Cite this review

Pith. "Pith review of The Landscape of Data Reuse in Interactive Information Retrieval: Motivations, Sources, and Evaluation of Reusability." pith.science (2026). https://pith.science/paper/V7PDPDM6

@misc{pith2026241115430,
  author       = {Pith},
  title        = {Pith review of: The Landscape of Data Reuse in Interactive Information Retrieval: Motivations, Sources, and Evaluation of Reusability},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V7PDPDM6}},
  note         = {Machine review of arXiv:2411.15430}
}
read the original abstract

Sharing and reusing research data can effectively reduce redundant efforts in data collection and curation, especially for small labs and research teams conducting human-centered system research, and enhance the replicability of evaluation experiments. Building a sustainable data reuse process and culture relies on frameworks that encompass policies, standards, roles, and responsibilities, all of which must address the diverse needs of data providers, curators, and reusers. To advance the knowledge and accumulate empirical understandings on data reuse, this study investigated the data reuse practices of experienced researchers from the area of Interactive Information Retrieval (IIR) studies, where data reuse has been strongly advocated but still remains a challenge. To enhance the knowledge on data reuse behavior and reusability assessment strategies within IIR community, we conducted 21 semi-structured in-depth interviews with IIR researchers from varying demographic backgrounds, institutions, and stages of careers on their motivations, experiences, and concerns over data reuse. We uncovered the reasons, strategies of reusability assessments, and challenges faced by data reusers within the field of IIR as they attempt to reuse researcher data in their studies. The empirical finding improves our understanding of researchers' motivations for reusing data, their approaches to discovering reusable research data, as well as their concerns and criteria for assessing data reusability, and also enriches the on-going discussions on evaluating user-generated data and research resources and promoting community-level data reuse culture and standards.

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

16 extracted references · 14 canonical work pages

  1. [1]

    What are the main intents/motivations for researchers to reuse data?

  2. [2]

    How do IIR researchers discover and obtain reusable research data in their research practices?

  3. [3]

    How do IIR researchers assess the reusability of research data shared by others?

  4. [4]

    reusability

    What are the main concerns that demotivate IIR researchers from reusing others’ data? 4 METHODS Survey and interview represent the two most commonly used methods in examining data sharing and reuse. Survey research enables researchers to identify common practices, behaviors, and perceptions related to data sharing and reuse within a targe t population. It...

  5. [5]

    Data format and compatibility with your research tools

  6. [6]

    Data recency and relevance to your research objectives

  7. [7]

    Documentation and annotations accompanying the dataset 20 JIANG T. ET AL

  8. [8]

    Imagine you are browsing a data repository with the intention of finding data that could enhanc your research, for what purpose would you try to do the data? 9.2 Codebook TABLE 4 The Codebook Theme Sub-theme Code Purposes and motivations Purposes for reusing others’ data Validate their own data or research results Test applicability of their findings Answ...

Show all 16 references
  1. [9]

    Data quality and reliability

  2. [10]

    Metadata completeness and accuracy

  3. [11]

    Data source credibility and origin

  4. [12]

    Licensing and permissions for data usage

  5. [227]

    data thrifting

    https://doi.org/10.1007/s00799-015-0157-z Borgman, C. L., & Groth, P. T. (2024). From data creator to data reuser: Distance matters. https://doi.org/10.48550/ARXIV. 2402.07926 Borgman, C. L., Scharnhorst, A., & Golshan, M. S. (2019). Digital data archives as knowledge infrastr...

  6. [304]

    thick description: Toward an interpretive theory of culture

    https://doi.org/10.1145/2467696.2467712 Faniel, I. M., & Jacobsen, T. E. (2010). Reusing scientific data: How earthquake engineering researchers assess the reusability of colleagues’ data. Computer Supported Cooperative Work (CSCW), 19(3), 355–375. https://doi.org/10.1007/s106...

  7. [381]

    V., Borgman, C

    https://doi.org/10.1016/S0306-4573(00)00053-4 Pasquetto, I. V., Borgman, C. L., & Wofford, M. F. (2019). Uses and reuses of scientific data: The data creators’ advantage. Harvard Data Science Review, 1(2). https://doi.org/10.1162/99608f92.fc14bf2d 18 JIANG T. ET AL Pasquetto, ...

  8. [2016]

    These efforts provided valuable insights and firsthand experiences on the challenges and opportunities associated with sharing and reusing IIR research materials

    Tracks, the INEX Interactive Track (iTrack) (Pharo et al., 2010) , Cultural Heritage in CLEF (CHiC) Interactive Task Petras et al., 2013, the Repository of Assigned Search Tasks (RepAST) (Freund & Wildemuth, 2014), and the INEX 2014 Social Book Search Track (Bellot et al., 201...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.