Pith. sign in

REVIEW 3 major objections 4 minor 101 references

Measuring the State of Open Science in Transportation Using Large Language Models

T0 review · 3 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read After an LLM sweep of 10,724 transportation papers, only 5% share code and 4% share data, and sharing earns no citation or review-time reward.

desk verdict First large-scale LLM-based baseline of code/data sharing in transportation; rates are directionally credible but the 96-paper validation and unpropagated measurement error make the exact headline numbers softer than they look. read the letter →

arxiv 2601.14429 v2 pith:GR4XEEH5 submitted 2026-01-20 cs.DL cs.AIcs.CYcs.ET

classification cs.DLcs.AIcs.CYcs.ET
keywords opensciencemonitoringtransportationresearchlargelanguagemodelsdataavailabilitycodereproducibilityinter-rateragreement
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that the state of open science in transportation research can be measured automatically, at scale, and reliably, using large language models to read full-text articles. Analyzing 10,724 papers from the Transportation Research Part journals (2019–2024), it reports that only about 5% of quantitative papers make code available and 4% make data available in a repository, with about 3% doing both. It also finds that sharing is uneven across journals, topics, and world regions, and that it brings no measurable benefit in citations or review speed—code-sharing papers actually took 2 to 27 days longer to review. The authors conclude that journals and funding agencies, not author self-interest, must drive structural change, and that the pipeline can serve as a repeatable monitor for the field.

What carries the argument

The key machinery is an automated three-stage pipeline: it parses the journals' structured XML full texts into markdown, prompts a large language model with separate single-task prompts to extract availability flags (code used, code publicly available, data cited, data repository available, quantitative study), and validates those flags against two human annotators on a 96-paper manual set using inter-rater agreement metrics (kappa). Only features meeting agreement thresholds are retained, and link-checking verifies that stated repository URLs are live and contain the claimed artifact. The validated flags feed two logit choice models that estimate how journal, region, topic, and paper charac

What would settle it

Re-annotate a fresh random sample of roughly 300 papers from the same corpus with two human annotators using the paper's own definitions, then compare the human code/data repository rates with the pipeline's rates on those same papers; the central measurement collapses if the disagreement is large enough to move the 5%/4% figures—especially for 'data repository available,' which already showed only Fair inter-rater agreement on the 96-paper validation set.

Watch

Extended reading notes

Core claim

The central discovery is a measured baseline for open-science practice in a flagship journal series. Using an LLM pipeline validated against two human annotators on a 96-paper manual set, the authors estimate that of 10,480 quantitative research articles, 5% share a code repository, 4% share a data repository, and about 3% share both, while 29% cite or link existing public datasets and only a small fraction contribute new data. Sharing concentrates in particular journals (TR-B and TR-C), in data-driven modeling topics, and among corresponding authors outside Asia, and it increases for newer papers. The paper's headline negative result is that papers sharing code or data do not receive more c

Load-bearing premise

The load-bearing premise is that the LLM's agreement with two human annotators on 96 papers—Almost Perfect for code availability but only Fair for data-repository availability—transfers intact to the full 10,724-paper corpus, so the extracted features can be treated as ground truth; if that transfer fails, the 5% and 4% rates could shift materially.

Editorial extensions

If this is right

  • The field now has a repeatable baseline: future audits can rerun the pipeline to track whether open-science rates rise over time.
  • Because sharing brings no citation or review-time benefit, journal policies—structured availability statements, open-science checklists, badges—and funder requirements are the levers most likely to move the rates.
  • At 3–5% availability, most transportation papers fail a necessary condition for computational reproducibility, since sharing data and code is a prerequisite for others to verify results.
  • The large gap between citing others' data (29%) and sharing new data (<4%) implies that a handful of datasets supports a large share of the field's empirical work, so publishing new open datasets could have outsized impact.
  • The pipeline is transferable to other journals in the same publisher and to other publishers with text-mining access, allowing cross-field comparison of open-science practice.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Because the pipeline only counts repositories explicitly mentioned in the full text, the 5%/4% figures are lower bounds under the paper's scope; adding an external web search for shared artifacts would likely raise the measured rates.
  • The no-citation/no-review-time result is purely observational; isolating the causal effect of sharing would require a quasi-experimental design, such as comparing citation trajectories before and after a journal's open-science policy change.
  • The 96-paper validation set and the Fair inter-rater agreement on 'data repository available' mean the headline rates carry nontrivial measurement error; a larger, stratified validation sample would be needed to detect year-over-year trend changes with confidence.
  • The paper's choice-model findings suggest a testable extension: apply the same pipeline to the journals' policy changes (e.g., a journal that begins requiring availability statements) to see whether mandated statements, rather than soft encouragement, shift the 5%/4% figures.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper presents an LLM-based pipeline for extracting code and data availability indicators from full-text articles, applies it to 10,724 Transportation Research Part A–F and TR-IP articles published 2019–2024, and validates the extraction on a 96-paper manual validation dataset (MVD) via inter-rater agreement analysis. The headline findings are that about 5% of quantitative papers share a code repository, 4% share a data repository, and about 3% share both; that sharing varies by journal, topic, and corresponding-author region; and that sharing is not associated with higher citation counts or shorter review times. The paper also provides topic models, logit choice models, and an interactive explorer.

Significance. If the measurement is reliable, this is a valuable first large-scale, field-wide snapshot of open science practices in transportation research. The pipeline is reproducible and the authors make code and permissible data available; the comparison against text-search baselines (Appendix C) and the explicit inter-rater agreement analysis are commendable and go beyond what is common in LLM-as-annotator studies. The finding that data/code availability is uncorrelated with citation outcomes, if confirmed with proper inference, has clear policy relevance for journals and funders. However, the central claim rests critically on the validity of the LLM-extracted features at scale, and the evidence for that validity is the agreement on only 96 papers—with the data-repository feature showing only Fair inter-rater agreement (Fleiss κ=0.399, H1–H2 Cohen κ=0.333). This uncertainty is not propagated into the headline rates or the choice models, which is the main load-bearing weakness.

major comments (3)
  1. [§4.4.2, Table 4; §5.1] The headline 4% data-repository rate is based on an LLM feature whose agreement is weak: Fleiss κ=0.399 (Fair) and, crucially, human-human Cohen κ=0.333 (Table 4). Table 3 shows prevalence differences that are material at this rate: H1=8.33%, H2=3.12%, LLM=3.12%. With a 96-paper validation set, the LLM's apparent agreement with H2 on this feature (Cohen κ=0.656) could reflect rater-specific bias rather than accuracy. The paper acknowledges this ('retained with caution') but does not quantify the impact on the 5%/4%/3% estimates. Please report the MVD confusion matrix for this feature, compute misclassification-corrected prevalence estimates and confidence intervals (e.g., via bootstrap or a latent-class model), and re-run the §5 choice models with measurement-error propagation or at least a sensitivity analysis. Without this, the exact headline rates are not statistically grounded, even
  2. [§4.5, Appendix D, Appendix E] The pipeline is described as LLM-based and validated against the 96-paper MVD, but the postprocessing includes steps outside that validation: Appendix E lists 'uses_data_bool' as 'after manual correction,' and Appendix D excludes records with 'irreconcilable inconsistencies' (10,724 → 10,480) without reporting how many papers or which flags were affected. If manual correction changed availability flags, or if the excluded papers are non-random with respect to availability, the agreement statistics in §4.4 do not cover the final analysis dataset. Please quantify: how many papers were manually corrected, how many were excluded, and whether any headline rates change materially when the corrected/excluded cases are handled differently.
  3. [§5.4] The abstract and conclusions state that there is 'no significant difference in citation counts or review duration between papers that provided data and code and those that did not.' The support in §5.4 is descriptive: LOESS curves in Figure 4 and a comparison of mean review times (269 vs. 254 days, CI 2–27 days). No regression or hypothesis test is reported for citations that controls for paper age, journal, topic, or other confounders, and the figure only shows visual overlap. The review-time claim is also based on a single unadjusted comparison. Please provide a formal statistical analysis (e.g., a citation-count model with controls, and a regression of review time on availability indicators including journal fixed effects) before making this incentive-gap claim. This is load-bearing for the policy recommendation that current incentives do not reward sharing.
minor comments (4)
  1. [§4.4.1] Typo: 'We also calculate' should be 'We also calculated.'
  2. [§9] The statement 'Given that some of these repositories contain repackaged datasets, less than 4% of the papers we reviewed contributed new data to the transportation community' is presented as a conclusion, but the paper does not systematically classify whether shared repositories contain new vs. repackaged data. Please mark this as an interpretive inference and support it with evidence or soften the claim.
  3. [§6.2] The limitation that temperature was not set to 0 during extraction is acknowledged, but the practical impact on reproducibility is not discussed. Since the code and data are to be released, please also release the exact prompt versions and model snapshot identifiers, and consider stating whether rerunning with temperature 0 changes any of the headline rates.
  4. [Table 3 and Table 4] The caption for Table 3 says 'Prev. (H1), Prev. (H2), Prev. (LLM)' but the column headers show 'Prev.' only; please align headers with the caption for clarity.

Circularity Check

0 steps flagged · score 1.0 of 10

No significant circularity: LLM measurement is validated against independent human annotations; headline rates are outputs, not fitted inputs.

full rationale

The paper's derivation chain is a measurement pipeline rather than a fitted-then-predicted loop. Section 2 fixes operational definitions for code/data availability; Section 4.3 uses Gemini to extract features; Section 4.4 validates those extractions against a 96-paper manual validation set with two independent human annotators, reporting Fleiss/Cohen kappa values that are external agreement benchmarks, not quantities derived from the final rates. The headline 5%/4%/3% statistics in §5.1 and the choice models in §5.2–5.3 are computed from the extracted features; no equation is algebraically equivalent to an input, and no parameter is fitted to the headline rates and then renamed as a prediction. The logit models estimate associations from the extracted labels rather than assuming them. Concerns raised in the text—small MVD size, Fair agreement on is_data_repository_available (Fleiss κ=0.399), the authors' own admission that LLMs are imperfect and that temperature was not 0, and the §5.1 note that results are 'subject to estimation error'—are measurement-validity limitations, not circularity. The self-citations to RERITE tutorials and community efforts ([50], [95], [97]) are contextual and are not load-bearing for the central measurement. The Appendix C comparison against text search and the independent human annotators provide an external benchmark. Thus no circular step is exhibited.

Assumptions & free parameters 3 free parameters · 5 assumptions · 0 invented entities

The central measurement rests on the operational definition of availability (text-stated links), the representativeness of the 96-paper validation set, and the LDA/logit modeling choices. No new physical or conceptual entities are invented; the listed items are domain assumptions and fitted model parameters.

free parameters (3)
  • Number of LDA topics (K=15) = 15
    Chosen by coherence score on the corpus; affects topic assignments used as covariates in the choice models (Section 4.2, Appendix F).
  • Logit model variable groupings = e.g., TRB+TRC combined; Europe+SouthAmerica+Africa combined
    Categories combined/excluded based on statistical significance in the same data (Sections 5.2-5.3), so the factor claims are fit to the data.
  • LLM temperature = non-zero (not specified)
    Acknowledged in Section 6.2; affects exact extraction outputs and reproducibility of the pipeline.
assumptions (5)
  • domain assumption The 96-paper manual validation set is representative of the full 10,724-paper corpus, so LLM agreement measured on it transfers to the whole corpus.
    Used to justify applying LLM-extracted features to all papers without full manual verification (Section 4.4, Section 6.2).
  • domain assumption Data/code availability stated in the full text (presence of links) is an appropriate operationalization of open science practice; absence of a link indicates unavailability (no external search).
    Defined in Sections 2 and 3.1; drives all downstream features and the headline rates.
  • standard math i.i.d. extreme-value (Gumbel) errors for the binary/multinomial logit models.
    Equations (1)-(8) in Sections 5.2-5.3; standard MNL assumption.
  • domain assumption LDA topic model with K=15 fitted via gensim provides meaningful topic assignments for papers.
    Section 4.2, Appendix F; topics used as covariates; acknowledged misclassification of COVID topic in Section 6.2.
  • domain assumption Citation counts from SCOPUS and review time computed from XML dates are accurate measures of academic reward.
    Section 5.4 incentive analysis.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Measuring the State of Open Science in Transportation Using Large Language Models." pith.science (2026). https://pith.science/paper/GR4XEEH5

@misc{pith2026260114429,
  author       = {Pith},
  title        = {Pith review of: Measuring the State of Open Science in Transportation Using Large Language Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GR4XEEH5}},
  note         = {Machine review of arXiv:2601.14429}
}
read the original abstract

Open science initiatives have strengthened scientific integrity and accelerated research progress across many fields, but the state of their practice within transportation research remains under-investigated. Key features of open science, defined here as data and code availability, are difficult to extract due to the inherent complexity of the field. Previous work has either been limited to small-scale studies due to the labor-intensive nature of manual analysis or has relied on large-scale bibliometric approaches that sacrifice contextual richness. This paper introduces an automatic and scalable feature-extraction pipeline to measure code and data availability in transportation research. We employ Large Language Models (LLMs) for this task and validate their performance against a manually curated dataset and through an inter-rater agreement analysis. We applied this pipeline to examine 10,724 research articles published in the Transportation Research Part series of journals between 2019 and 2024. Our analysis found that only 5% of quantitative papers shared a code repository, 4% of quantitative papers shared a data repository, and about 3% of papers shared both, with trends differing across journals, topics, and geographic regions. We found no significant difference in citation counts or review duration between papers that provided data and code and those that did not, suggesting a misalignment between open science efforts and traditional academic metrics. Consequently, encouraging these practices will likely require structural interventions from journals and funding agencies to supplement the lack of direct author incentives. The pipeline developed in this study can be readily scaled to other journals, representing a critical step toward the automated measurement and monitoring of open science practices in transportation research.

Figures

Figures reproduced from arXiv: 2601.14429 by the authors.

Figure 1
Figure 1. Features considered for data and code availabilities. [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗
Figure 2
Figure 2. Number of articles published in Transportation Research journals (TR-A, TR-B, TR-C, TR-D, TR-E, TR-F, and TR-IP) [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Overview of the features extracted from our pipeline [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Citations per year of each paper over time by code and data-sharing practice, with a LOESS moving average line [PITH_FULL_IMAGE:figures/full_fig_p016_4.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

101 extracted references · 1 canonical work pages

  1. [1]

    Reproducible research in computational science,

    R. D. Peng, “Reproducible research in computational science,”Science, vol. 334, no. 6060, pp. 1226– 1227, 2011

  2. [2]

    Estimating the reproducibility of psychological science,

    O. S. Collaboration, “Estimating the reproducibility of psychological science,”Science, vol. 349, no. 6251, p. aac4716, 2015

  3. [3]

    The reproducible research movement in statistics,

    V . Stodden, “The reproducible research movement in statistics,”Statistical Journal of the IAOS, vol. 30, no. 2, pp. 91–93, 2014

  4. [4]

    Guide for authors - Transportation Research Part C: Emerging technologies

    “Guide for authors - Transportation Research Part C: Emerging technologies.”https://www. sciencedirect.com/journal/transportation-research-part-c-emerging-technologies/ publish/guide-for-authors, 2025. Accessed: 2025-06-16

  5. [5]

    Aims & scope - Transportation Research Part C: Emerging technologies

    “Aims & scope - Transportation Research Part C: Emerging technologies.”https://www. sciencedirect.com/journal/transportation-research-part-c-emerging-technologies/ about/aims-and-scope, 2025. Accessed: 2025-06-16

  6. [6]

    Open science is a research accelerator,

    M. Woelfle, P . Olliaro, and M. H. Todd, “Open science is a research accelerator,”Nature chemistry, vol. 3, no. 10, pp. 745–748, 2011

  7. [7]

    Understanding open science

    UNESCO, “Understanding open science.”https://doi.org/10.54677/UTCD9302, 2022. Document code: SC-PBS-STIP/2022/OST/1, 6 pages

  8. [8]

    Equity, transparency, and accountability: open science for the 21st century,

    M. A. Winker, T. Bloom, S. Onie, and J. Tumwine, “Equity, transparency, and accountability: open science for the 21st century,”The Lancet, vol. 402, p. 1206–1209, Oct. 2023

Show all 101 references
  1. [9]

    Why nasa and federal agencies are declaring this the year of open science,

    C. Gentemann, “Why nasa and federal agencies are declaring this the year of open science,”Nature, vol. 613, no. 7943, p. 217, 2023

  2. [10]

    The principles of open science monitoring,

    E. Bobrov, L. Bracco, M. Dacos, N. Fressengeas, I. Hrynaszkiewicz, A. Iarkaeva, A. Per ˇsi´c, V . Proud- man, L. Romary, and R. Sabo, “The principles of open science monitoring,” July 2025

  3. [11]

    N. A. of Sciences, Medicine, Policy, G. Affairs, B. on Research Data, Information, D. on Engineering, P . Sciences, C. on Applied, T. Statistics,et al.,Reproducibility and replicability in science. National Academies Press, 2019

  4. [12]

    Seven easy steps to open science,

    S. Cr ¨uwell, J. van Doorn, A. Etz, M. C. Makel, H. Moshontz, J. C. Niebaum, A. Orben, S. Parsons, and M. Schulte-Mecklenbeck, “Seven easy steps to open science,”Zeitschrift f¨ ur Psychologie, 2019

  5. [13]

    The unreasonable effectiveness of open science in ai: A replication study,

    O. E. Gundersen, O. Cappelen, M. Møln ˚a, and N. G. Nilsen, “The unreasonable effectiveness of open science in ai: A replication study,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 39, pp. 26211–26219, 2025

  6. [14]

    Assess- ing data availability and research reproducibility in hydrology and water resources,

    J. H. Stagge, D. E. Rosenberg, A. M. Abdallah, H. Akbar, N. A. Attallah, and R. James, “Assess- ing data availability and research reproducibility in hydrology and water resources,”Scientific data, vol. 6, no. 1, pp. 1–12, 2019

  7. [15]

    Estimating the deep replicability of scientific findings using human and artificial intelligence,

    Y. Yang, W. Youyou, and B. Uzzi, “Estimating the deep replicability of scientific findings using human and artificial intelligence,”Proceedings of the National Academy of Sciences, vol. 117, no. 20, pp. 10762–10768, 2020

  8. [16]

    A discipline-wide investigation of the replicability of psychology papers over the past two decades,

    W. Youyou, Y. Yang, and B. Uzzi, “A discipline-wide investigation of the replicability of psychology papers over the past two decades,”Proceedings of the National Academy of Sciences, vol. 120, no. 6, p. e2208863120, 2023. 29 (a) Optimization (b) Mobility policy (c) Driving be...

  9. [17]

    Improving reproducibility in machine learning research (a report from the neurips 2019 reproducibility program),

    J. Pineau, P . Vincent-Lamarre, K. Sinha, V . Larivi `ere, A. Beygelzimer, F. d’Alch ´e Buc, E. Fox, and H. Larochelle, “Improving reproducibility in machine learning research (a report from the neurips 2019 reproducibility program),”Journal of machine learning research, vol. ...

  10. [18]

    A computational reproducibility study of plos one articles featuring longitudinal data analyses,

    H. Seibold, S. Czerny, S. Decke, R. Dieterle, T. Eder, S. Fohr, N. Hahn, R. Hartmann, C. Heindl, P . Kopper,et al., “A computational reproducibility study of plos one articles featuring longitudinal data analyses,”PLoS One, vol. 16, no. 6, p. e0251194, 2021

  11. [19]

    ”get in researchers; we’re measuring reproducibility

    D. Olszewski, A. Lu, C. Stillman, K. Warren, C. Kitroser, A. Pascual, D. Ukirde, K. Butler, and P . Traynor, “”get in researchers; we’re measuring reproducibility”: A reproducibility study of ma- chine learning papers in tier 1 security conferences,” inProceedings of the 2023 ...

  12. [20]

    Revisiting reproducibility in transportation simulation studies,

    K. Riehl, A. Kouvelas, and M. A. Makridis, “Revisiting reproducibility in transportation simulation studies,”European Transport Research Review, vol. 17, no. 1, p. 22, 2025

  13. [21]

    Next generation simula- tion (ngsim) vehicle trajectories and supporting data. [dataset]. provided by its datahub through data.transportation.gov,

    U.S. Department of Transportation Federal Highway Administration, “Next generation simula- tion (ngsim) vehicle trajectories and supporting data. [dataset]. provided by its datahub through data.transportation.gov,” 2016

  14. [22]

    Open science now: A systematic literature review for an integrated definition,

    R. Vicente-Saez and C. Martinez-Fuentes, “Open science now: A systematic literature review for an integrated definition,”Journal of business research, vol. 88, pp. 428–436, 2018

  15. [23]

    Research data guidelines,

    Elsevier, “Research data guidelines,” 2025. Accessed: 2025-12-23

  16. [24]

    Reasons, challenges, and some tools for doing reproducible transportation research,

    Z. Zheng, “Reasons, challenges, and some tools for doing reproducible transportation research,” Communications in Transportation Research, vol. 1, p. 100004, 2021

  17. [25]

    Journals Data XML

    Elsevier, “Journals Data XML.”https://supportcontent.elsevier.com/Support%20Hub/DaaS/ Journals_Data_XML.pdf, Nov. 2024. Data as a Service documentation. Accessed 17 Jun 2025

  18. [26]

    Latent Dirichlet allocation,

    D. M. Blei, A. Y. Ng, and M. I. Jordan, “Latent Dirichlet allocation,”J. Mach. Learn. Res., vol. 3, no. 4-5, pp. 993–1022, 2003

  19. [27]

    Discovering themes and trends in transportation research using topic modeling,

    L. Sun and Y. Yin, “Discovering themes and trends in transportation research using topic modeling,” Transportation Research Part C: Emerging Technologies, vol. 77, pp. 49–66, 2017

  20. [28]

    Identifying regional characteristics of transportation research with transport research international documentation (trid) data,

    Y. Sun and S. Kirtonia, “Identifying regional characteristics of transportation research with transport research international documentation (trid) data,”Transportation Research Part A: Policy and Practice, vol. 137, pp. 111–130, 2020

  21. [29]

    Visualizing data using t-sne,

    L. v. d. Maaten and G. Hinton, “Visualizing data using t-sne,”Journal of machine learning research, vol. 9, no. Nov, pp. 2579–2605, 2008

  22. [30]

    Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,

    G. Team, “Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities,” tech. rep., Google, 2025. Accessed: 2025-06-20

  23. [31]

    Sci- assess: Benchmarking llm proficiency in scientific literature analysis,

    H. Cai, X. Cai, J. Chang, S. Li, L. Yao, W. Changxin, Z. Gao, H. Wang, L. Yongge, M. Lin,et al., “Sci- assess: Benchmarking llm proficiency in scientific literature analysis,” inFindings of the Association for Computational Linguistics: NAACL 2025, pp. 2335–2357, 2025

  24. [32]

    Multi-task inference: Can large language models follow multiple instructions at once?,

    G. Son, S. Baek, S. Nam, I. Jeong, and S. Kim, “Multi-task inference: Can large language models follow multiple instructions at once?,” inProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 5606–5627, 2024

  25. [33]

    Comparative analysis of prompt strategies for large language models: Single-task vs. multitask prompts,

    M. Gozzi and F. Di Maio, “Comparative analysis of prompt strategies for large language models: Single-task vs. multitask prompts,”Electronics, vol. 13, no. 23, p. 4712, 2024. 31

  26. [34]

    Interrater reliability: the kappa statistic,

    M. L. McHugh, “Interrater reliability: the kappa statistic,”Biochemia medica, vol. 22, no. 3, pp. 276– 282, 2012

  27. [35]

    Measuring nominal scale agreement among many raters,

    J. L. Fleiss, “Measuring nominal scale agreement among many raters,”Psychological Bulletin, vol. 76, no. 5, pp. 378–382, 1971

  28. [36]

    The measurement of observer agreement for categorical data,

    J. R. Landis and G. G. Koch, “The measurement of observer agreement for categorical data,”Biomet- rics, vol. 33, no. 1, pp. 159–174, 1977

  29. [37]

    A coefficient of agreement for nominal scales,

    J. Cohen, “A coefficient of agreement for nominal scales,”Educational and psychological measurement, vol. 20, no. 1, pp. 37–46, 1960

  30. [38]

    A short introduction to biogeme,

    M. Bierlaire, “A short introduction to biogeme,” Tech. Rep. TRANSP-OR 230620, Transport and Mo- bility Laboratory, ´Ecole Polytechnique F´ed´erale de Lausanne, 2023. Version prepared using Biogeme 3.2.11 (April 20, 2023)

  31. [39]

    Engaging new partners in transportation research: Integrating the publishing, archiving, and indexing of technical literature into the research process,

    M. P . Newton, D. M. Bullock, C. Watkinson, P . J. Bracke, and D. K. Horton, “Engaging new partners in transportation research: Integrating the publishing, archiving, and indexing of technical literature into the research process,”Transportation research record, vol. 2291, no....

  32. [40]

    Evolution of data creation, management, publication, and curation in the research process,

    L. D. Zilinski, D. A. Scherer Jr, D. M. Bullock, D. Horton, and C. E. Matthews, “Evolution of data creation, management, publication, and curation in the research process,”Transportation Research Record, vol. 2414, no. 1, pp. 9–19, 2014

  33. [41]

    Using the future wheel methodology to assess the impact of open science in the transport sector,

    A. F. Nielsen, J. Michelmann, A. Akac, K. Palts, A. Zilles, A. Anagnostopoulou, and O. Langeland, “Using the future wheel methodology to assess the impact of open science in the transport sector,” Scientific reports, vol. 13, no. 1, p. 6000, 2023

  34. [42]

    Data sharing of transport research data,

    H. Gellerman, E. Svanberg, and Y. Barnard, “Data sharing of transport research data,”Transportation Research Procedia, vol. 14, pp. 2227–2236, 2016

  35. [43]

    Udrive: the european naturalistic driving study,

    R. Eenink, Y. Barnard, M. Baumann, X. Augros, and F. Utesch, “Udrive: the european naturalistic driving study,” inProceedings of Transport Research Arena, IFSTTAR, 2014

  36. [44]

    Design of open source framework for traffic and travel simulation,

    G. Tamminga, M. Miska, E. Santos, H. Van Lint, A. Nakasone, H. Prendinger, and S. Hoogendoorn, “Design of open source framework for traffic and travel simulation,”Transportation research record, vol. 2291, no. 1, pp. 44–52, 2012

  37. [45]

    Open traffic: a toolbox for traffic research,

    G. Tamminga, P . Knoppers, and J. Van Lint, “Open traffic: a toolbox for traffic research,”Procedia Computer Science, vol. 32, pp. 788–795, 2014

  38. [46]

    About calibration of car-following dynamics of automated and human-driven vehicles: Methodology, guidelines and codes,

    V . Punzo, Z. Zheng, and M. Montanino, “About calibration of car-following dynamics of automated and human-driven vehicles: Methodology, guidelines and codes,”Transportation Research Part C: Emerging Technologies, vol. 128, p. 103165, 2021

  39. [47]

    Enabling reproducible research in sensor-based transportation mode recognition with the sussex-huawei dataset,

    L. Wang, H. Gjoreski, M. Ciliberto, S. Mekki, S. Valentin, and D. Roggen, “Enabling reproducible research in sensor-based transportation mode recognition with the sussex-huawei dataset,”IEEE Access, vol. 7, pp. 10870–10891, 2019

  40. [48]

    Comparing hundreds of machine learning and dis- crete choice models for travel demand modeling: An empirical benchmark,

    S. Wang, B. Mo, Y. Zheng, S. Hess, and J. Zhao, “Comparing hundreds of machine learning and dis- crete choice models for travel demand modeling: An empirical benchmark,”Transportation Research Part B: Methodological, vol. 190, p. 103061, 2024

  41. [49]

    Sharing, collaborating, and bench- marking to advance travel demand research: A demonstration of short-term ridership prediction,

    J. D. Caicedo, C. Guirado, M. C. Gonz ´alez, and J. L. Walker, “Sharing, collaborating, and bench- marking to advance travel demand research: A demonstration of short-term ridership prediction,” Transport Policy, 2025

  42. [50]

    Reproducibility in transportation research: A hands-on tutorial,

    C. Wu, B. Ghosh, Z. Zheng, and I. Mart´ınez, “Reproducibility in transportation research: A hands-on tutorial,” 2024. 32

  43. [51]

    Freeway performance measurement system: operational analysis tool,

    T. Choe, A. Skabardonis, and P . Varaiya, “Freeway performance measurement system: operational analysis tool,”Transportation research record, vol. 1811, no. 1, pp. 67–75, 2002

  44. [52]

    Transportation networks for research

    T. N. for Research Core Team, “Transportation networks for research.”https://github.com/ bstabler/TransportationNetworks, 2025. Accessed: November 13, 2025

  45. [53]

    The acceptance of modal innovation: The case of swiss- metro,

    M. Bierlaire, K. Axhausen, and G. Abay, “The acceptance of modal innovation: The case of swiss- metro,” inSwiss transport research conference, vol. 1, 2001

  46. [54]

    nuscenes: A multimodal dataset for autonomous driving,

    H. Caesar, V . Bankiti, A. H. Lang, S. Vora, V . E. Liong, Q. Xu, A. Krishnan, Y. Pan, G. Baldan, and O. Beijbom, “nuscenes: A multimodal dataset for autonomous driving,” inProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11621–11631, 2020

  47. [55]

    Scalability in perception for autonomous driving: Waymo open dataset,

    P . Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V . Patnaik, P . Tsui, J. Guo, Y. Zhou, Y. Chai, B. Caine,et al., “Scalability in perception for autonomous driving: Waymo open dataset,” inProceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, p...

  48. [56]

    Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset,

    S. Ettinger, S. Cheng, B. Caine, C. Liu, H. Zhao, S. Pradhan, Y. Chai, B. Sapp, C. R. Qi, Y. Zhou, et al., “Large scale interactive motion forecasting for autonomous driving: The waymo open motion dataset,” inProceedings of the IEEE/CVF international conference on computer vis...

  49. [57]

    On the new era of urban traffic monitoring with massive drone data: The pneuma large-scale field experiment,

    E. Barmpounakis and N. Geroliminis, “On the new era of urban traffic monitoring with massive drone data: The pneuma large-scale field experiment,”Transportation research part C: emerging tech- nologies, vol. 111, pp. 50–71, 2020

  50. [58]

    Evaluation of large-scale complete vehicle trajectories dataset on two kilometers highway segment for one hour duration: Zen traffic data,

    T. Seo, Y. Tago, N. Shinkai, M. Nakanishi, J. Tanabe, D. Ushirogochi, S. Kanamori, A. Abe, T. Ko- dama, S. Yoshimura,et al., “Evaluation of large-scale complete vehicle trajectories dataset on two kilometers highway segment for one hour duration: Zen traffic data,” in2020 Inte...

  51. [59]

    I-24 motion: An instrument for freeway traffic science,

    D. Gloudemans, Y. Wang, J. Ji, G. Zachar, W. Barbour, E. Hall, M. Cebelak, L. Smith, and D. B. Work, “I-24 motion: An instrument for freeway traffic science,”Transportation Research Part C: Emerging Technologies, vol. 155, p. 104311, 2023

  52. [60]

    Enhancing human mobility research with open and standardized datasets,

    T. Yabe, M. Luca, K. Tsubouchi, B. Lepri, M. C. Gonzalez, and E. Moro, “Enhancing human mobility research with open and standardized datasets,”Nature Computational Science, vol. 4, no. 7, pp. 469– 472, 2024

  53. [61]

    Yjmob100k: City-scale and longitudinal dataset of anonymized human mobility trajectories,

    T. Yabe, K. Tsubouchi, T. Shimizu, Y. Sekimoto, K. Sezaki, E. Moro, and A. Pentland, “Yjmob100k: City-scale and longitudinal dataset of anonymized human mobility trajectories,”Scientific Data, vol. 11, no. 1, p. 397, 2024

  54. [62]

    The dlr highway traffic dataset (dlr-ht): Longest road user trajectories on a german highway,

    C. Schicktanz, L. Klitzke, K. Gimm, R. L ¨udtke, K. Liesner, H. H. Mosebach, F. Heuer, A. Wodtke, and L. Asbach, “The dlr highway traffic dataset (dlr-ht): Longest road user trajectories on a german highway,”Authorea Preprints, 2025

  55. [63]

    Main challenges and opportunities, constraints and barriers of open science in transport research,

    K. Folla, A. Anagnostopoulou, and G. Yannis, “Main challenges and opportunities, constraints and barriers of open science in transport research,”International Journal of Research in Humanities and Social Studies, vol. 8, no. 5, 2021

  56. [64]

    Synshrp2: A synthetic multimodal benchmark for driving safety-critical events derived from real-world driving data,

    L. Shi, B. Jiang, Z. Yuan, M. A. Perez, and F. Guo, “Synshrp2: A synthetic multimodal benchmark for driving safety-critical events derived from real-world driving data,” inProceedings of the IEEE/CVF International Conference on Computer Vision, pp. 4586–4596, 2025

  57. [65]

    Big data in public transportation: a review of sources and methods,

    T. F. Welch and A. Widita, “Big data in public transportation: a review of sources and methods,” Transport reviews, vol. 39, no. 6, pp. 795–818, 2019. 33

  58. [66]

    Collect it all: National security, big data and governance,

    J. W. Crampton, “Collect it all: National security, big data and governance,”GeoJournal, vol. 80, no. 4, pp. 519–531, 2015

  59. [67]

    Threats related to open geospatial data in the uncertain geopolitical environment,

    J. Nikander, T. Jama, and H. Tenkanen, “Threats related to open geospatial data in the uncertain geopolitical environment,”International Archives of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. 48, no. 4/W12-2024, pp. 121–126, 2024

  60. [68]

    China’s privacy protection strategy and its geopolitical implications,

    C. Zhang, “China’s privacy protection strategy and its geopolitical implications,”Asian Review of Political Economy, vol. 3, no. 1, p. 6, 2024

  61. [69]

    A brief history of travel forecasting,

    M. Nie, “A brief history of travel forecasting,”Transportation, pp. 1–26, 2025

  62. [70]

    2024 iatbr lifetime achievement lecture – travel behavior research: Science, em- pirics, models and applications

    H. S. Mahmassani, “2024 iatbr lifetime achievement lecture – travel behavior research: Science, em- pirics, models and applications.” YouTube, TBD National Center, 2025. Event recording. Accessed: 2025-10-13

  63. [71]

    Eager: How well does traffic control theory and optimization research work on public roads?,

    M. Levin and R. Stern, “Eager: How well does traffic control theory and optimization research work on public roads?,” Tech. Rep. Award No. 2437781, National Science Foundation, September

  64. [72]

    Reproducible research in signal processing,

    P . Vandewalle, J. Kovacevic, and M. Vetterli, “Reproducible research in signal processing,”IEEE Signal Processing Magazine, vol. 26, no. 3, pp. 37–47, 2009

  65. [73]

    What does research reproducibility mean?,

    S. N. Goodman, D. Fanelli, and J. P . Ioannidis, “What does research reproducibility mean?,”Science translational medicine, vol. 8, no. 341, pp. 341ps12–341ps12, 2016

  66. [74]

    Chatgpt outperforms crowd workers for text-annotation tasks,

    F. Gilardi, M. Alizadeh, and M. Kubli, “Chatgpt outperforms crowd workers for text-annotation tasks,”Proceedings of the National Academy of Sciences, vol. 120, no. 30, p. e2305016120, 2023

  67. [75]

    Reproscreener: Leveraging llms for assessing computational repro- ducibility of machine learning pipelines,

    A. Bhaskar and V . Stodden, “Reproscreener: Leveraging llms for assessing computational repro- ducibility of machine learning pipelines,” inProceedings of the 2nd ACM Conference on Reproducibility and Replicability, ACM REP ’24, (New York, NY, USA), p. 101–109, Association for...

  68. [76]

    Core-bench: Fostering the credibil- ity of published research through a computational reproducibility agent benchmark,

    Z. S. Siegel, S. Kapoor, N. Nadgir, B. Stroebl, and A. Narayanan, “Core-bench: Fostering the credibil- ity of published research through a computational reproducibility agent benchmark,”Transactions on Machine Learning Research, 2024

  69. [77]

    The role of large language models in the peer-review process: opportu- nities and challenges for medical journal reviewers and editors,

    J. Lee, J. Lee, and J.-J. Yoo, “The role of large language models in the peer-review process: opportu- nities and challenges for medical journal reviewers and editors,”Journal of Educational Evaluation for Health Professions, vol. 22, 2025

  70. [78]

    The promise and challenges of using llms to accelerate the screening process of systematic reviews,

    A. Huotala, M. Kuutila, P . Ralph, and M. M ¨antyl¨a, “The promise and challenges of using llms to accelerate the screening process of systematic reviews,” inProceedings of the 28th International Con- ference on Evaluation and Assessment in Software Engineering, pp. 262–271, 2024

  71. [79]

    A critical assessment of large language models for systematic reviews: utilizing chatgpt for complex data extraction,

    H. Mahmoudi, D. Chang, H. Lee, N. Ghaffarzadegan, and M. S. Jalali, “A critical assessment of large language models for systematic reviews: utilizing chatgpt for complex data extraction,”Available at SSRN 4797024, 2024

  72. [80]

    Comparative analysis of smart city scientific research trends in the usa and china,

    Y. Lai and H. Zhao, “Comparative analysis of smart city scientific research trends in the usa and china,”Nature Cities, pp. 1–9, 2025

  73. [81]

    Artificial intelligence faces reproducibility crisis,

    M. Hutson, “Artificial intelligence faces reproducibility crisis,” 2018

  74. [82]

    Croissant: A metadata format for ml-ready datasets,

    M. Akhtar, O. Benjelloun, C. Conforti, L. Foschini, J. Giner-Miguelez, P . Gijsbers, S. Goswami, N. Jain, M. Karamousadakis, M. Kuchnik,et al., “Croissant: A metadata format for ml-ready datasets,”Advances in Neural Information Processing Systems, vol. 37, pp. 82133–82148, 2024. 34

  75. [83]

    Reproducibility in management science,

    M. Fi ˇsar, B. Greiner, C. Huber, E. Katok, A. I. Ozkes, and M. S. R. Collaboration, “Reproducibility in management science,”Management Science, vol. 70, no. 3, pp. 1343–1356, 2024

  76. [84]

    Welcome to the ijoc software and data repositories,

    INFORMS Journal on Computing, “Welcome to the ijoc software and data repositories,” 2025. Ac- cessed 2025-12-24

  77. [85]

    Information infrastructure for research col- laboration in land use, transportation, and environmental planning,

    J. Ferreira Jr, M. Diao, Y. Zhu, W. Li, and S. Jiang, “Information infrastructure for research col- laboration in land use, transportation, and environmental planning,”Transportation research record, vol. 2183, no. 1, pp. 85–93, 2010

  78. [86]

    The oxford radar robotcar dataset: A radar extension to the oxford robotcar dataset,

    D. Barnes, M. Gadd, P . Murcutt, P . Newman, and I. Posner, “The oxford radar robotcar dataset: A radar extension to the oxford robotcar dataset,” in2020 IEEE international conference on robotics and automation (ICRA), pp. 6433–6438, IEEE, 2020

  79. [87]

    The ind dataset: A drone dataset of naturalistic road user trajectories at german intersections,

    J. Bock, R. Krajewski, T. Moers, S. Runde, L. Vater, and L. Eckstein, “The ind dataset: A drone dataset of naturalistic road user trajectories at german intersections,” in2020 IEEE Intelligent Vehicles Symposium (IV), pp. 1929–1934, IEEE, 2020

  80. [88]

    Flow: A Modular Learning Frame- work for Mixed Autonomy Traffic,

    C. Wu, A. R. Kreidieh, K. Parvate, E. Vinitsky, and A. M. Bayen, “Flow: A Modular Learning Frame- work for Mixed Autonomy Traffic,”IEEE Transactions on Robotics (T-RO), vol. 38, pp. 1270–1286, July 2021

  81. [89]

    Dl- traff: Survey and benchmark of deep learning models for urban traffic prediction,

    R. Jiang, D. Yin, Z. Wang, Y. Wang, J. Deng, H. Liu, Z. Cai, J. Deng, X. Song, and R. Shibasaki, “Dl- traff: Survey and benchmark of deep learning models for urban traffic prediction,” inProceedings of the 30th ACM international conference on information & knowledge management...

  82. [90]

    Reproducible research in transportation,

    “Reproducible research in transportation,” 2025. focuses on reproducible practices in transportation research

  83. [91]

    How can we ensure visibility and diversity in research contributions? how the contributor role taxonomy (credit) is helping the shift from authorship to contributorship,

    L. Allen, A. O’Connell, and V . Kiermer, “How can we ensure visibility and diversity in research contributions? how the contributor role taxonomy (credit) is helping the shift from authorship to contributorship,”Learned Publishing, vol. 32, no. 1, pp. 71–74, 2019

  84. [92]

    Artifact review and badging,

    ACM, “Artifact review and badging,” 2020. Version 1.1

  85. [93]

    Reproducible, reusable, and robust reinforcement learning,

    J. Pineau, “Reproducible, reusable, and robust reinforcement learning,”Advances in Neural Informa- tion Processing Systems, 2018

  86. [94]

    The State of Open Data 2024: Special Report Bridging policy and practice in data sharing,

    M. Hahnel, G. Smith, and A. Campbell, “The State of Open Data 2024: Special Report Bridging policy and practice in data sharing,” 12 2024

  87. [95]

    REproducible research in transportation engineering (rerite) working group

    “REproducible research in transportation engineering (rerite) working group.”https://www. rerite.org/, 2024. Accessed: 2026-01-05

  88. [96]

    Towards wide-scale adoption of open science practices: The role of open science communities,

    K. Armeni, L. Brinkman, R. Carlsson, A. Eerland, R. Fijten, R. Fondberg, V . E. Heininga, S. Heunis, W. Q. Koh, M. Masselink,et al., “Towards wide-scale adoption of open science practices: The role of open science communities,”Science and Public Policy, vol. 48, no. 5, pp. 605...

  89. [97]

    Reproducibil- ity in transportation research: A hands-on tutorial 2.0 – reproducible manuscripts

    Z. Zheng, D. Levinson, S. Mohammadian, B. Ghosh, I. Mart ´ınez, and C. Wu, “Reproducibil- ity in transportation research: A hands-on tutorial 2.0 – reproducible manuscripts.”https:// zuduozheng.github.io/RERITE_TUTORIAL_2.0, 2025. Accessed: 2026-01-05

  90. [98]

    Generative ai poses ethical challenges for open science,

    L. Acion, M. Rajngewerc, G. Randall, and L. Etcheverry, “Generative ai poses ethical challenges for open science,”Nature Human Behaviour, vol. 7, no. 11, pp. 1800–1801, 2023

  91. [99]

    Text and data mining license

    Elsevier, “Text and data mining license.”https://www.elsevier.com/about/ policies-and-standards/text-and-data-mining/license, 2025. Accessed: 2025-10-22

  92. [100]

    Software Framework for Topic Modelling with Large Corpora,

    R. ˇReh ˚uˇrek and P . Sojka, “Software Framework for Topic Modelling with Large Corpora,” inProceed- ings of the LREC 2010 Workshop on New Challenges for NLP Frameworks, (Valletta, Malta), pp. 45–50, ELRA, May 2010. 35

  93. [2024]

    Grant awarded to University of Minnesota–Twin Cities, effective 1 September 2024; expires 31 August 2026

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.