{"id":"d160158d-3576-47bd-8828-0295b2429a67","arxiv_id":"2601.14429","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Only ~5% of quantitative transportation papers share code and ~4% share data repositories, a measured baseline from an LLM pipeline applied to 10,724 articles.","lead":"This paper uses an LLM-based pipeline to check 10,724 transportation journal articles for publicly shared code and data, finding only 5% share code and 4% share a data repository. A smart generalist should read it because it turns the vague concern 'transportation research isn't open' into a measured, field-wide baseline with clear policy implications for journals and funders.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline availability rates rest on unquantified transfer from a 96-paper validation set; the Fair-agreement data-repository feature (Fleiss κ=0.399) can materially move a 4% rate, so a larger validation sample and misclassification correction are needed.","rationale":"The reader's weakest assumption is exactly the load-bearing point: agreement measured on the 96-paper MVD is assumed to transfer to the full corpus, and the weakest-validity feature (data repository, Fleiss κ=0.399) directly feeds the 4% headline. I agree with that framing. My stress-test adds two manuscript-internal details that sharpen it. First, the MVD prevalence table shows the LLM agrees with H2 (3.12%) but sits 5.2 points below H1 (8.33%); with only 96 observations, that spread is wide relative to a 4% headline rate. Second, the data dictionary in Appendix E describes uses_data_bool as 'after manual correction,' and Appendix D filters out 'irreconcilable inconsistencies' without reporting how many or which records; either step could affect the final availability flags and would make the pipeline something other than a pure LLM measurement. The right response is not rejection—the pipeline is transparent, the code/data-repo rates are plausibly low, and the human-human agreement is also weak on the same feature—but the conditional verdict should be retained until the transfer assumption is tested more directly. My concrete check would resolve whether the rates are robust to measurement error and to the undocumented postprocessing steps.","tokens_in":31157,"tokens_out":5932,"duration_ms":74801,"concrete_test":"Re-annotate a stratified random sample of 400 papers from the 10,724 (not the original 96) with two annotators, following Appendix A definitions; compute the LLM-vs-human misclassification matrix for is_data_repository_available and is_code_publicly_available; apply it to the full corpus and report corrected rates with 95% bootstrap CIs. If the corrected data-repository rate remains below 8% and the code rate within 3–8%, the headline claim survives; if not, the validation-transfer assumption fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central measurement claims (5%/4%/3%) treat LLM outputs on 10,724 papers as ground truth. The only evidence for this is agreement on 96 random papers (§3.2, §4.4). Table 4 shows is_data_repository_available has Fleiss κ=0.399 (Fair) and human-human κ=0.333; is_data_cited κ=0.522; is_quantitative κ=0.459. The LLM's prevalence for data repositories is 3.12%, compared with H1=8.33% and H2=3.12% (Table 3). Because the headline data-repo rate is 4%, these errors are not small relative to the signal, and no confidence interval or misclassification correction is reported for §5.1 rates or the choice models. The problem is compounded by two postprocessing steps: Appendix E lists uses_data_bool as 'after manual correction,' and Appendix D excludes records with 'irreconcilable inconsistencies' (10,724→10,480), without quantifying what is dropped. If either step affects the availability flags, the pipeline is not the purely LLM-based measurement the validation section assumes.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an LLM-based pipeline for extracting code and data availability indicators from full-text articles, applies it to 10,724 Transportation Research Part A–F and TR-IP articles published 2019–2024, and validates the extraction on a 96-paper manual validation dataset (MVD) via inter-rater agreement analysis. The headline findings are that about 5% of quantitative papers share a code repository, 4% share a data repository, and about 3% share both; that sharing varies by journal, topic, and corresponding-author region; and that sharing is not associated with higher citation counts or shorter review times. The paper also provides topic models, logit choice models, and an interactive explorer.","tokens_in":31481,"tokens_out":3053,"duration_ms":38510,"significance":"If the measurement is reliable, this is a valuable first large-scale, field-wide snapshot of open science practices in transportation research. The pipeline is reproducible and the authors make code and permissible data available; the comparison against text-search baselines (Appendix C) and the explicit inter-rater agreement analysis are commendable and go beyond what is common in LLM-as-annotator studies. The finding that data/code availability is uncorrelated with citation outcomes, if confirmed with proper inference, has clear policy relevance for journals and funders. However, the central claim rests critically on the validity of the LLM-extracted features at scale, and the evidence for that validity is the agreement on only 96 papers—with the data-repository feature showing only Fair inter-rater agreement (Fleiss κ=0.399, H1–H2 Cohen κ=0.333). This uncertainty is not propagated into the headline rates or the choice models, which is the main load-bearing weakness.","major_comments":[{"comment":"The headline 4% data-repository rate is based on an LLM feature whose agreement is weak: Fleiss κ=0.399 (Fair) and, crucially, human-human Cohen κ=0.333 (Table 4). Table 3 shows prevalence differences that are material at this rate: H1=8.33%, H2=3.12%, LLM=3.12%. With a 96-paper validation set, the LLM's apparent agreement with H2 on this feature (Cohen κ=0.656) could reflect rater-specific bias rather than accuracy. The paper acknowledges this ('retained with caution') but does not quantify the impact on the 5%/4%/3% estimates. Please report the MVD confusion matrix for this feature, compute misclassification-corrected prevalence estimates and confidence intervals (e.g., via bootstrap or a latent-class model), and re-run the §5 choice models with measurement-error propagation or at least a sensitivity analysis. Without this, the exact headline rates are not statistically grounded, even","section":"§4.4.2, Table 4; §5.1"},{"comment":"The pipeline is described as LLM-based and validated against the 96-paper MVD, but the postprocessing includes steps outside that validation: Appendix E lists 'uses_data_bool' as 'after manual correction,' and Appendix D excludes records with 'irreconcilable inconsistencies' (10,724 → 10,480) without reporting how many papers or which flags were affected. If manual correction changed availability flags, or if the excluded papers are non-random with respect to availability, the agreement statistics in §4.4 do not cover the final analysis dataset. Please quantify: how many papers were manually corrected, how many were excluded, and whether any headline rates change materially when the corrected/excluded cases are handled differently.","section":"§4.5, Appendix D, Appendix E"},{"comment":"The abstract and conclusions state that there is 'no significant difference in citation counts or review duration between papers that provided data and code and those that did not.' The support in §5.4 is descriptive: LOESS curves in Figure 4 and a comparison of mean review times (269 vs. 254 days, CI 2–27 days). No regression or hypothesis test is reported for citations that controls for paper age, journal, topic, or other confounders, and the figure only shows visual overlap. The review-time claim is also based on a single unadjusted comparison. Please provide a formal statistical analysis (e.g., a citation-count model with controls, and a regression of review time on availability indicators including journal fixed effects) before making this incentive-gap claim. This is load-bearing for the policy recommendation that current incentives do not reward sharing.","section":"§5.4"}],"minor_comments":[{"comment":"Typo: 'We also calculate' should be 'We also calculated.'","section":"§4.4.1"},{"comment":"The statement 'Given that some of these repositories contain repackaged datasets, less than 4% of the papers we reviewed contributed new data to the transportation community' is presented as a conclusion, but the paper does not systematically classify whether shared repositories contain new vs. repackaged data. Please mark this as an interpretive inference and support it with evidence or soften the claim.","section":"§9"},{"comment":"The limitation that temperature was not set to 0 during extraction is acknowledged, but the practical impact on reproducibility is not discussed. Since the code and data are to be released, please also release the exact prompt versions and model snapshot identifiers, and consider stating whether rerunning with temperature 0 changes any of the headline rates.","section":"§6.2"},{"comment":"The caption for Table 3 says 'Prev. (H1), Prev. (H2), Prev. (LLM)' but the column headers show 'Prev.' only; please align headers with the caption for clarity.","section":"Table 3 and Table 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is well motivated and the pipeline is a useful contribution, but the central quantitative claims need to be made robust to the measurement error documented in the paper itself. The 96-paper validation set is small, and the data-repository feature—central to the 4% headline—has only Fair agreement even between humans. The presence of manual correction in the postprocessing (Appendix E) and exclusions (Appendix D) further complicates the claim that the LLM pipeline is validated. I would not reject, but the authors need to present corrected estimates or clear sensitivity analyses, and they need to replace the visual/no-formal-test incentive analysis with a proper regression before the conclusions can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is the first large-scale, field-wide measurement of code and data availability in transportation research, and that alone makes it worth a serious look. They built an LLM pipeline, validated it against two human annotators on 96 papers, and ran it on 10,724 TR journal articles. The headline findings — 5% code, 4% data repository, 3% both, no citation or review-time incentive — are new and directionally credible. The paper also does something right that most bibliometric work skips: it formalizes decision rules, publishes the prompts, compares against regex baselines, and reports inter-rater agreement instead of just claiming accuracy. The comparison showing regex/text search badly overestimates availability is genuinely useful for the open-science monitoring community. The authors are also honest about limitations: they note the small validation set, the non-zero temperature, and that they don't measure reproducibility itself.\n\nThe main soft spot is exactly what the stress-test note flags. The headline percentages rest on LLM labels treated as ground truth for 10,724 papers, validated on only 96 papers. The weakest feature, is_data_repository_available, has Fleiss kappa 0.399 and human-human kappa 0.333. That is the feature feeding the 4% headline number. With a base rate around 4%, a few dozen misclassifications can move the rate by a full percentage point, and the paper does not report confidence intervals or a misclassification correction. The postprocessing also drops 244 records with 'irreconcilable inconsistencies' and Appendix E mentions manual correction of uses_data_bool. Those steps may be fine, but they are not quantified, so the pipeline is not purely the automated measurement the validation section describes. The choice models are built on the same noisy features, and the significance-driven specification search is not stress-tested. None of this breaks the central claim — rates are low even with large measurement error — but it does mean the exact 5%/4%/3% figures should be treated as estimates with meaningful uncertainty, not precise baselines.\n\nThe work is a solid contribution to open-science monitoring in transportation. It deserves peer review. The authors should be asked to release code and data (they say they will), expand the validation set or at least quantify misclassification impact on the headline rates, and report the number and nature of excluded records. For a reader who wants the first field-wide baseline of sharing practices in transportation, this is the paper to cite; for anyone who needs precise point estimates, wait for the measurement-error analysis.\n\nSend to peer review. Conditional accept in principle, but require the validation and uncertainty fixes before it becomes a baseline others build on.","headline":"First large-scale LLM-based baseline of code/data sharing in transportation; rates are directionally credible but the 96-paper validation and unpropagated measurement error make the exact headline numbers softer than they look.","tokens_in":31959,"tokens_out":1283,"would_cite":true,"duration_ms":14493,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"After an LLM sweep of 10,724 transportation papers, only 5% share code and 4% share data, and sharing earns no citation or review-time reward.","keywords":["open science","open science monitoring","transportation research","large language models","data availability","code availability","reproducibility","inter-rater agreement"],"falsifier":"Re-annotate a fresh random sample of roughly 300 papers from the same corpus with two human annotators using the paper's own definitions, then compare the human code/data repository rates with the pipeline's rates on those same papers; the central measurement collapses if the disagreement is large enough to move the 5%/4% figures—especially for 'data repository available,' which already showed only Fair inter-rater agreement on the 96-paper validation set.","tokens_in":31078,"feed_emoji":"🚗","tokens_out":7453,"duration_ms":72079,"temperature":0.7,"pith_summary":"The paper sets out to show that the state of open science in transportation research can be measured automatically, at scale, and reliably, using large language models to read full-text articles. Analyzing 10,724 papers from the Transportation Research Part journals (2019–2024), it reports that only about 5% of quantitative papers make code available and 4% make data available in a repository, with about 3% doing both. It also finds that sharing is uneven across journals, topics, and world regions, and that it brings no measurable benefit in citations or review speed—code-sharing papers actually took 2 to 27 days longer to review. The authors conclude that journals and funding agencies, not author self-interest, must drive structural change, and that the pipeline can serve as a repeatable monitor for the field.","feed_headline":"Only 5% of transport papers share code, 4% data","feed_subtitle":"LLM audit of 10,724 Transportation Research articles finds sharing earns no citation edge and slower reviews.","key_machinery":"The key machinery is an automated three-stage pipeline: it parses the journals' structured XML full texts into markdown, prompts a large language model with separate single-task prompts to extract availability flags (code used, code publicly available, data cited, data repository available, quantitative study), and validates those flags against two human annotators on a 96-paper manual set using inter-rater agreement metrics (kappa). Only features meeting agreement thresholds are retained, and link-checking verifies that stated repository URLs are live and contain the claimed artifact. The validated flags feed two logit choice models that estimate how journal, region, topic, and paper charac","core_discovery":"The central discovery is a measured baseline for open-science practice in a flagship journal series. Using an LLM pipeline validated against two human annotators on a 96-paper manual set, the authors estimate that of 10,480 quantitative research articles, 5% share a code repository, 4% share a data repository, and about 3% share both, while 29% cite or link existing public datasets and only a small fraction contribute new data. Sharing concentrates in particular journals (TR-B and TR-C), in data-driven modeling topics, and among corresponding authors outside Asia, and it increases for newer papers. The paper's headline negative result is that papers sharing code or data do not receive more c","pith_inferences":["Because the pipeline only counts repositories explicitly mentioned in the full text, the 5%/4% figures are lower bounds under the paper's scope; adding an external web search for shared artifacts would likely raise the measured rates.","The no-citation/no-review-time result is purely observational; isolating the causal effect of sharing would require a quasi-experimental design, such as comparing citation trajectories before and after a journal's open-science policy change.","The 96-paper validation set and the Fair inter-rater agreement on 'data repository available' mean the headline rates carry nontrivial measurement error; a larger, stratified validation sample would be needed to detect year-over-year trend changes with confidence.","The paper's choice-model findings suggest a testable extension: apply the same pipeline to the journals' policy changes (e.g., a journal that begins requiring availability statements) to see whether mandated statements, rather than soft encouragement, shift the 5%/4% figures."],"forward_implications":["The field now has a repeatable baseline: future audits can rerun the pipeline to track whether open-science rates rise over time.","Because sharing brings no citation or review-time benefit, journal policies—structured availability statements, open-science checklists, badges—and funder requirements are the levers most likely to move the rates.","At 3–5% availability, most transportation papers fail a necessary condition for computational reproducibility, since sharing data and code is a prerequisite for others to verify results.","The large gap between citing others' data (29%) and sharing new data (<4%) implies that a handful of datasets supports a large share of the field's empirical work, so publishing new open datasets could have outsized impact.","The pipeline is transferable to other journals in the same publisher and to other publishers with text-mining access, allowing cross-field comparison of open-science practice."],"fun_headline_variants":["LLM audit: 5% of transport papers share code, 4% data","No citation reward for sharing code or data in transport research","Open science gap: 95% of transport papers don't share code","Transport research fails open science: only 5% share code","Sharing code in transport papers yields no citation boost"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the LLM's agreement with two human annotators on 96 papers—Almost Perfect for code availability but only Fair for data-repository availability—transfers intact to the full 10,724-paper corpus, so the extracted features can be treated as ground truth; if that transfer fails, the 5% and 4% rates could shift materially.","fun_headline_variants_meta":{"raw":{"variants":["LLM audit: 5% of transport papers share code, 4% data","No citation reward for sharing code or data in transport research","Open science gap: 95% of transport papers don't share code","Transport research fails open science: only 5% share code","Sharing code in transport papers yields no citation boost"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000777,"raw_usage":{"total_tokens":3309,"prompt_tokens":815,"completion_tokens":2494,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":2404}},"tokens_in":559,"tokens_out":2494,"duration_ms":15516,"temperature":1.0,"reasoning_tokens":2404,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T09:11:43.938401+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a fresh random sample of roughly 300 papers from the same corpus with two human annotators using the paper's own definitions, then compare the human code/data repository rates with the pipeline's rates on those same papers; the central measurement collapses if the disagreement is large enough to move the 5%/4% figures—especially for 'data repository available,' which already showed only Fair inter-rater agreement on the 96-paper validation set.","supporting_citations":[],"review_version":1}