{"id":"845e6b88-3981-4ac8-bf51-eda0fbb828c4","arxiv_id":"2606.01109","paper_version":1,"verdict":"CONDITIONAL","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"FOSSIL is a multilingual dataset of 96 annotated SSH articles with 7,600+ footnote references and a Grobid specialization that raises micro-F1 from 0.36 to 0.72.","lead":"The authors created FOSSIL, a new open dataset of 96 law and humanities articles containing over 7,600 annotated footnote references, plus an annotation tool and a specialized Grobid pipeline. Researchers building citation tools or digital libraries for SSH fields may find it useful because standard tools fail on footnote-based citations mixed with commentary.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Reported F1 gains depend on a gold standard whose completeness and consistency remain undocumented","rationale":"The reader’s weakest assumption (reliability of the seven-annotator gold standard) is exactly the load-bearing point for the empirical claim. Because the paper itself flags ongoing annotation work and the abstract supplies no agreement statistics, the concern is internal to the reported evaluation rather than an external consensus issue. No other technical assumption (e.g., Grobid specialization mechanics) is shown to be weaker.","tokens_in":1724,"tokens_out":324,"duration_ms":14065,"concrete_test":"Finish all segmentation/parsing/cross-reference annotations on the 96 articles, compute pairwise agreement (Cohen’s κ or Fleiss’ κ) on a 20-article overlap set, then re-evaluate both pipelines on the completed labels; if κ < 0.75 or the new micro-F1 differs by >0.05, the doubling claim must be qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline claim (specialized Grobid reaching micro-F1 0.72 vs. 0.36) is evaluated end-to-end on the FOSSIL labels. The abstract states that segmentation, parsing, and cross-reference annotations are still ongoing and supplies no inter-annotator agreement figures for the seven-annotator workflow. Without those quantities it is impossible to know whether the 0.72 figure reflects stable reference boundaries or a provisional labeling that will shift once the remaining tasks finish.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces FOSSIL, an openly licensed multilingual dataset of 96 annotated scholarly articles from law and humanities containing over 7,600 footnote-embedded references. It describes the PDF-TEI Editor annotation tool, a documented seven-annotator workflow, and a Grobid specialization for footnote-based citations. In end-to-end evaluation the specialized pipeline achieves micro-F1 of 0.72 versus 0.36 for default Grobid, driven mainly by recall gains, while noting substantial remaining headroom for cross-references and mixed-content footnotes. The manuscript is presented as an extended abstract of work in progress, with segmentation, parsing, and cross-reference annotations still ongoing.","tokens_in":1816,"tokens_out":442,"duration_ms":19207,"significance":"Creation of a new, openly licensed gold-standard resource for footnote citation extraction in SSH fields addresses a clear gap, as most existing tools target end-of-document bibliographies in the sciences. If the annotations prove reliable, the reported near-doubling of extraction quality provides concrete evidence that domain-specific specialization can yield substantial gains and supplies a reusable benchmark for future work on complex, multilingual, mixed-content footnotes.","major_comments":[{"comment":"Abstract: The headline result (micro-F1 rising from 0.36 to 0.72) is evaluated end-to-end against the FOSSIL labels, yet the abstract states that 'annotations of citations segmentation and parsing, and cross-reference resolution are ongoing.' Because the evaluation depends on these still-incomplete labels, the reported performance figure and the claim of nearly doubled quality cannot yet be treated as stable.","section":"Abstract"},{"comment":"Abstract: The seven-annotator workflow is presented as the foundation for the dataset and evaluation, but no inter-annotator agreement statistics or consistency metrics are supplied. Without these quantities it is impossible to determine whether the 0.72 micro-F1 reflects reliable reference boundaries or provisional labeling that may shift once the remaining annotation tasks are finished.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful review and for recognizing the value of FOSSIL as a resource for footnote citation extraction. We address the two major comments point by point below, with proposed revisions where appropriate.","responses":[{"response":"The reported evaluation measures reference boundary detection (the core extraction task) against the completed annotations of the 7,600+ footnote references. The ongoing work concerns only the subsequent tasks of intra-footnote segmentation, bibliographic field parsing, and cross-reference resolution; these are not required for the boundary-level evaluation presented. The 0.72 micro-F1 therefore rests on the stable reference-identification labels. We will revise the abstract to explicitly separate the completed annotation components from the ongoing ones and to qualify the performance claim accordingly.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The headline result (micro-F1 rising from 0.36 to 0.72) is evaluated end-to-end against the FOSSIL labels, yet the abstract states that 'annotations of citations segmentation and parsing, and cross-reference resolution are ongoing.' Because the evaluation depends on these still-incomplete labels, the reported performance figure and the claim of nearly doubled quality cannot yet be treated as stable."},{"response":"We agree that quantitative inter-annotator agreement metrics would strengthen confidence in the labels. The documented workflow incorporated multiple rounds of joint review and adjudication, but formal IAA statistics were not computed for this extended abstract. We will expand the description of the workflow to detail the consistency measures employed and will include any available agreement figures or a clear statement of this limitation in the revised manuscript.","revision_made":"partial","referee_comment":"[Abstract] Abstract: The seven-annotator workflow is presented as the foundation for the dataset and evaluation, but no inter-annotator agreement statistics or consistency metrics are supplied. Without these quantities it is impossible to determine whether the 0.72 micro-F1 reflects reliable reference boundaries or provisional labeling that may shift once the remaining annotation tasks are finished."}],"tokens_in":1410,"tokens_out":444,"duration_ms":25079,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is a new openly licensed dataset of 96 articles with over 7,600 footnote references, aimed at a domain where standard citation tools fail because references sit inside footnotes mixed with commentary and cross-references.\n\nWhat the work actually delivers is the dataset itself, a documented seven-annotator workflow, the PDF-TEI Editor tool, and a Grobid model retrained on footnote material. The reported end-to-end gain is real on the numbers given: recall improves enough to nearly double the score. That addresses a documented gap in SSH digital library work.\n\nThe soft spot is that this is explicitly work in progress. Segmentation, parsing, and cross-reference annotations are still ongoing, and the abstract gives no inter-annotator agreement numbers or error analysis on the current labels. The 0.72 figure therefore rests on a partial gold standard that could shift once the remaining tasks finish. No circularity or invented results, just an incomplete evaluation.\n\nThis paper is for people building citation parsers or digital libraries for legal and humanities texts who need training data in that style. A reader working on general scientific bibliography tools will get less from it.\n\nIt deserves peer review once the full annotation set and consistency checks are released, because the resource fills a real hole even if the current numbers are provisional.","headline":"FOSSIL supplies the first open multilingual dataset of footnote citations from law and humanities papers plus a Grobid tweak that lifts micro-F1 from 0.36 to 0.72 on their labels, but the labels are still incomplete.","tokens_in":2348,"tokens_out":363,"would_cite":false,"duration_ms":11877,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A new dataset of footnote citations and a specialized Grobid pipeline nearly double extraction quality over the default tool.","keywords":["citation extraction","footnotes","law and humanities","dataset","Grobid","reference parsing","digital libraries","annotation workflow"],"falsifier":"Re-running the end-to-end evaluation on a fresh collection of law and humanities articles whose footnotes were annotated independently of the original seven-annotator process and finding that the specialized pipeline no longer outperforms default Grobid.","tokens_in":2592,"feed_emoji":"","tokens_out":675,"duration_ms":16273,"temperature":0.7,"pith_summary":"The paper introduces FOSSIL, an openly licensed multilingual dataset of 96 law and humanities articles containing more than 7,600 annotated footnote references. It supplies a collaborative annotation tool, a seven-annotator workflow, and a customized Grobid model trained to handle the interleaved commentary, cross-references, and stylistic variety typical of footnotes. End-to-end tests show the specialized pipeline raises micro-F1 from 0.36 to 0.72, mainly by lifting recall, while leaving room for improvement on mixed-content cases. Existing citation tools were built for end-of-document bibliographies in the sciences and perform poorly on the footnote-heavy style of law and humanities scholarship.","feed_headline":"Specialized pipeline doubles citation extraction quality from footnotes","feed_subtitle":"New FOSSIL dataset and Grobid model raise micro-F1 from 0.36 to 0.72 on law and humanities articles with footnote references.","key_machinery":"FOSSIL dataset of annotated footnote-embedded references paired with the specialized Grobid pipeline for extracting citations from footnotes rather than end-of-document lists.","core_discovery":"The authors create FOSSIL together with PDF-TEI Editor and a documented annotation workflow, then train a Grobid specialization for footnote-based citations; in end-to-end evaluation the specialized pipeline raises micro-F1 from 0.36 to 0.72 over default Grobid, driven by higher recall, although cross-references and mixed-content footnotes remain challenging.","pith_inferences":["The same specialization approach could be tested on historical or non-English footnote corpora outside the current 96 articles.","Higher-quality footnote extraction might enable large-scale mapping of citation practices across entire law and humanities literatures.","The dataset could support studies that compare citation density or style differences between law, history, and philosophy subfields."],"forward_implications":["Automated processing of legal and humanities literature can now reach substantially higher recall for footnote references.","Openly licensed training data becomes available for further refinement of citation parsers aimed at SSH fields.","Remaining performance gaps on cross-references and mixed-content footnotes are now quantifiable and can guide targeted improvements.","The annotation workflow offers a reusable template for creating gold standards in additional languages or citation styles."],"fun_headline_variants":["FOSSIL dataset improves micro-F1 for footnote citations to 0.72","New Grobid specialization for law and humanities footnote citations","FOSSIL provides multilingual gold standard for citation extraction","Annotation workflow raises citation extraction quality from footnotes","FOSSIL and PDF-TEI Editor for footnote reference labeling"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The seven-annotator workflow and its labels form a reliable gold standard that captures the range of footnote styles, languages, and mixed-content cases found in actual law and humanities articles.","fun_headline_variants_meta":{"raw":{"variants":["FOSSIL dataset improves micro-F1 for footnote citations to 0.72","New Grobid specialization for law and humanities footnote citations","FOSSIL provides multilingual gold standard for citation extraction","Annotation workflow raises citation extraction quality from footnotes","FOSSIL and PDF-TEI Editor for footnote reference labeling"]},"model":"grok-4.3","cost_usd":0.006601,"raw_usage":{"total_tokens":3070,"prompt_tokens":644,"num_sources_used":0,"completion_tokens":82,"cost_in_usd_ticks":66012000,"prompt_tokens_details":{"text_tokens":644,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2344,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":644,"tokens_out":82,"duration_ms":21920,"temperature":1.0,"reasoning_tokens":2344,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T16:14:00.380827+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Re-running the end-to-end evaluation on a fresh collection of law and humanities articles whose footnotes were annotated independently of the original seven-annotator process and finding that the specialized pipeline no longer outperforms default Grobid.","supporting_citations":[],"review_version":1}