{"id":"f71a50d0-c693-4409-8a86-c63573399d96","arxiv_id":"2412.13140","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A complete manual audit of Tobacco3482 finds 11.7% label errors, 16.7% multi-label samples, and that roughly a third of a top model's errors are actually valid labels.","lead":"The paper manually re-reviewed every image in the Tobacco3482 document classification benchmark and reports that 11.7% of samples are mislabeled or unclassifiable and 16.7% have multiple valid labels. It also shows that about 35% of a strong model's mistakes vanish under the corrected labels, so published accuracy numbers on this benchmark are partly artifacts of label noise.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central percentages rest on a single annotator's application of self-authored guidelines; without inter-annotator agreement the 11.7%, 16.7%, and 196/554 correction figures are not established as reproducible ground truth.","rationale":"The reader's weakest assumption identifies the same load-bearing concern: the correctness of the central percentages depends on the Table 1 guidelines and their consistent application by a single annotator. My stress-test finds no additional internal inconsistency that would change the verdict. The paper is transparent about its method, provides code and annotations, and does not overstate its ability to distinguish between subjective and objective label judgments. However, the paper does present the numbers as definitive findings rather than as one-annotator estimates, and the model-impact analysis directly inherits that subjectivity. The appropriate verdict remains CONDITIONAL: accept the direction and the release of artifacts, but require publication of inter-annotator reliability (or at minimum an explicit statement that the figures are single-annotator estimates) before the precise percentages are cited as ground truth. The proposed replication check is feasible because the dataset is public and the annotation artifacts are released; it directly tests whether the claim is reproducible under the same guidelines. I therefore leave the reader's verdict unchanged and do not escalate to rejection, since the concern is about uncertainty quantification rather than a demonstrated fatal flaw.","tokens_in":5470,"tokens_out":3540,"duration_ms":35655,"concrete_test":"Run an independent replication study: recruit two annotators who have not seen the paper's annotation output, give them only the Tobacco3482 category names and the paper's Table 1 guidelines, and have each independently label a stratified random sample of at least 500 documents (50 per category). Compute per-category percent agreement and Cohen's kappa between the original annotation set and each new annotator, and between the two new annotators. If kappa is below 0.8, or if the estimated unknown/mislabeled/multi-label rates differ from 11.7%/16.7% by more than 3 percentage points, report the paper's figures as single-annotator estimates and revise the DiT accuracy correction accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's quantitative claims—151 unknown, 258 mislabeled, 583 multi-label, and the resulting 196 of 554 DiT mistakes counted as non-mistakes—are all derived from one team's own labeling guidelines (Table 1) applied in a single pass, with no reported inter-annotator agreement, adjudication, or external validation. The guidelines contain explicitly subjective exclusion rules (e.g., ADVE excludes 'magazine covers and content drafts'; Email excludes 'HTML code'; Letter excludes memos using 'letter' in its subject). These decisions are load-bearing: e.g., the Letter category's 24.0% mislabel rate depends on whether a document is judged a physical letter versus a memo addressed to an organization based on ambiguous cues like letterhead and 'To/From'. A second annotator applying reasonable but different boundary decisions could shift the unknown/mislabeled counts substantially. The model-impact analysis inherits this fragility: 147 of the 196 'not actually mistakes' are cases where the model predicted an alternative label that the authors' multi-label annotation deemed valid, and 49 are cases where the authors deemed the true label 'unknown'. If the annotation set changes, the 89.7% corrected accuracy and the '35% of mistakes' attribution change accordingly. The paper provides code and annotation files, which is helpful, but the artifacts do not by themselves prove reliability; they merely make the single annotator's judgments inspectable. The central claim is not internally inconsistent, but its precision (11.7%, 16.7%, 89.7%) is not yet warranted by the evidence presented.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a manual audit of all 3,482 images in the Tobacco3482 document classification dataset. The authors first define a set of category guidelines (Table 1), then re-annotate every document, finding 151 samples with no valid Tobacco3482 category, 258 samples with a wrong original label, and 583 samples with multiple valid labels. They further analyze the mistakes of a DiT model evaluated under 4-fold cross-validation, reporting 554 total mistakes, of which 196 are claimed to be valid alternative labels or unlabeled/unknown cases, raising the model's accuracy from 84.1% to 89.7% under their reannotation criteria.","tokens_in":5761,"tokens_out":6577,"duration_ms":62996,"significance":"If the audit is reliable, this is a useful contribution to the growing literature on benchmark label quality, providing the first systematic error count for Tobacco3482 and extending the authors' earlier RVL-CDIP analysis. The paper is transparent about its annotation artifacts, releasing the new labels and code, and its headline numbers are internally consistent (409/3482 = 11.7%, 583/3482 = 16.7%, 196/554 = 35.4%). The main value is conditional on the reproducibility of the annotation procedure: all quantitative claims rest on a single annotator's application of self-authored guidelines, with no inter-annotator agreement or external adjudication. With an additional reliability study or a substantially weakened framing, the paper could be a valuable dataset-quality reference; in its current form, the central percentages are plausible but not established as stable ground truth.","major_comments":[{"comment":"The quantitative claims (151 unknown, 258 mislabeled, 583 multi-label) are derived from a single annotator applying guidelines that contain explicitly subjective boundary decisions, for example 'Letter excludes memos using letter in its subject', 'Note excludes Reports with Note in title', and 'ADVE excludes magazine covers and content drafts'. The paper reports no inter-annotator agreement, no independent adjudication, and no estimate of annotation variability. Because every downstream number—11.7%, 16.7%, and the 196/554 correction—inherits this single annotator's judgment, this is a load-bearing methodological gap. Please add an inter-annotator agreement study on a sample (reporting, for example, Cohen's kappa for unknown, mislabeled, and multi-label decisions) with adjudication of disagreements, or alternatively reframe all counts as one annotator's assessment and add a prominent limitations paragraph that removes the implication of ground truth.","section":"§2.2, Table 1"},{"comment":"The '89.7% accuracy' is computed by post-hoc reclassifying 196 of the model's predictions as correct; it is not a model trained or evaluated on a relabeled version of Tobacco3482. The subsequent comparison with DocXclassifier's 90.7% accuracy in [16] is therefore not a controlled comparison, since the latter method was trained and evaluated under its own protocol on the original labels. Please state explicitly that 89.7% is a counterfactual reannotation upper bound, not a measured accuracy from a retrained model, and soften the claim that the corrected score 'brings it closer to' a state-of-the-art method.","section":"§3"},{"comment":"The three problematic label types (unknown, mis-labeled, multi-label) are presented as separate counts, but the paper never states whether these categories are mutually exclusive. In particular, a document whose original label is wrong can have two or more valid alternative labels, making it both 'mis-labeled' and 'multi-label'. As written, the reader cannot compute the unique number of documents with any label issue, and the sum 151 + 258 + 583 = 992 may overstate the total number of affected samples. Please report the pairwise and triple intersections among the three categories and give a unified rate for documents with at least one annotation problem.","section":"§2.2"},{"comment":"The TTID cross-check identifies 134 of 1,707 located documents (roughly 7.8%) as having multiple TTID category annotations, yet the authors' own review finds 583 multi-label samples (16.7% of the dataset). The paper does not reconcile this large discrepancy or use the TTID annotations as an independent partial check on the authors' multi-label counts. A short discussion of whether the difference reflects intentionally narrow TTID annotations, different category schemas, or annotator interpretation would substantially strengthen confidence in the audit.","section":"§2.1"}],"minor_comments":[{"comment":"The sentence 'our findings highlight flaws in the Tobacco3482 dataset concerining data quality' contains a typo: 'concerining' should be 'concerning'.","section":"§1"},{"comment":"The same paragraph reports '84.1% top-1 accuracy' on one line and 'the original 84.0% accuracy' two sentences later; please make the reported baseline accuracy consistent.","section":"§3"},{"comment":"The subsection heading 'Multiple Lables' should be 'Multiple Labels'.","section":"§2.2"},{"comment":"The figure caption begins 'un-problematicproblematic', which appears to be a formatting error and should read 'un-problematic and problematic'.","section":"Figure 1"},{"comment":"For reproducibility, please report the DiT fine-tuning setup used in the 4-fold cross-validation (e.g., whether folds were stratified, number of epochs, learning rate, batch size, and random seeds).","section":"§3"}],"recommendation":"major_revision","confidential_remarks":"The paper is best viewed as a dataset-quality audit rather than a modeling contribution. The single-annotator issue is the central obstacle; the authors' existing release of annotation files is a strength, but it makes the judgments inspectable rather than validated. The self-citation to the authors' prior RVL-CDIP audit is not by itself a problem, but it means the annotation methodology is an in-family standard that has not yet been independently replicated. If the authors can add even a modest inter-annotator reliability study on a sample, or clearly reframe the percentages as one annotator's assessment, the paper would be acceptable; otherwise the title and abstract overstate the certainty of the numerical claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First useful audit of Tobacco3482. The headline numbers—11.7% unknown/mislabeled, 16.7% multi-label, 35% of DiT mistakes not really mistakes—are plausible and internally consistent, and the release of annotations and code is a genuine contribution to a subfield that has been flying blind. The independent check via TTID's original multi-label annotations (3.8% of documents carry multiple category labels there) gives some external support for the multi-label finding, which is more than most audits offer.\n\nThe soft spots are real but not fatal: the entire count rests on one annotator applying the authors' own Table 1 guidelines, with no inter-annotator agreement, no adjudication, and no error bars. Some guidelines are explicitly subjective (e.g., Letter excludes memos using 'letter' in the subject; ADVE excludes magazine covers). A second reasonable annotator could shift the numbers several points. The paper should present these as estimates, not ground truth, and ideally run a second annotator on a sample to bound the uncertainty. The model-impact analysis inherits this fragility, so the 89.7% corrected accuracy should also be labeled as a single-annotator estimate.\n\nThat said, the internal arithmetic checks out (151+258=409=11.7%; 196/554=35.4%), the code and annotations are inspectable, and the qualitative conclusion—that Tobacco3482 has meaningful label noise—is well-supported by the examples and the TTID cross-check. The paper is short and honest about being a first pass.\n\nWho should read it: anyone using Tobacco3482 or RVL-CDIP for evaluation, and people working on dataset auditing methodology. It deserves a serious referee. I'd recommend acceptance with a mandatory revision that adds a reliability analysis or at minimum careful caveats, and possibly a second annotator on a subset to quantify agreement.","headline":"A genuinely useful first audit of Tobacco3482 that gets the direction right; the exact percentages need a reliability caveat because they rest on one annotator's judgment.","tokens_in":6286,"tokens_out":1877,"would_cite":true,"duration_ms":16489,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that 11.7% of Tobacco3482 documents are improperly annotated—151 with no valid category, 258 mislabeled—and that 583 documents (16.7%) have multiple valid labels; re-scoring a strong model's errors with that knowledge…","keywords":["Tobacco3482","label errors","document image classification","annotation quality","multi-label images","benchmark evaluation","model mistakes"],"falsifier":"Have a second annotator, blinded to the original labels and working only from the paper's published category guidelines, re-annotate all 3,482 images independently; if the counts of unknown, mislabeled, and multi-label documents diverge materially from 151, 258, and 583—say, by more than a few percentage points—the paper's headline error rates and its '35% of mistakes' attribution are not stable.","tokens_in":5255,"feed_emoji":"🏷️","tokens_out":9452,"duration_ms":81211,"temperature":0.7,"pith_summary":"The paper tries to establish that Tobacco3482, a standard benchmark for document image classification, is substantially noisier than its users assume. A complete manual re-annotation of all 3,482 images, guided by category definitions the authors wrote down first, yields 151 documents with no valid category, 258 documents assigned to the wrong category, and 583 documents with more than one valid label. The authors then use their re-annotation to re-score the mistakes of a top-performing DiT document transformer: 196 of the model's 554 errors are valid alternative labels or uncategorizable documents, so its accuracy rises from 84.1% to 89.7% when those are counted as correct. If the audit is right, published accuracy figures on Tobacco3482 understate model capability and should not be read as clean measures of classification skill.","feed_headline":"11.7% of Tobacco3482 labels are wrong or unknown","feed_subtitle":"Manual audit finds 583 multi-label images; 35% of a top model's 'errors' are actually correct","key_machinery":"The central object is a re-annotation protocol with three error types: unknown, mislabeled, and multiple labels. The paper first fixes explicit definitions for each of the ten Tobacco3482 categories (for example, a Letter must have physical addressing and a sign-off, while a Memo is addressed with To/From), then applies those definitions to all 3,482 images and records every case where the original single label fails. This protocol is what generates the 151/258/583 counts and the overlap structure between categories, and it is also how the authors convert a model's predicted labels into a verdict of 'valid alternative label' or 'no valid label' when re-scoring the DiT model's mistakes.","core_discovery":"On the paper's own terms, the central discovery is that the label layer of Tobacco3482 is not a single ground truth. The authors inspect every image and report 151 unknown-label cases, 258 wrong-label cases, and 583 multi-label cases, with errors concentrated in categories such as Letter (about 24% mislabeled, mostly documents they would call Memos) and Scientific (about 24% mislabeled and 14.6% unknown). To show the practical effect, they evaluate a DiT model with four-fold cross-validation on the original labels, collect its 554 errors, and find that 147 are predictions of a valid alternative label and 49 are predictions for documents that have no valid label at all. Re-scoring those 196 cases as correct moves the model from 84.1% to 89.7% top-1 accuracy.","pith_inferences":["If the same single-annotator audit were repeated on other document benchmarks drawn from the same source collection, similarly high label-error rates might appear, so reported gains on those benchmarks should be read with caution.","A corrected, multi-label version of Tobacco3482 could serve as a testbed for deciding whether improvements in published accuracy come from better models or from better fit to noisy labels.","Because 49 of the model's 'mistakes' are on documents with no valid category, a practical classifier deployed on such documents would need a rejection or 'other' option; without one, accuracy metrics are unfairly harsh.","One testable implication: a model trained or evaluated with multi-label structure (for example, predicting both Report and Memo) should beat a single-label model on the paper's re-annotated labels even if they tie on the original labels."],"forward_implications":["Reported top-1 accuracy on Tobacco3482 understates model capability; the paper's own DiT model jumps from 84.1% to 89.7% once valid alternative labels and unknown cases are counted as correct.","Benchmark comparisons on the raw dataset can mis-rank models, since a large share of measured errors are annotation artifacts rather than classification failures.","Future evaluations should report multi-label-aware and unknown-aware metrics instead of single-label accuracy on the original labels.","The 583 multi-label documents imply that multi-label classification is a more faithful task for this benchmark than the original single-label task.","Error rates are concentrated in certain categories (Letter, Scientific), so per-category accuracy numbers on the original labels are especially unreliable."],"supporting_citations":[{"why":"Defines the Tobacco3482 dataset, the object of the entire audit.","marker":"[19]"},{"why":"Supplies the re-annotation-with-guidelines method and the precedent of finding label errors in a related document benchmark.","marker":"[7]"},{"why":"The DiT document transformer whose 554 errors on Tobacco3482 are re-scored in the impact analysis.","marker":"[11]"},{"why":"Provides the general argument that label errors in test sets destabilize benchmark measurements.","marker":"[15]"},{"why":"Defines the document-classification benchmark whose pretrained weights the DiT model uses.","marker":"[4]"},{"why":"Describes the source test collection the Tobacco3482 documents come from, used to look up original multi-category annotations.","marker":"[9]"},{"why":"The recent method achieving 90.7% accuracy, the reference point the re-scored DiT accuracy approaches.","marker":"[16]"},{"why":"The orthogonal finding of feature biases in the same datasets, cited to reinforce doubts about published performance as true capability.","marker":"[17]"}],"fun_headline_variants":["Over 400 Tobacco3482 labels are wrong or unknown","35% of top model's errors on Tobacco3482 are actually correct","16.7% of Tobacco3482 have multiple valid labels","Accounting for label errors lifts DiT accuracy to 89.7%","Tobacco3482 label audit: 11.7% mislabeled, 16.7% multi-label"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the paper's self-authored labeling guidelines and the single annotator's application of them are the correct way to decide what each Tobacco3482 document really is, and the paper reports no second annotation or adjudication to test that judgment.","fun_headline_variants_meta":{"raw":{"variants":["Over 400 Tobacco3482 labels are wrong or unknown","35% of top model's errors on Tobacco3482 are actually correct","16.7% of Tobacco3482 have multiple valid labels","Accounting for label errors lifts DiT accuracy to 89.7%","Tobacco3482 label audit: 11.7% mislabeled, 16.7% multi-label"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00132,"raw_usage":{"total_tokens":5326,"prompt_tokens":847,"completion_tokens":4479,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":463,"completion_tokens_details":{"reasoning_tokens":4389}},"tokens_in":463,"tokens_out":4479,"duration_ms":31259,"temperature":1.0,"reasoning_tokens":4389,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T13:21:58.432297+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a second annotator, blinded to the original labels and working only from the paper's published category guidelines, re-annotate all 3,482 images independently; if the counts of unknown, mislabeled, and multi-label documents diverge materially from 151, 258, and 583—say, by more than a few percentage points—the paper's headline error rates and its '35% of mistakes' attribution are not stable.","supporting_citations":[{"cited_title":"Automatic document logo detection","cited_arxiv_id":null,"evidence_quote":"Defines the Tobacco3482 dataset, the object of the entire audit."},{"cited_title":"On evaluation of document classification with RVL-CDIP","cited_arxiv_id":null,"evidence_quote":"Supplies the re-annotation-with-guidelines method and the precedent of finding label errors in a related document benchmark."},{"cited_title":"DiT: Self-supervised pre-training for docu- ment image transformer","cited_arxiv_id":null,"evidence_quote":"The DiT document transformer whose 554 errors on Tobacco3482 are re-scored in the impact analysis."},{"cited_title":"Northcutt, Anish Athalye, and Jonas Mueller","cited_arxiv_id":null,"evidence_quote":"Provides the general argument that label errors in test sets destabilize benchmark measurements."},{"cited_title":"Harley, Alex Ufkes, and Konstantinos G","cited_arxiv_id":null,"evidence_quote":"Defines the document-classification benchmark whose pretrained weights the DiT model uses."},{"cited_title":"Grossman, and Jefferson Heard","cited_arxiv_id":null,"evidence_quote":"Describes the source test collection the Tobacco3482 documents come from, used to look up original multi-category annotations."},{"cited_title":"DocXclassifier: towards a robust and interpretable deep neu- ral network for document image classification","cited_arxiv_id":null,"evidence_quote":"The recent method achieving 90.7% accuracy, the reference point the re-scored DiT accuracy approaches."},{"cited_title":"The reality of high performing deep learning models: A case study on document image classification","cited_arxiv_id":null,"evidence_quote":"The orthogonal finding of feature biases in the same datasets, cited to reinforce doubts about published performance as true capability."}],"review_version":1}