{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2019:EMCX3KNYQSYAMONWQMUKITMAQZ","short_pith_number":"pith:EMCX3KNY","schema_version":"1.0","canonical_sha256":"23057da9b884b00639b68328a44d80867d59be79695e1c06d66dbe2e3142337e","source":{"kind":"arxiv","id":"1907.11692","version":1},"attestation_state":"computed","paper":{"title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"A careful retraining of BERT — longer, on more data, with dynamic masking and no next-sentence loss — matches or beats every model published after it on GLUE, SQuAD, and RACE.","cross_cats":[],"primary_cat":"cs.CL","authors_text":"Danqi Chen, Jingfei Du, Luke Zettlemoyer, Mandar Joshi, Mike Lewis, Myle Ott, Naman Goyal, Omer Levy, Veselin Stoyanov, Yinhan Liu","submitted_at":"2019-07-26T17:48:29Z","abstract_excerpt":"Language model pretraining has led to significant performance gains but careful comparison between different approaches is challenging. Training is computationally expensive, often done on private datasets of different sizes, and, as we will show, hyperparameter choices have significant impact on the final results. We present a replication study of BERT pretraining (Devlin et al., 2019) that carefully measures the impact of many key hyperparameters and training data size. We find that BERT was significantly undertrained, and can match or exceed the performance of every model published after it"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":true,"formal_links_present":true},"canonical_record":{"source":{"id":"1907.11692","kind":"arxiv","version":1},"metadata":{"license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","primary_cat":"cs.CL","submitted_at":"2019-07-26T17:48:29Z","cross_cats_sorted":[],"title_canon_sha256":"a6658a1fd9390b3fb8d3fcc8e7edeaea97c38c2de7818cf45883a5ff37d20dc4","abstract_canon_sha256":"28bcebc417de7b07736f5b8236aee8283e3f9f07155471aeb3002e6eb8878753"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-04T23:49:50.769069Z","signature_b64":"4iAsrntTSgrIB+p2U9zW8b1Pkvq8Y416cxMjI9433tlfELo/xOCjXn7dsyLdeV0jarfq+jZ8PQqbenLEBPd9Dg==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"23057da9b884b00639b68328a44d80867d59be79695e1c06d66dbe2e3142337e","last_reissued_at":"2026-07-04T23:49:50.768566Z","signature_status":"signed_v1","first_computed_at":"2026-07-04T23:49:50.768566Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"A careful retraining of BERT — longer, on more data, with dynamic masking and no next-sentence loss — matches or beats every model published after it on GLUE, SQuAD, and RACE.","cross_cats":[],"primary_cat":"cs.CL","authors_text":"Danqi Chen, Jingfei Du, Luke Zettlemoyer, Mandar Joshi, Mike Lewis, Myle Ott, Naman Goyal, Omer Levy, Veselin Stoyanov, Yinhan Liu","submitted_at":"2019-07-26T17:48:29Z","abstract_excerpt":"Language model pretraining has led to significant performance gains but careful comparison between different approaches is challenging. Training is computationally expensive, often done on private datasets of different sizes, and, as we will show, hyperparameter choices have significant impact on the final results. We present a replication study of BERT pretraining (Devlin et al., 2019) that carefully measures the impact of many key hyperparameters and training data size. We find that BERT was significantly undertrained, and can match or exceed the performance of every model published after it"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Under controlled comparison, BERT's masked-language-modeling objective with the original architecture, when trained longer on more data with larger batches, dynamic masking, no NSP loss, and byte-level BPE, matches or exceeds the downstream performance of every published post-BERT method (XLNet, SpanBERT, MT-DNN, etc.) on GLUE, SQuAD, and RACE — implying that previously reported gains over BERT are substantially attributable to training budget rather than architectural or objective novelty.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That holding \"architecture and objective\" fixed while varying data, steps, batch size, and masking constitutes a fair attribution of credit. The XLNet comparison in particular conflates multiple axes (RoBERTa uses 160GB vs. XLNet's 126GB, different step counts, different vocabularies), and the authors acknowledge they did not retune XLNet under matched compute. The claim that MLM is \"competitive\" with permutation LM rests on this, and the paper itself notes (footnote 2) that other methods could likely also improve with more tuning.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"With better hyperparameters, more data, and longer training, an unchanged BERT-Large architecture matches or exceeds XLNet and other successors on GLUE, SQuAD, and RACE.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"A careful retraining of BERT — longer, on more data, with dynamic masking and no next-sentence loss — matches or beats every model published after it on GLUE, SQuAD, and RACE.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"65d31102b27e2e035ef5aca6c91d7f6f0917d17497f16d06af4df5920b14e84b"},"source":{"id":"1907.11692","kind":"arxiv","version":1},"verdict":{"id":"ec2cff2e-7d40-46ff-bb02-1922b601c8b2","model_set":{"reader":"claude-opus-4-7"},"created_at":"2026-05-09T01:10:08.388354Z","strongest_claim":"Under controlled comparison, BERT's masked-language-modeling objective with the original architecture, when trained longer on more data with larger batches, dynamic masking, no NSP loss, and byte-level BPE, matches or exceeds the downstream performance of every published post-BERT method (XLNet, SpanBERT, MT-DNN, etc.) on GLUE, SQuAD, and RACE — implying that previously reported gains over BERT are substantially attributable to training budget rather than architectural or objective novelty.","one_line_summary":"With better hyperparameters, more data, and longer training, an unchanged BERT-Large architecture matches or exceeds XLNet and other successors on GLUE, SQuAD, and RACE.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That holding \"architecture and objective\" fixed while varying data, steps, batch size, and masking constitutes a fair attribution of credit. The XLNet comparison in particular conflates multiple axes (RoBERTa uses 160GB vs. XLNet's 126GB, different step counts, different vocabularies), and the authors acknowledge they did not retune XLNet under matched compute. The claim that MLM is \"competitive\" with permutation LM rests on this, and the paper itself notes (footnote 2) that other methods could likely also improve with more tuning.","pith_extraction_headline":"A careful retraining of BERT — longer, on more data, with dynamic masking and no next-sentence loss — matches or beats every model published after it on GLUE, SQuAD, and RACE."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/1907.11692/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":51,"sample":[{"doi":"","year":2007,"title":"Eneko Agirre, Llu' i s M`arquez, and Richard Wicentowski, editors. 2007. Proceedings of the Fourth International Workshop on Semantic Evaluations (SemEval-2007)","work_id":"e1e558e1-e4d3-45df-96aa-ce583b1efdb5","ref_index":1,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2019,"title":"Cloze-driven Pretraining of Self-attention Networks","work_id":"bfe64b17-f035-49cc-a296-5a20504947f4","ref_index":2,"cited_arxiv_id":"1903.07785","is_internal_anchor":false},{"doi":"","year":2006,"title":"Roy Bar-Haim, Ido Dagan, Bill Dolan, Lisa Ferro, Danilo Giampiccolo, Bernardo Magnini, and Idan Szpektor. 2006. The second PASCAL recognising textual entailment challenge. In Proceedings of the second","work_id":"360397b4-2aff-48ff-bfa2-a1b84fdb8f01","ref_index":3,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2009,"title":"Luisa Bentivogli, Ido Dagan, Hoa Trang Dang, Danilo Giampiccolo, and Bernardo Magnini. 2009. The fifth PASCAL recognizing textual entailment challenge","work_id":"53ae3d9b-cd0a-42c8-9b72-30f5109721dc","ref_index":4,"cited_arxiv_id":"","is_internal_anchor":false},{"doi":"","year":2015,"title":"Samuel R Bowman, Gabor Angeli, Christopher Potts, and Christopher D Manning. 2015. A large annotated corpus for learning natural language inference. In Empirical Methods in Natural Language Processing","work_id":"4e74678c-96e0-4ad9-8fb6-b307bd5276a8","ref_index":5,"cited_arxiv_id":"","is_internal_anchor":false}],"resolved_work":51,"snapshot_sha256":"d0bfed0f2f4ef9471d7b8ba9433b35f069d8647063c604d6e873b40b1bbd390d","internal_anchors":3},"formal_canon":{"evidence_count":1,"snapshot_sha256":"af265921ee5818a773a28788bec339f08bdcbc93e626184249e020798fe44e1c"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"1907.11692","created_at":"2026-07-04T23:49:50.768628+00:00"},{"alias_kind":"arxiv_version","alias_value":"1907.11692v1","created_at":"2026-07-04T23:49:50.768628+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.1907.11692","created_at":"2026-07-04T23:49:50.768628+00:00"},{"alias_kind":"pith_short_12","alias_value":"EMCX3KNYQSYA","created_at":"2026-07-04T23:49:50.768628+00:00"},{"alias_kind":"pith_short_16","alias_value":"EMCX3KNYQSYAMONW","created_at":"2026-07-04T23:49:50.768628+00:00"},{"alias_kind":"pith_short_8","alias_value":"EMCX3KNY","created_at":"2026-07-04T23:49:50.768628+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":575,"internal_anchor_count":575,"sample":[{"citing_arxiv_id":"2607.05689","citing_title":"UCSC NLP at SemEval-2026 Task 10: Boundary-Aware Span Extraction and RoBERTa Classification for Conspiracy Detection","ref_index":1,"is_internal_anchor":true},{"citing_arxiv_id":"2607.07159","citing_title":"Recovering Latent Structures after Variational Bayesian Variable Selection: Fit Assessment and Factor-Number Selection in Partially Exploratory Factor Analysis","ref_index":25,"is_internal_anchor":true},{"citing_arxiv_id":"2607.07527","citing_title":"A Unified Detection Framework for AI-Related Content and Artifacts","ref_index":29,"is_internal_anchor":true},{"citing_arxiv_id":"2607.07573","citing_title":"Multi-Class vs. Multi-Label BERT for CVE-to-CWE Mapping: How Taxonomy Structure Shapes the Errors","ref_index":11,"is_internal_anchor":true},{"citing_arxiv_id":"2607.07611","citing_title":"Asymmetric Focal Loss Improves Graph Neural Network Prediction of Drug-Drug Interactions","ref_index":17,"is_internal_anchor":true},{"citing_arxiv_id":"2607.05908","citing_title":"Drift Happens: An Empirical Study of Neural Architecture Robustness to Temporal Distribution Shift","ref_index":27,"is_internal_anchor":true},{"citing_arxiv_id":"2607.05937","citing_title":"Is Domain Adaptation Always Helpful? A Frozen-Backbone Study of Cross-Domain Sentiment Transfer","ref_index":2,"is_internal_anchor":true},{"citing_arxiv_id":"2607.06213","citing_title":"FDIFormer:Protocol-Aware Transformer Learning for False Data Injection Attack Detection in Smart Grid Networks","ref_index":27,"is_internal_anchor":true},{"citing_arxiv_id":"2606.07522","citing_title":"Community-Specific Slang and Entity Detection via Semantic Shift in Fine-Tuned Language Models","ref_index":5,"is_internal_anchor":true},{"citing_arxiv_id":"2607.00004","citing_title":"Why Advanced Encoders Lag on Sparse Retrieval? The Answer and an Approach to Bridging Vocabulary Gaps","ref_index":34,"is_internal_anchor":true},{"citing_arxiv_id":"2604.21534","citing_title":"UKP_Psycontrol at SemEval-2026 Task 2: Modeling Valence and Arousal Dynamics from Text","ref_index":3,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26062","citing_title":"When Certainty Is an Artifact: Keyword Lexicon Blindness and the (Mis)Measurement of Rhetorical Stance","ref_index":5,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25343","citing_title":"Invoice Haystack: Benchmarking Document Retrieval and Visual Question Answering Under Strong Visual Homogeneity","ref_index":46,"is_internal_anchor":true},{"citing_arxiv_id":"2606.25325","citing_title":"Omni-Perception Policy Optimization for Multimodal Emotion Reasoning","ref_index":116,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24331","citing_title":"Transformer-Based Language Models Across Domain Verticals: Architectures, Applications and Critical Assessment","ref_index":17,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24259","citing_title":"SURGELLM: Rethinking Multi-Task Evaluation through Task-Aware Feature Gating with Class-Balanced Normalization","ref_index":6,"is_internal_anchor":true},{"citing_arxiv_id":"2606.27342","citing_title":"Understanding Domain-Aware Distribution Alignment in Budgeted Entity Matching","ref_index":9,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26613","citing_title":"EcoTable: Cost-effective Table Integration in Data Lakes for Natural Language Queries","ref_index":44,"is_internal_anchor":true},{"citing_arxiv_id":"2606.27215","citing_title":"Vulnerability of Natural Language Classifiers to Evolutionary Generated Adversarial Text","ref_index":22,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26994","citing_title":"Event-Aware Instructed Assistant for Referring Video Segmentation","ref_index":45,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26749","citing_title":"Structure Before Collapse: Transient semantic geometry in next-token prediction","ref_index":168,"is_internal_anchor":true},{"citing_arxiv_id":"2606.27274","citing_title":"BetXplain: An Explanation-Annotated Dataset for Detecting Manipulative Betting Advertisements on Social Media","ref_index":2,"is_internal_anchor":true},{"citing_arxiv_id":"2606.26489","citing_title":"Comparing BERT Sentence-Pair Classification and Few-Shot LLM Prompting for Detecting Threat and Solution Framing in German Climate News","ref_index":18,"is_internal_anchor":true},{"citing_arxiv_id":"2606.23919","citing_title":"Unified Multi-Task Relevance Modeling for E-Commerce: Comparing Task Routing Architectures Across LLMs and Cross-Encoders","ref_index":33,"is_internal_anchor":true},{"citing_arxiv_id":"2606.24937","citing_title":"The Hitchhiker's Guide to Agentic AI: From Foundations to Systems","ref_index":38,"is_internal_anchor":true}]},"formal_canon":{"evidence_count":1,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/EMCX3KNYQSYAMONWQMUKITMAQZ","json":"https://pith.science/pith/EMCX3KNYQSYAMONWQMUKITMAQZ.json","graph_json":"https://pith.science/api/pith-number/EMCX3KNYQSYAMONWQMUKITMAQZ/graph.json","events_json":"https://pith.science/api/pith-number/EMCX3KNYQSYAMONWQMUKITMAQZ/events.json","paper":"https://pith.science/paper/EMCX3KNY"},"agent_actions":{"view_html":"https://pith.science/pith/EMCX3KNYQSYAMONWQMUKITMAQZ","download_json":"https://pith.science/pith/EMCX3KNYQSYAMONWQMUKITMAQZ.json","view_paper":"https://pith.science/paper/EMCX3KNY","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=1907.11692&json=true","fetch_graph":"https://pith.science/api/pith-number/EMCX3KNYQSYAMONWQMUKITMAQZ/graph.json","fetch_events":"https://pith.science/api/pith-number/EMCX3KNYQSYAMONWQMUKITMAQZ/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/EMCX3KNYQSYAMONWQMUKITMAQZ/action/timestamp_anchor","attest_storage":"https://pith.science/pith/EMCX3KNYQSYAMONWQMUKITMAQZ/action/storage_attestation","attest_author":"https://pith.science/pith/EMCX3KNYQSYAMONWQMUKITMAQZ/action/author_attestation","sign_citation":"https://pith.science/pith/EMCX3KNYQSYAMONWQMUKITMAQZ/action/citation_signature","submit_replication":"https://pith.science/pith/EMCX3KNYQSYAMONWQMUKITMAQZ/action/replication_record"}},"created_at":"2026-07-04T23:49:50.768628+00:00","updated_at":"2026-07-04T23:49:50.768628+00:00"}