{"paper":{"title":"Diffusion-Inspired Masked Fine-Tuning for Knowledge Injection in Autoregressive LLMs","license":"http://creativecommons.org/licenses/by-nc-sa/4.0/","headline":"Masked fine-tuning lets autoregressive LLMs absorb new facts without paraphrases and without reversal-curse failures.","cross_cats":["cs.AI"],"primary_cat":"cs.CL","authors_text":"Ely Hahami, Haim Sompolinsky, Jingxuan Fan, Xu Pan, Ziqian Xie","submitted_at":"2025-10-10T21:43:50Z","abstract_excerpt":"Large language models (LLMs) are often used in environments where facts evolve, yet factual knowledge updates via fine-tuning on unstructured text often suffer from 1) reliance on compute-heavy paraphrasing augmentation and 2) the reversal curse. Recent studies show diffusion large language models (dLLMs) require fewer training samples to achieve lower loss in pre-training and are more resistant to the reversal curse, suggesting dLLMs may learn new knowledge more easily than autoregressive LLMs (arLLMs). We test this hypothesis in controlled knowledge fine-tuning experiments and find that whil"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"The masked fine-tuning for arLLMs substantially improves the efficacy of knowledge injection, i.e. no paraphrase needed and resistant to the reversal curse, closing the gap between arLLMs and dLLMs.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the demasking objective alone induces the observed knowledge-injection advantage in arLLMs independent of diffusion-specific architecture details, and that the controlled experiments isolate this effect without confounding differences in model scale, data distribution, or masking implementation.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Masked fine-tuning enables autoregressive LLMs to inject new factual knowledge without paraphrases and with reversal-curse resistance, matching diffusion LLM advantages on QA tasks.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Masked fine-tuning lets autoregressive LLMs absorb new facts without paraphrases and without reversal-curse failures.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"d9d6daa9402ffd87a5266e86cad8b4d6e476d3026019d589a317f4fe0cdb35bf"},"source":{"id":"2510.09885","kind":"arxiv","version":6},"verdict":{"id":"4305538a-0d9a-4d9c-855d-0ec86d924738","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-18T07:18:49.407331Z","strongest_claim":"The masked fine-tuning for arLLMs substantially improves the efficacy of knowledge injection, i.e. no paraphrase needed and resistant to the reversal curse, closing the gap between arLLMs and dLLMs.","one_line_summary":"Masked fine-tuning enables autoregressive LLMs to inject new factual knowledge without paraphrases and with reversal-curse resistance, matching diffusion LLM advantages on QA tasks.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the demasking objective alone induces the observed knowledge-injection advantage in arLLMs independent of diffusion-specific architecture details, and that the controlled experiments isolate this effect without confounding differences in model scale, data distribution, or masking implementation.","pith_extraction_headline":"Masked fine-tuning lets autoregressive LLMs absorb new facts without paraphrases and without reversal-curse failures."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2510.09885/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}