{"paper":{"title":"Grounding Multi-Hop Reasoning in Structural Causal Models via Group Relative Policy Optimization","license":"http://creativecommons.org/licenses/by/4.0/","headline":"Grounding multi-hop fact verification in a structural causal model and optimizing it with group relative policy optimization yields more accurate and less hallucinated reasoning than standard chain-of-thought methods.","cross_cats":[],"primary_cat":"cs.AI","authors_text":"Askar Hamdulla, Baohua Zhang, Chunxiao Gao, Guotong Geng, Huaping Zhang, Juan Wang, Qiuchi Li, Quan Zhang, Shuai Lei, Yunbo Cao, Yunhan Bu, Zhunchen Luo","submitted_at":"2026-05-02T15:05:38Z","abstract_excerpt":"Multi-Hop Fact Verification requires complex reasoning across disparate evidence, posing significant challenges for Large Language Models , which may suffer from hallucinations and fractured logical chains. Existing methods, while improving transparency via Chain-of-Thought , often lack explicit modeling of the structural dependencies between evidence and claims. In this work, we introduce an SCM-inspired framework that grounds reasoning in explicit directed dependency graphs, treating verification as a constructive structural reasoning process rather than full causal inference with interventi"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Extensive experiments on HoVer and EX-FEVER demonstrate that our SCM-GRPO framework significantly outperforms state-of-the-art baselines, offering a reliable and interpretable solution for complex fact verification.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That explicitly modeling verification as a constructive causal inference process inside a structural causal model will produce more accurate and less hallucinated reasoning chains than standard chain-of-thought methods without introducing new modeling errors.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"SCM-GRPO grounds multi-hop fact verification in structural causal models and applies GRPO reinforcement learning to optimize reasoning chain length, outperforming baselines on HoVer and EX-FEVER.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Grounding multi-hop fact verification in a structural causal model and optimizing it with group relative policy optimization yields more accurate and less hallucinated reasoning than standard chain-of-thought methods.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"636cbfab7982b04d7adc59da33ab930f5dad07a1306fd149fdd29a6729eb74b1"},"source":{"id":"2605.01482","kind":"arxiv","version":3},"verdict":{"id":"28a751db-2df2-43aa-9963-4ae9fb9fa3ef","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-11T02:12:57.703359Z","strongest_claim":"Extensive experiments on HoVer and EX-FEVER demonstrate that our SCM-GRPO framework significantly outperforms state-of-the-art baselines, offering a reliable and interpretable solution for complex fact verification.","one_line_summary":"SCM-GRPO grounds multi-hop fact verification in structural causal models and applies GRPO reinforcement learning to optimize reasoning chain length, outperforming baselines on HoVer and EX-FEVER.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That explicitly modeling verification as a constructive causal inference process inside a structural causal model will produce more accurate and less hallucinated reasoning chains than standard chain-of-thought methods without introducing new modeling errors.","pith_extraction_headline":"Grounding multi-hop fact verification in a structural causal model and optimizing it with group relative policy optimization yields more accurate and less hallucinated reasoning than standard chain-of-thought methods."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2605.01482/integrity.json","findings":[],"available":true,"detectors_run":[{"name":"ai_meta_artifact","ran_at":"2026-05-20T17:41:07.373009Z","status":"completed","version":"1.0.0","findings_count":0},{"name":"doi_compliance","ran_at":"2026-05-19T17:14:37.062627Z","status":"completed","version":"1.0.0","findings_count":0}],"snapshot_sha256":"24620f8e0b5165dab834ba9b01c0cec52100a6cea05b6ea10314ac478cc9440a"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":2,"snapshot_sha256":"e7e4508c0f2eaf08a47bda52dfce33fbe70217368bacb5502d074829a5e3046f"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}