{"paper":{"title":"Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Rule-VLN adds 177 regulatory categories to a 29k-node urban graph to test whether navigation agents can obey semantic rules instead of only reaching goals.","cross_cats":["cs.CV","cs.RO"],"primary_cat":"cs.AI","authors_text":"Jiawen Wen, Penglei Sun, Suixuan Qiu, Weisheng Xu, Wenjie Zhang, Xiaofei Yang, Xiaowen Chu","submitted_at":"2026-04-18T13:41:46Z","abstract_excerpt":"As embodied AI transitions to real-world deployment, the success of the Vision-and-Language Navigation (VLN) task tends to evolve from mere reachability to social compliance. However, current agents suffer from a \"goal-driven trap\", prioritizing physical geometry (\"can I go?\") over semantic rules (\"may I go?\"), frequently overlooking subtle regulatory constraints. To bridge this gap, we establish Rule-VLN, the first large-scale urban benchmark for rule-compliant navigation. Spanning a massive 29k-node environment, it injects 177 diverse regulatory categories into 8k constrained nodes across fo"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Experiments demonstrate that while Rule-VLN challenges state-of-the-art models, SNRM significantly restores navigation capabilities, reducing CVR by 19.26% and boosting TC by 5.97%.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the 177 injected regulatory categories and the coarse-to-fine VLM perception in SNRM accurately capture and interpret real-world semantic and behavioral constraints in a zero-shot manner without domain-specific fine-tuning or additional supervision.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Rule-VLN is the first large-scale benchmark injecting 177 regulatory categories into an urban environment, and the proposed SNRM module equips pre-trained VLN agents with zero-shot semantic reasoning and detour planning to reduce constraint violations by 19.26% and improve task completion.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Rule-VLN adds 177 regulatory categories to a 29k-node urban graph to test whether navigation agents can obey semantic rules instead of only reaching goals.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"97874106a3b60e225dfadccf103fc5531da9af28abd310de2b1080450f3e8bc7"},"source":{"id":"2604.16993","kind":"arxiv","version":2},"verdict":{"id":"a2eeabda-b463-4617-9906-524dea7ee88d","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-10T06:52:20.124678Z","strongest_claim":"Experiments demonstrate that while Rule-VLN challenges state-of-the-art models, SNRM significantly restores navigation capabilities, reducing CVR by 19.26% and boosting TC by 5.97%.","one_line_summary":"Rule-VLN is the first large-scale benchmark injecting 177 regulatory categories into an urban environment, and the proposed SNRM module equips pre-trained VLN agents with zero-shot semantic reasoning and detour planning to reduce constraint violations by 19.26% and improve task completion.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the 177 injected regulatory categories and the coarse-to-fine VLM perception in SNRM accurately capture and interpret real-world semantic and behavioral constraints in a zero-shot manner without domain-specific fine-tuning or additional supervision.","pith_extraction_headline":"Rule-VLN adds 177 regulatory categories to a 29k-node urban graph to test whether navigation agents can obey semantic rules instead of only reaching goals."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2604.16993/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}