Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:25:31.762978Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2412.08972.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T17:25:31.762978Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T12:58:02.335060Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-11T12:58:02.541932Z
51 of 51 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fe74d52f-9e9a-4981-9381-23d17cb9ce5a · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9c95e2b-5f9e-46d6-bf7c-2a2d276398b7 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 599ee7b7-00f5-4437-a3b6-b38fe3bb4ac1 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 472af47e-a485-4b67-bd82-325ffc628e56 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8c645659-9cea-455c-896a-efcc98ce5329 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9fcdded-cafc-41bd-aeb4-2870f989629b · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Benchmarking Large Language Models on Controllable Generation under Diversified Instructions
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b3ab5312-a0a8-4a93-8215-efe0af9a4d2f · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Training Verifiers to Solve Math Word Problems
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75cfa14a-b8f9-4487-896f-7ae1df5c893d · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios A Survey on In-context Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab56d529-30ee-4481-861e-93a4c9952e9d · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cef593c2-1dbc-4197-adc3-f9923440cff7 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcd5b8b4-e197-468e-b98f-c474912133a6 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Specializing Smaller Language Models towards Multi-Step Reasoning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faa6572b-b8ad-4e9f-8cc2-f61ace86641f · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Neural Module Networks for Reasoning over Text
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80815c01-31e8-47f4-862e-111eaa9ec44e · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios FOLIO: Natural Language Reasoning with First-Order Logic
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a36dd340-b6a0-4515-8470-c043cb3bfde8 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Can Large Language Models Understand Real-World Complex Instructions?
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b3b2d29-82b6-4140-8d85-e1d892dbcc71 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Measuring Mathematical Problem Solving With the MATH Dataset
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dad0c5e5-51b0-4dea-a78e-47e6e0f51821 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Disentangling Logic: The Role of Context in Large Language Model Reasoning Capabilities
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bcfb457-4e42-403d-b33e-24e07603a894 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Fine-tuning and Utilization Methods of Domain-specific LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84047906-9d58-4a8b-b19f-970db3f7ca96 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 554c006a-5a8d-431e-90e0-d7d6e714a2b1 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Large Language Models are Zero-Shot Reasoners
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c110deb-f1f6-4d3d-b68f-216350ef4e65 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e8113bb3-d5cd-4f4a-aaa8-2ac2ecb38777 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66f26733-5eae-44f1-b1d6-7627e89aeeab · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios AlignBench: Benchmarking Chinese Alignment of Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3cafca9-863a-4946-9151-f0ca0cfdaa72 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e8f1c96-c593-444a-ba49-79d753438df8 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios GPT-4 Technical Report
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7fe685c-4e24-41a8-9943-9522fe5cded6 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a2997379-5792-4bc2-ba04-86a4b9da5cdb · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Can LLMs Follow Simple Rules?
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d07f299e-107c-41cc-a9d8-1f3fb7fbe234 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54d7ec07-72ff-4892-8fef-63be4a50f400 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29a4a05b-dc45-474b-8d11-7feff824a42e · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Evaluating Large Language Models on Controlled Generation Tasks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff534545-f620-4513-9392-6df483194e48 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93c81ae0-192f-4a1e-a76b-4770122d047f · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios ProofWriter: Generating Implications, Proofs, and Abductive Statements over Natural Language
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cfbde95-e41a-4638-a9fc-2b37c86c2a4e · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecbe2bb7-1901-4741-8d75-912bb2e577b5 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Gemini: A Family of Highly Capable Multimodal Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0baaebf-292d-4d06-a7f2-7c8db95538c2 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios LLaMA: Open and Efficient Foundation Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1b92c73-8178-4984-90f9-730491d49f0c · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c88a85ba-37a6-4566-a606-1ce458d5726c · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Symbolic Working Memory Enhances Language Models for Complex Rule Application
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c782645d-ffea-4ab0-b09b-eed7f272ade5 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a85ad85-4774-4a3b-9fa8-dffe70884ab0 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Larger language models do in-context learning differently
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cdb2072-3836-4bf3-93a8-6fbb1f18f0f0 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Benchmarking Complex Instruction-Following with Multiple Constraints Composition
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89f1b186-2986-43fc-b091-9e364538064d · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5c57647e-2132-4c7d-9d04-3f09215c7e4b · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ffddade-a12f-4d0c-822a-75a995be22b8 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9aa2e2f-1cc9-4ae8-9b06-958f5effb3e4 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8f55908-4f7a-4154-a59b-33998ccec510 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 844f10de-a47f-4635-bcba-4a3297884042 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios IDEAL: Influence-Driven Selective Annotations Empower In-Context Learners in Large Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 384d583b-315a-4a8a-adf5-48678e0a6465 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Supervised Chain of Thought
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73562581-8cd3-4c46-b5f0-be198b988482 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85d4fd0c-8c37-4866-b2a1-3e3105cc69c4 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios AR-LSAT: Investigating Analytical Reasoning of Text
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74a6d8bb-5218-4579-9c27-2f087c4131ca · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Instruction-Following Evaluation for Large Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab0a5149-1d39-483d-bf07-d5ecd39f7b4f · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios TRAD: Enhancing LLM Agents with Step-Wise Thought Retrieval and Aligned Decision
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b10c5b1-eb13-437b-a2d3-e8c4a2892122 · outbound
RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd2a529d-24b1-45f7-8885-b74ad43c7ffc · inbound
AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.