Pith. sign in

Paper Citation Record · LEDGER

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios

As of 12 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2412.08972.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08972 v2

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:25:31.762978Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:58:02.335060Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T12:58:02.541932Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe74d52f-9e9a-4981-9381-23d17cb9ce5a · outbound

This paper cites online" 'onlinestring :=.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.466590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.466590Z digest=sha256:c99ad9817cd6514bf185a14492c6568448a9af9d0e014c8c708942aec4eba766

Observation c9c95e2b-5f9e-46d6-bf7c-2a2d276398b7 · outbound

This paper cites write newline.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.472716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.472716Z digest=sha256:a3cd86f92727f398c9f2ae17ae04a71dc59109e1193b3490936fb41fd76f6d8f

Observation 599ee7b7-00f5-4437-a3b6-b38fe3bb4ac1 · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.478406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.478406Z digest=sha256:0d34d3e7fc2389340896ec1bf31478f05724072f881f9eb505ee34f740e9ea24

Observation 472af47e-a485-4b67-bd82-325ffc628e56 · outbound

This paper cites an unresolved cited work.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:25:32.634419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:25:31.484261Z digest=sha256:9493763932dfb1601cd2458b97491dafc50326e2431086baa489e1ed230c0c6a

Observation 8c645659-9cea-455c-896a-efcc98ce5329 · outbound

This paper cites Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.490688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.490688Z digest=sha256:a2eb9dda9c8ed349096611658755a120a0245849fe181a81a6947f93c6280d0c

Observation c9fcdded-cafc-41bd-aeb4-2870f989629b · outbound

This paper cites Benchmarking Large Language Models on Controllable Generation under Diversified Instructions.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Benchmarking Large Language Models on Controllable Generation under Diversified Instructions

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:25:32.521547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:25:31.502542Z digest=sha256:5a803d29463919e81b0fe6b166e4b0f8e9c445800a300423baa59c6a2896f1d7

Observation b3ab5312-a0a8-4a93-8215-efe0af9a4d2f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.507934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.507934Z digest=sha256:d103b12808219c17bef5da67430aae32bd1566058b75d9a7a178e7adcddb68d4

Observation 75cfa14a-b8f9-4487-896f-7ae1df5c893d · outbound

This paper cites A Survey on In-context Learning.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios A Survey on In-context Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.514193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.514193Z digest=sha256:9321797cceaeffcdf6d0d4c1f7ad5c0b227ce73b7d21825d3f531f3631379f7f

Observation ab56d529-30ee-4481-861e-93a4c9952e9d · outbound

This paper cites The Llama 3 Herd of Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.522554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.522554Z digest=sha256:481a66714de796fd8d1e7f0376b5cc7cfb39ac6c146c5eb462f24676124278fd

Observation cef593c2-1dbc-4197-adc3-f9923440cff7 · outbound

This paper cites NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios NPHardEval: Dynamic Benchmark on Reasoning Ability of Large Language Models via Complexity Classes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.528740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.528740Z digest=sha256:f92d57ffa4e7baefb1ce21784b696aeea7c789e96a2b7505c7fd9cb435444a8f

Observation dcd5b8b4-e197-468e-b98f-c474912133a6 · outbound

This paper cites Specializing Smaller Language Models towards Multi-Step Reasoning.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Specializing Smaller Language Models towards Multi-Step Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.534205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.534205Z digest=sha256:edba6fed26d6dbf4979cf1eb0c450dedc64c3ea3ab35110ecc855f937e1d51d5

Observation faa6572b-b8ad-4e9f-8cc2-f61ace86641f · outbound

This paper cites Neural Module Networks for Reasoning over Text.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Neural Module Networks for Reasoning over Text

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.539635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.539635Z digest=sha256:8db7757f79fa0243e5b2c7395388376f68d1e6021a2e19f36fcdfba522027daa

Observation 80815c01-31e8-47f4-862e-111eaa9ec44e · outbound

This paper cites FOLIO: Natural Language Reasoning with First-Order Logic.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios FOLIO: Natural Language Reasoning with First-Order Logic

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.545315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.545315Z digest=sha256:b90c844d3452c1cf8bb7535370614e5868c21c9d58aeba5a1286de9d060521e1

Observation a36dd340-b6a0-4515-8470-c043cb3bfde8 · outbound

This paper cites Can Large Language Models Understand Real-World Complex Instructions?.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Can Large Language Models Understand Real-World Complex Instructions?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.551256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.551256Z digest=sha256:7908dd593e5a5a54b2ea0073ba0ac1d8cf7e59a90dbf1931b7502f641384cbe8

Observation 3b3b2d29-82b6-4140-8d85-e1d892dbcc71 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.556838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.556838Z digest=sha256:105d1f3ef5798f171dac22491b0fb529d1ef00057ecf355790148273732abc7d

Observation dad0c5e5-51b0-4dea-a78e-47e6e0f51821 · outbound

This paper cites Disentangling Logic: The Role of Context in Large Language Model Reasoning Capabilities.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Disentangling Logic: The Role of Context in Large Language Model Reasoning Capabilities

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.562964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.562964Z digest=sha256:c71d2783a35afaa9e97a83d8f3888a3fded31867f9dcad7ac0f757de6b97c50d

Observation 1bcfb457-4e42-403d-b33e-24e07603a894 · outbound

This paper cites Fine-tuning and Utilization Methods of Domain-specific LLMs.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Fine-tuning and Utilization Methods of Domain-specific LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.568737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.568737Z digest=sha256:6c24765825ecaaebfb4583c9b46756c403bf5490ae6e1422c3fa5276be58d94d

Observation 84047906-9d58-4a8b-b19f-970db3f7ca96 · outbound

This paper cites FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.573876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.573876Z digest=sha256:ff87ac4e6693378eb9087e6581a65262ddb687981423fc16443048228d629614

Observation 554c006a-5a8d-431e-90e0-d7d6e714a2b1 · outbound

This paper cites Large Language Models are Zero-Shot Reasoners.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Large Language Models are Zero-Shot Reasoners

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.579584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.579584Z digest=sha256:f3da1dbdbfc7136828f4b8a83acf9955f8fbcc8c47cdd75007cc9c9dd12573ae

Observation 8c110deb-f1f6-4d3d-b68f-216350ef4e65 · outbound

This paper cites an unresolved cited work.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:25:32.618511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:25:31.590606Z digest=sha256:be9fbf519038f97a96a14329d77b01523302697591b775d857bfad2d40b115cd

Observation e8113bb3-d5cd-4f4a-aaa8-2ac2ecb38777 · outbound

This paper cites Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Program Induction by Rationale Generation : Learning to Solve and Explain Algebraic Word Problems

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.596604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.596604Z digest=sha256:04fc060287fd048a3787e8d1a2d013671098372afa4ffbc99c724fa363bcf8e4

Observation 66f26733-5eae-44f1-b1d6-7627e89aeeab · outbound

This paper cites AlignBench: Benchmarking Chinese Alignment of Large Language Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.603431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.603431Z digest=sha256:d7d66bad39e589184fd3c74bc6ab49aa3bdde08fb7d2c2135e86badf0563e5ba

Observation d3cafca9-863a-4946-9151-f0ca0cfdaa72 · outbound

This paper cites The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios The Neuro-Symbolic Concept Learner: Interpreting Scenes, Words, and Sentences From Natural Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.609161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.609161Z digest=sha256:357d3019082a688e633690d9aa3b80fb8daead8c44414dcb73ba5ca762e27f60

Observation 8e8f1c96-c593-444a-ba49-79d753438df8 · outbound

This paper cites GPT-4 Technical Report.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios GPT-4 Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.620525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.620525Z digest=sha256:c6cef52c67fb202a620bd2473d5a3b75b1f7a0fb0c588c64ef426b1c92d722b3

Observation a7fe685c-4e24-41a8-9943-9522fe5cded6 · outbound

This paper cites an unresolved cited work.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:25:32.601585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:25:31.626065Z digest=sha256:bb00d40d3ad93565278e780423088e602ad092ef4fdaf289847ba5fdb3f5d710

Observation a2997379-5792-4bc2-ba04-86a4b9da5cdb · outbound

This paper cites Can LLMs Follow Simple Rules?.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Can LLMs Follow Simple Rules?

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.631000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.631000Z digest=sha256:df0cac8c39b08702ef2d7f05a8746947c1407bf6ea93f348fbaed61dcfe2f517

Observation d07f299e-107c-41cc-a9d8-1f3fb7fbe234 · outbound

This paper cites Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Logic-LM: Empowering Large Language Models with Symbolic Solvers for Faithful Logical Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.635627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.635627Z digest=sha256:7af41e387881de5cf047590885690df4c27f9c236d25fa442e052b7713fc069a

Observation 54d7ec07-72ff-4892-8fef-63be4a50f400 · outbound

This paper cites an unresolved cited work.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.640961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.640961Z digest=sha256:b861584cc24a91a5244bac6ca56b7f26282123d6655002526aecce6b8058a8f6

Observation 29a4a05b-dc45-474b-8d11-7feff824a42e · outbound

This paper cites Evaluating Large Language Models on Controlled Generation Tasks.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Evaluating Large Language Models on Controlled Generation Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.645790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.645790Z digest=sha256:6f3233c5b5c5267c622fce7590fa112bc8d1a4bb568753e399f657bb793120b9

Observation ff534545-f620-4513-9392-6df483194e48 · outbound

This paper cites Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Beyond Instruction Following: Evaluating Inferential Rule Following of Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.650877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.650877Z digest=sha256:f82658e6ae27f0077bfed832da469058c3fc71811c93767ebe057cbc0b4c7959

Observation 93c81ae0-192f-4a1e-a76b-4770122d047f · outbound

This paper cites ProofWriter: Generating Implications, Proofs, and Abductive Statements over Natural Language.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios ProofWriter: Generating Implications, Proofs, and Abductive Statements over Natural Language

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.657742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.657742Z digest=sha256:591d2621a755d67f8417ac4d160acae5378b33e2e3ce7bd136695f085bbad90e

Observation 1cfbde95-e41a-4638-a9fc-2b37c86c2a4e · outbound

This paper cites Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Struc-Bench: Are Large Language Models Really Good at Generating Complex Structured Data?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.663743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.663743Z digest=sha256:e4c8702ab87c38a48933ba2228eb847a22c70b9884abbf26d06abc1366c131b0

Observation ecbe2bb7-1901-4741-8d75-912bb2e577b5 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Gemini: A Family of Highly Capable Multimodal Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.669363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.669363Z digest=sha256:c924052f4b1215e40378ada0a338b2171dfb634190968e7f4fbaa91d4547151f

Observation e0baaebf-292d-4d06-a7f2-7c8db95538c2 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios LLaMA: Open and Efficient Foundation Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.674596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.674596Z digest=sha256:822b8c3307f0aa177744c1f76a3e8122f32c18c25732bbd5d9432e2bb5be14e9

Observation d1b92c73-8178-4984-90f9-730491d49f0c · outbound

This paper cites Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Interleaving Retrieval with Chain-of-Thought Reasoning for Knowledge-Intensive Multi-Step Questions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.679713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.679713Z digest=sha256:476964f6c5d7d8e1070013c657f90e2adc828d25527f2628c1a1404fc9974f5f

Observation c88a85ba-37a6-4566-a606-1ce458d5726c · outbound

This paper cites Symbolic Working Memory Enhances Language Models for Complex Rule Application.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Symbolic Working Memory Enhances Language Models for Complex Rule Application

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.685026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.685026Z digest=sha256:131eac0fb023b2bcb0971694a55e0a739c1c04f18f8c5b906d9cdf45cf7e6e3f

Observation c782645d-ffea-4ab0-b09b-eed7f272ade5 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.690218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.690218Z digest=sha256:398dba7b1d2bde49b849ca54f5cda0c54e9f356e0b04fc7a1555b7005807121d

Observation 1a85ad85-4774-4a3b-9fa8-dffe70884ab0 · outbound

This paper cites Larger language models do in-context learning differently.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Larger language models do in-context learning differently

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.695578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.695578Z digest=sha256:10c4c9e4d80c995efc970aab3d859fbfc6219905a6e36488100fd7693fa954c0

Observation 4cdb2072-3836-4bf3-93a8-6fbb1f18f0f0 · outbound

This paper cites Benchmarking Complex Instruction-Following with Multiple Constraints Composition.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Benchmarking Complex Instruction-Following with Multiple Constraints Composition

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.700799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.700799Z digest=sha256:2832589c1622fadc4c16b5d3d0bfeebc89891ac5d10bbe4ecb35bf4fe7dbd3c6

Observation 89f1b186-2986-43fc-b091-9e364538064d · outbound

This paper cites an unresolved cited work.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-11T17:25:32.573102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T17:25:31.706510Z digest=sha256:5977a6c0d5a8fa5082dd9e20d18899c24385d4dc565bc2d9937afa23addae931

Observation 5c57647e-2132-4c7d-9d04-3f09215c7e4b · outbound

This paper cites AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.711425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.711425Z digest=sha256:7e98a7f33f68fb07b94b6bc1947c3fd916574777851525bf05ecf298aa75f5e8

Observation 3ffddade-a12f-4d0c-822a-75a995be22b8 · outbound

This paper cites FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.716582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.716582Z digest=sha256:52370eca9fe9ca41bb2a6cb4f31abf66dc5922efebb4ea5b0c5fa79a95483c3d

Observation f9aa2e2f-1cc9-4ae8-9b06-958f5effb3e4 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.721684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.721684Z digest=sha256:8c63899a0cf5e96aa8f7857215d3e91d2ea177cae3b41efcc17263b8ce4ca766

Observation a8f55908-4f7a-4154-a59b-33998ccec510 · outbound

This paper cites Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.727116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.727116Z digest=sha256:31d590811c6fa5c143b7717e1e9ce9734ad2228730e2dddb15b32b63adb4ea87

Observation 844f10de-a47f-4635-bcba-4a3297884042 · outbound

This paper cites IDEAL: Influence-Driven Selective Annotations Empower In-Context Learners in Large Language Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios IDEAL: Influence-Driven Selective Annotations Empower In-Context Learners in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.732283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.732283Z digest=sha256:7fd47bb91f8e8e93fe2c6814ea4ee5c1b3f0f028304928502c2f26c5b767ad0c

Observation 384d583b-315a-4a8a-adf5-48678e0a6465 · outbound

This paper cites Supervised Chain of Thought.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Supervised Chain of Thought

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.737463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.737463Z digest=sha256:2066fe80eae518a14aa606b78882f0569ed2fdf50ad1c271186339f4adbdfd90

Observation 73562581-8cd3-4c46-b5f0-be198b988482 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.742515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.742515Z digest=sha256:becdf2088d437606e9ffcadfad5a9a2513233d72f390b159ebccd2c5e2ffca63

Observation 85d4fd0c-8c37-4866-b2a1-3e3105cc69c4 · outbound

This paper cites AR-LSAT: Investigating Analytical Reasoning of Text.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios AR-LSAT: Investigating Analytical Reasoning of Text

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.747587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.747587Z digest=sha256:9381b6fcc8b4ab363d9e303d882255e760718a992b9c44f8bb661cb9848ed2bc

Observation 74a6d8bb-5218-4579-9c27-2f087c4131ca · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios Instruction-Following Evaluation for Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.752696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.752696Z digest=sha256:fe165655b8e029cecfab730b54d0e0c73d4af2b43c2301f91fa9a1b722615f74

Observation ab0a5149-1d39-483d-bf07-d5ecd39f7b4f · outbound

This paper cites TRAD: Enhancing LLM Agents with Step-Wise Thought Retrieval and Aligned Decision.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios TRAD: Enhancing LLM Agents with Step-Wise Thought Retrieval and Aligned Decision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.757823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.757823Z digest=sha256:227c10948847bfe9791b919a55ab3a95fd78478ac1e0f4b5627a6859d68fdb13

Observation 7b10c5b1-eb13-437b-a2d3-e8c4a2892122 · outbound

This paper cites DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios DyVal: Dynamic Evaluation of Large Language Models for Reasoning Tasks

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.762978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.762978Z digest=sha256:0d8a147fb7e83ae8cce3802a323e1bf3890436de7ed5d7ea2f7b436763828ac3

Pith citing papers

Observation dd2a529d-24b1-45f7-8885-b74ad43c7ffc · inbound

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge cites this paper.

AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World Knowledge RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:58:02.549983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T12:58:02.335060Z digest=sha256:8169d772eec43373a569f8d4c1003984beb81430c62eba08e02e572bc0966ed2