Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2404.01869.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:26.602021Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T23:27:26.723998Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation d1afb69a-6236-47d2-8128-61707e539603 · inbound
TS-Reasoner: Domain-Oriented Time Series Inference Agents for Reasoning and Automated Analysis Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c436901c-6679-44ae-b83f-896fa2f89fb7 · inbound
Large Language Model assisted Hybrid Fuzzing Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 18926fbe-dfa6-4174-bbb5-84dc97c31bec · inbound
PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5eae57af-4b47-48db-ac84-34593590a22b · inbound
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee5f1057-1385-4871-b8bd-00f9d536a1e2 · inbound
Propositional Logic for Probing Generalization in Neural Networks Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3608d722-55aa-45b5-8579-e4e6c896ec9e · inbound
The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0665d821-a98b-4193-949d-93d704d56573 · inbound
Unveiling Causal Reasoning in Large Language Models: Reality or Mirage? Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c30afdeb-77fb-4f63-aa54-b3bce31efbbd · inbound
Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0201270e-fd6e-4c99-a112-338a45d35d92 · inbound
Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d2d8d73-ee8e-4180-82a9-32fcbcc1dac6 · inbound
SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9029285-c2f9-4d4c-82fa-3c3c26941291 · inbound
HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f16f27d-be57-49ad-8229-e64dba374242 · inbound
LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18f575ae-5a69-4698-a01e-48370b28e9fa · inbound
ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d019ff42-c5a4-4de4-a718-a161f8264ee1 · inbound
Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e57b035f-3313-4c7f-9912-7c9a63b2e998 · inbound
Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5e95860f-12fd-4d31-8e4b-8cc8a0e2b51c · inbound
Measuring AI Reasoning: A Guide for Researchers Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 138
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5477db6f-4f37-47e1-9e7d-fbf2e56e0112 · inbound
Reasoning emerges from constrained inference manifolds in large language models Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2700da80-4028-4f46-96d9-30ebecdca0d8 · inbound
Cross-Lingual Consensus: Aligning Multilingual Cultural Knowledge via Multilingual Self-Consistency Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2af6ec53-3384-4746-b4df-6e62da89311e · inbound
Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a18e703d-5e5b-4251-957c-40f7849ac7ab · inbound
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 480fcf42-782a-47c0-903a-d240184b144c · inbound
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 157347d4-b946-48ce-b33b-6149dfa131a4 · inbound
Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.