Pith. sign in

Paper Citation Record · LEDGER

Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 22 inbound Pith citation observations for arXiv:2404.01869.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.01869 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 22 of 22 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:17:26.602021Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:27:26.723998Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d1afb69a-6236-47d2-8128-61707e539603 · inbound

TS-Reasoner: Domain-Oriented Time Series Inference Agents for Reasoning and Automated Analysis cites this paper.

TS-Reasoner: Domain-Oriented Time Series Inference Agents for Reasoning and Automated Analysis Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:45:47.180451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T19:45:39.130509Z digest=sha256:b8674b101735816605c4c6708c699d90ce981baa9726b5978f61f237607280d7

Observation c436901c-6679-44ae-b83f-896fa2f89fb7 · inbound

Large Language Model assisted Hybrid Fuzzing cites this paper.

Large Language Model assisted Hybrid Fuzzing Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:15:28.694964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T07:13:13.613187Z digest=sha256:268af5a4259487d08bbecda1482cd89af33d24318afebee9c2f6f70b5abe0ec8

Observation 18926fbe-dfa6-4174-bbb5-84dc97c31bec · inbound

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts cites this paper.

PuzzleWorld: A Benchmark for Multimodal, Open-Ended Reasoning in Puzzlehunts Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-19T10:47:15.149914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T10:43:02.601014Z digest=sha256:c7d2bbab578690ab304fb316a6d6b03a6e966ca40c202d9fc2e0ab3b2b919a5d

Observation 5eae57af-4b47-48db-ac84-34593590a22b · inbound

Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation cites this paper.

Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 21

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:51:08.241334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:51:08.241334Z digest=sha256:b653ff46efe3f5b52026e3dac94968eee19ceb7a11f6305dc6f24e0dde016ad2

Observation ee5f1057-1385-4871-b8bd-00f9d536a1e2 · inbound

Propositional Logic for Probing Generalization in Neural Networks cites this paper.

Propositional Logic for Probing Generalization in Neural Networks Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:08:57.700850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:08:57.700850Z digest=sha256:b2c3a027376c8e8f88bcefcf66210e9fe96dba7adb71494c7fc5d6a33bd0cc27

Observation 3608d722-55aa-45b5-8579-e4e6c896ec9e · inbound

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs cites this paper.

The Scales of Justitia: A Comprehensive Survey on Safety Evaluation of LLMs Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:26.602021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:26.602021Z digest=sha256:8ee9dc187eb548a1b75e06e2ae0309629a642123e817da1b46b920cf07e12f4d

Observation 0665d821-a98b-4193-949d-93d704d56573 · inbound

Unveiling Causal Reasoning in Large Language Models: Reality or Mirage? cites this paper.

Unveiling Causal Reasoning in Large Language Models: Reality or Mirage? Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:23.526141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:23.526141Z digest=sha256:670f8a1c9460404a0b4abbc8b8c2aa02280ec540cf7a3727a45742ddd2abf31a

Observation c30afdeb-77fb-4f63-aa54-b3bce31efbbd · inbound

Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models cites this paper.

Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:22.769487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:58:22.769487Z digest=sha256:a082e34f9884c465180db5cc8dd7e85a4d91a2a850197d080c55709f2ce72411

Observation 0201270e-fd6e-4c99-a112-338a45d35d92 · inbound

Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies cites this paper.

Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:36.232363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:58:36.232363Z digest=sha256:c0f07858d0cd587dc9f80aa2d2ba9db7326d1c8e244338ff1c5c06d79ce0fd2e

Observation 8d2d8d73-ee8e-4180-82a9-32fcbcc1dac6 · inbound

SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models cites this paper.

SCOPE: Stochastic and Counterbiased Option Placement for Evaluating Large Language Models Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:45:42.582894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:45:42.582894Z digest=sha256:3b291cabac4cd9211502d61a3d6933f33fc209f517297430121b1975ada93d02

Observation a9029285-c2f9-4d4c-82fa-3c3c26941291 · inbound

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation cites this paper.

HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T05:41:42.094226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:41:42.094226Z digest=sha256:a5a56c1ffaa445b00153d8fcab7b99da571bb4e9bf23510ea26285e4eb395ac9

Observation 3f16f27d-be57-49ad-8229-e64dba374242 · inbound

LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue cites this paper.

LLMs and their Limited Theory of Mind: Evaluating Mental State Annotations in Situated Dialogue Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T11:46:53.218414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:46:53.218414Z digest=sha256:f86462400a7e98dff045cd477c3eb5d3637f31244796ce8e37b5a9337974275b

Observation 18f575ae-5a69-4698-a01e-48370b28e9fa · inbound

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning cites this paper.

ReasonBENCH: Benchmarking the (In)Stability of LLM Reasoning Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T17:53:55.547037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:53:55.547037Z digest=sha256:e4c058479b58e05e5c4d7082a3e0a2b7045be3cf47c1d18505173f5beddf5f41

Observation d019ff42-c5a4-4de4-a718-a161f8264ee1 · inbound

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations cites this paper.

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:47:15.358111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T03:43:18.987241Z digest=sha256:f147eb09ccb15f3b088bb5057604b6d3114e8425efab15d2439c8fd31406077b

Observation e57b035f-3313-4c7f-9912-7c9a63b2e998 · inbound

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints cites this paper.

Empirical Evidence of Complexity-Induced Limits in Large Language Models on Finite Discrete State-Space Problems with Explicit Validity Constraints Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:15:29.419676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T14:12:45.438246Z digest=sha256:7bc4b088046cf33e1524b51e59666a5b7e57d915c3f8a710a6e6d45818b00522

Observation 5e95860f-12fd-4d31-8e4b-8cc8a0e2b51c · inbound

Measuring AI Reasoning: A Guide for Researchers cites this paper.

Measuring AI Reasoning: A Guide for Researchers Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 138

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:05:36.704850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T18:53:18.586923Z digest=sha256:6af85728080cd29bc9975e235174eec193c1fdb6be830d308a73bf8a4139fb9c

Observation 5477db6f-4f37-47e1-9e7d-fbf2e56e0112 · inbound

Reasoning emerges from constrained inference manifolds in large language models cites this paper.

Reasoning emerges from constrained inference manifolds in large language models Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:01:15.379951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T01:58:30.713839Z digest=sha256:2abe85e1fb1f8e51ac73cad7ffd03d2c5f87809fcc87fe472115f418de12c0b7

Observation 2700da80-4028-4f46-96d9-30ebecdca0d8 · inbound

Cross-Lingual Consensus: Aligning Multilingual Cultural Knowledge via Multilingual Self-Consistency cites this paper.

Cross-Lingual Consensus: Aligning Multilingual Cultural Knowledge via Multilingual Self-Consistency Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.575582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T17:31:39.958584Z digest=sha256:ceaba87d0f84ba9a5e3ca6aea6ae2d467e95b5d2abb26562ff93c7b579f645d3

Observation 2af6ec53-3384-4746-b4df-6e62da89311e · inbound

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild cites this paper.

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.106736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T14:41:07.354007Z digest=sha256:e76b1cc5f01f6074b73bf1dd55e6c9df35364bab487a2c82778facc2750ab18b

Observation a18e703d-5e5b-4251-957c-40f7849ac7ab · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:34:40.593743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T13:27:50.367497Z digest=sha256:898bbc68c4aa435e120cd43f4aeaa57abd423c32d5384d08d4de955cd93db8ef

Observation 480fcf42-782a-47c0-903a-d240184b144c · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:05:31.924167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T07:35:51.017797Z digest=sha256:122fdb7a0b1d1872621ee3a7465eca7aa4bcc8c29b0e2ccb22a305c5ee66c915

Observation 157347d4-b946-48ce-b33b-6149dfa131a4 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:27:26.725939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T23:22:08.567841Z digest=sha256:f401e51992d0962f14865e12391920dbe5a53c835d82d00c4a3afa55784450af