Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Reasoning Robustness in Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2503.04550.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.04550 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:52:41.500608Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T16:53:40.544660Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ece21f9f-633a-43c6-9300-958f80c86e96 · inbound

Wireless Agentic AI with Retrieval-Augmented Multimodal Semantic Perception cites this paper.

Wireless Agentic AI with Retrieval-Augmented Multimodal Semantic Perception Benchmarking Reasoning Robustness in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:41.500608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:41.500608Z digest=sha256:3d2362e8a7ed1a8654e7b429928cdaf65084e161926925d41c0d1076fc11e4f6

Observation addce766-70cb-442f-9db5-af65c27dccac · inbound

Bayesian Social Deduction with Graph-Informed Language Models cites this paper.

Bayesian Social Deduction with Graph-Informed Language Models Benchmarking Reasoning Robustness in Large Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:32:09.373274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T07:27:20.027179Z digest=sha256:b1f3515540ddec99430423ef0267fc51af4d4510751b88729c9c43ce13a65c39

Observation a99acdcf-a909-44f6-afdf-ec0a1265bd8d · inbound

A Large Language Model-Empowered Agent for Reliable and Robust Structural Analysis cites this paper.

A Large Language Model-Empowered Agent for Reliable and Robust Structural Analysis Benchmarking Reasoning Robustness in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.584043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:21:11.584043Z digest=sha256:107d4997912ec7c55a86c1c4c0dd061a367db3a130c7575ec820cd044667b42e

Observation 3b6d29b8-3d4f-40aa-a26c-e2931ad03d7c · inbound

Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies cites this paper.

Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies Benchmarking Reasoning Robustness in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:37.741861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:58:37.741861Z digest=sha256:dc8bcedc0c53a23eacedba3d36e40d913d3152767feb33a1bb5b8be50eae474d

Observation 343f7b8a-0f47-48cc-a7ea-b84176a34a89 · inbound

Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need cites this paper.

Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need Benchmarking Reasoning Robustness in Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:41.606874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:41.606874Z digest=sha256:089223eb4561159c21b14627485666729f20ce6076a8160ec524d47ba4a037a0

Observation ab2b6813-c8b3-4b9b-b62b-dc614589eade · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models Benchmarking Reasoning Robustness in Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:10.164969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:10.164969Z digest=sha256:61314afa611110ec779b8ed4feb4c9827479b687478d74951072d3813c9e8c11

Observation 49485f00-064e-4416-8d89-69c43dac6161 · inbound

Throttling Web Agents Using Reasoning Gates cites this paper.

Throttling Web Agents Using Reasoning Gates Benchmarking Reasoning Robustness in Large Language Models

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:05.016955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:05.016955Z digest=sha256:c623cef776432491e622ef5786aea80f537f70a12460212bb27269d8cb17feaa

Observation e460ba93-291d-45bf-97f4-e0a883657636 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Benchmarking Reasoning Robustness in Large Language Models

Reference 189

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:07.143764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:07.143764Z digest=sha256:abd4fafb18d4926b367bc62ae5333e84ac011570ac5e437020212aa0a676557c

Observation 21109c5e-c722-4689-b483-84b58368f1a7 · inbound

Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation cites this paper.

Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation Benchmarking Reasoning Robustness in Large Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:20:06.816253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T12:18:18.407788Z digest=sha256:c6cfd5eec8e6ac1fa23e65becbd2c533717db38abb2413181b4c269252b89dcf

Observation 985e2736-6776-4983-a8be-cc8b35157f77 · inbound

Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning cites this paper.

Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning Benchmarking Reasoning Robustness in Large Language Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:22:01.794846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:49.761472Z digest=sha256:1e1be96d92971e4b33fbf59af6c70dfae65a8a062d194444a5e32e0d0d7b766e

Observation 6c9f0343-f6fb-446a-a119-2f5a828c2255 · inbound

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments cites this paper.

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments Benchmarking Reasoning Robustness in Large Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:53:40.546212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T16:51:36.524194Z digest=sha256:83dac10470730fc623da978cb809d081c7a31990f228e187f480202b3744ec59

Observation 827760a3-26cf-493a-a2ca-eebe1f86bf14 · inbound

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models cites this paper.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Benchmarking Reasoning Robustness in Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:55.417781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:55.417781Z digest=sha256:bdc3419f312fb57873e3af262f3ef3cda39d51b9877a12c55caa41051eafaac7

Observation 07a6f5f0-324f-4787-9690-08dd164475b6 · inbound

Implicit Reasoning Steering via Concept Chaining cites this paper.

Implicit Reasoning Steering via Concept Chaining Benchmarking Reasoning Robustness in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:22.985400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:44:22.985400Z digest=sha256:0b2f3d2daf3b8cdb765c12017a09558dab6841426ae5bf7d8b71c2794ed23350

Observation 81fb91db-cea3-4b1d-af33-11f627ce29fd · inbound

Implicit Reasoning Steering via Concept Chaining cites this paper.

Implicit Reasoning Steering via Concept Chaining Benchmarking Reasoning Robustness in Large Language Models

Reference 186

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:37.046778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:44:37.046778Z digest=sha256:e5519061295ab90a0f0183d5dd762d03fe523bea10629f0b2a7b523aa484ffa3