Pith. sign in

Paper Citation Record · LEDGER

Benchmarking Reasoning Robustness in Large Language Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2503.04550.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.04550 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:52:41.500608Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T16:53:40.544660Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ece21f9f-633a-43c6-9300-958f80c86e96 · inbound

Wireless Agentic AI with Retrieval-Augmented Multimodal Semantic Perception cites this paper.

Wireless Agentic AI with Retrieval-Augmented Multimodal Semantic Perception Benchmarking Reasoning Robustness in Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:41.500608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:52:41.500608Z digest=sha256:e86248b8948894645aa42fb88e98bdf17c51b226b0331d4cd3868c6fbca4ea3b

Observation addce766-70cb-442f-9db5-af65c27dccac · inbound

Bayesian Social Deduction with Graph-Informed Language Models cites this paper.

Bayesian Social Deduction with Graph-Informed Language Models Benchmarking Reasoning Robustness in Large Language Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:32:09.373274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T07:27:20.027179Z digest=sha256:d28cb7b7042efb10cafe3b18c78a0123a8cc5f112fe9786f21e67e3daef92c2e

Observation a99acdcf-a909-44f6-afdf-ec0a1265bd8d · inbound

A Large Language Model-Empowered Agent for Reliable and Robust Structural Analysis cites this paper.

A Large Language Model-Empowered Agent for Reliable and Robust Structural Analysis Benchmarking Reasoning Robustness in Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:21:11.584043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:21:11.584043Z digest=sha256:7bffe8bb7ea7d11da44948bf127d950eee8090abf0c8f92bad3277e5083d82dc

Observation 3b6d29b8-3d4f-40aa-a26c-e2931ad03d7c · inbound

Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies cites this paper.

Dissecting Clinical Reasoning in Language Models: A Comparative Study of Prompts and Model Adaptation Strategies Benchmarking Reasoning Robustness in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:58:37.741861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:58:37.741861Z digest=sha256:a3eff66b35640c467966d68c5081def42d3885f48c9260a7401dd3f1822fb28d

Observation 343f7b8a-0f47-48cc-a7ea-b84176a34a89 · inbound

Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need cites this paper.

Bottom-up Domain-specific Superintelligence: A Reliable Knowledge Graph is What We Need Benchmarking Reasoning Robustness in Large Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T16:17:41.606874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:17:41.606874Z digest=sha256:05c168511f8a1f04ecf8040b4e76134e08fec5cb4434269586cc10bd6f441935

Observation ab2b6813-c8b3-4b9b-b62b-dc614589eade · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models Benchmarking Reasoning Robustness in Large Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:10.164969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:10.164969Z digest=sha256:2916d26269d6d577674b87855f45df360b920a4ee89743bbfe41e9b5f9d43b6d

Observation 49485f00-064e-4416-8d89-69c43dac6161 · inbound

Throttling Web Agents Using Reasoning Gates cites this paper.

Throttling Web Agents Using Reasoning Gates Benchmarking Reasoning Robustness in Large Language Models

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-05T12:28:05.016955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:28:05.016955Z digest=sha256:66cec3e46350be24586ee04b560622faac8fddd4888c5807d32cc28e95e27643

Observation e460ba93-291d-45bf-97f4-e0a883657636 · inbound

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models cites this paper.

A Comprehensive Survey on Trustworthiness in Reasoning with Large Language Models Benchmarking Reasoning Robustness in Large Language Models

Reference 189

Resolution
unresolved
no resolver link, observed 2026-08-05T10:39:07.143764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:39:07.143764Z digest=sha256:72a49b4e615b896044d42e1bb6cb938ff8f10d1428dfdbbb751addfd2ab9d396

Observation 21109c5e-c722-4689-b483-84b58368f1a7 · inbound

Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation cites this paper.

Conv-FinRe: A Conversational and Longitudinal Benchmark for Utility-Grounded Financial Recommendation Benchmarking Reasoning Robustness in Large Language Models

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:20:06.816253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T12:18:18.407788Z digest=sha256:c1dc6cf2ec57909bfe6bdfa5a036e956cbf45229695218afa10621b10aac924f

Observation 985e2736-6776-4983-a8be-cc8b35157f77 · inbound

Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning cites this paper.

Seir\^enes: Adversarial Self-Play with Evolving Distractions for LLM Reasoning Benchmarking Reasoning Robustness in Large Language Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:22:01.794846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:19:49.761472Z digest=sha256:98263f5b02a51dfe0c5918033648819a9bdfca78b8d418549834515c3dc3b50f

Observation 6c9f0343-f6fb-446a-a119-2f5a828c2255 · inbound

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments cites this paper.

Learning to Act under Noise: Enhancing Agent Robustness via Noisy Environments Benchmarking Reasoning Robustness in Large Language Models

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:53:40.546212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T16:51:36.524194Z digest=sha256:c40a374cabcfd183c6248266afb6b33190699b2b4b7887fb0d6a2c2462a7ac35

Observation 827760a3-26cf-493a-a2ca-eebe1f86bf14 · inbound

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models cites this paper.

GSM-Plus-BN: A Perturbation-Based Benchmark for Bangla Mathematical Reasoning in Large Language Models Benchmarking Reasoning Robustness in Large Language Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T05:48:55.417781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:48:55.417781Z digest=sha256:a8bc8aee72ad0bb296b1a52fdba62bffc87fa7140e778cb29b3804340a9a2175

Observation 07a6f5f0-324f-4787-9690-08dd164475b6 · inbound

Implicit Reasoning Steering via Concept Chaining cites this paper.

Implicit Reasoning Steering via Concept Chaining Benchmarking Reasoning Robustness in Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:22.985400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:44:22.985400Z digest=sha256:5fc62a4b679ba558758e1de98988a1ba6fe2ab85cc918e59c77f510795364345

Observation 81fb91db-cea3-4b1d-af33-11f627ce29fd · inbound

Implicit Reasoning Steering via Concept Chaining cites this paper.

Implicit Reasoning Steering via Concept Chaining Benchmarking Reasoning Robustness in Large Language Models

Reference 186

Resolution
unresolved
no resolver link, observed 2026-08-02T02:44:37.046778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T02:44:37.046778Z digest=sha256:ecaf075ef11b000d3af5bb61ff1d2dc4ccd95a22056b724dea40d7e8d14e5dc4