Pith. sign in

Paper Citation Record · LEDGER

ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2304.10703.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2304.10703 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:33:47.710576Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:37:36.781188Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 818a946b-3f6c-4f31-a07f-6200ea3d67e2 · inbound

Retrieval Augmented Decision-Making: A Requirements-Driven, Multi-Criteria Framework for Structured Decision Support cites this paper.

Retrieval Augmented Decision-Making: A Requirements-Driven, Multi-Criteria Framework for Structured Decision Support ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:47.710576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:47.710576Z digest=sha256:9b8e6024eec4860c1a8e228f7330e5fbff530e0f1590f0265770c4d290ead5d8

Observation 9e83b7d7-971b-4220-8f1f-5ce5d603d714 · inbound

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM cites this paper.

MIRAGE: Assessing Hallucination in Multimodal Reasoning Chains of MLLM ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:33:10.405095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:33:10.405095Z digest=sha256:87fc01374974e5291beec7310d56f783d24c18fab4bc1e6f7b8eb154d9b53903

Observation 4d15d7d0-eef1-42aa-a693-9fbaf3263aff · inbound

Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains cites this paper.

Knowledge or Reasoning? A Close Look at How LLMs Think Across Domains ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:32.114341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:32.114341Z digest=sha256:ab7807b3d5d1ca60c64f28c2cde3f7c23e06e6bce8795707e49f481dab8d701b

Observation 7d6bf28a-8846-4d2c-a4c9-afa79093e908 · inbound

Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation cites this paper.

Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:51:08.254881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:51:08.254881Z digest=sha256:2e72d1af2b706ae6857acd9f21d496d9b04ac478035f495bf5d7b440425c9609

Observation b8e998bf-7eae-4b59-bb1d-43c97d681ccf · inbound

Rethinking Human Preference Evaluation of LLM Rationales cites this paper.

Rethinking Human Preference Evaluation of LLM Rationales ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T17:15:27.887155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:15:27.887155Z digest=sha256:e6b3f647338a1bd02804a3d1cd4a4b79eb5ac6a209a9fbca34c419321da47263

Observation 9fea0482-9d0f-4543-996a-c99c71027027 · inbound

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data cites this paper.

A Survey on Evaluating Quality and Trustworthiness in LLM-Generated Data ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness

Reference 173

Resolution
unresolved
no resolver link, observed 2026-08-03T08:15:28.060444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T08:15:28.060444Z digest=sha256:61b8c26385b01235e10a840dd81da35807eca369b4562a7d440c82facf82fdd0

Observation c123eb9d-781b-49eb-9ee1-5674ca531a17 · inbound

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations cites this paper.

Fragile Thoughts: How Large Language Models Handle Chain-of-Thought Perturbations ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:47:15.332782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T03:43:18.987241Z digest=sha256:5eb9fff25375c85bb6c82e81ec3d134c55d79d49be154b17dca6a54554a8ac7d

Observation 94b106a3-1bc7-4b3d-8944-042e81ff2d1f · inbound

Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks cites this paper.

Beyond Output Correctness: Benchmarking and Evaluating Large Language Model Reasoning in Coding Tasks ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:36:02.566440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:56:03.196091Z digest=sha256:6e470a51bc030586b38f5aaeee315a594b5b14fe8ac272248de00f4c953669b7

Observation cce65705-223c-4626-b753-e9dab7a3b6e1 · inbound

Supervised Fine-tuning with Synthetic Rationale Data Hurts Real-World Disease Prediction cites this paper.

Supervised Fine-tuning with Synthetic Rationale Data Hurts Real-World Disease Prediction ReCEval: Evaluating Reasoning Chains via Correctness and Informativeness

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:37:36.786700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T13:48:39.832057Z digest=sha256:a2564d2a4b9aaba53ac8ca90e4cab7bd1a2925dffee62ec22c8b57faf25e53f6