Pith. sign in

Paper Citation Record · LEDGER

MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2312.17080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.17080 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:22:56.090625Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4502f655-9c6c-40f9-939f-c6f49a5381a6 · inbound

LIMO: Less is More for Reasoning cites this paper.

LIMO: Less is More for Reasoning MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:11:37.140177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T02:11:36.932541Z digest=sha256:3718ba6f9c8261d78b298cec3cfb24d5429af106fd35e31733eabdce1e856754

Observation 3352a403-d713-4bc4-8714-334ec1debcbf · inbound

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos cites this paper.

VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T04:22:56.090625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:22:56.090625Z digest=sha256:5ef4acee3b08ba9119c4eb86a019c66cd4f80dd94a35708329cf919fc5424d56

Observation 6f78baf5-d5ec-4367-843b-8004c469db1f · inbound

Can You Trick the Grader? Adversarial Persuasion of LLM Judges cites this paper.

Can You Trick the Grader? Adversarial Persuasion of LLM Judges MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T21:56:01.377003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:56:01.377003Z digest=sha256:2603b1a1a3dfdfde9cc44b509d92fe122a657415b6a346641bfbf22e9dd1cced

Observation a493ec8a-69c0-4af4-8a3d-1a0ba82680a9 · inbound

LLMs cannot spot math errors, even when allowed to peek into the solution cites this paper.

LLMs cannot spot math errors, even when allowed to peek into the solution MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T12:40:45.229549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:40:45.229549Z digest=sha256:449855c49e01be3243e208746ce397def6d8d4244b1f7c6c8b70b9442064a98c

Observation b9f1217e-f53c-4cb7-8020-a8fb7bf01c78 · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:45.076669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:45.076669Z digest=sha256:63fad78101971f550a368a2aee47bdc65b848cea2032e2cab0173749af4793c4

Observation 676ed4a1-d5b1-4aa6-abde-c2d56aed9a2a · inbound

MedPRMBench: A Fine-grained Benchmark for Process Reward Models in Medical Reasoning cites this paper.

MedPRMBench: A Fine-grained Benchmark for Process Reward Models in Medical Reasoning MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:01:13.462381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:58:03.842336Z digest=sha256:a2d598f544a6debec3c3bd62c844428f7bee57aa39aa8bbaab1d22d1d2a23356

Observation 9210e3f8-2089-426b-9df2-480b33761759 · inbound

DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models cites this paper.

DRIFTLENS: Measuring Memory-Induced Reasoning Drift in Personalized Language Models MR-GSM8K: A Meta-Reasoning Benchmark for Large Language Model Evaluation

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:38:18.433521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-03T13:37:39.545003Z digest=sha256:5d036c657bd33f92ff8be53431f8e3713600e016ad04fa7fcdf78b971b2392e4