Pith. sign in

Paper Citation Record · LEDGER

A Primer in Post-Training Reasoning Data: What We Know About How It Works

As of 5 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2606.02113.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.02113 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T14:40:21.583101Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T09:56:36.860369Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch15

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 4ae6c1e7-fad5-4de1-90c8-13d89c3c5446 · outbound

This paper cites Miranda, Hanyang Zhao, Mohammad Rifqi Farhansyah, Garry Kuwanto, Derry Wijaya, and Genta Indra Winata.

A Primer in Post-Training Reasoning Data: What We Know About How It Works Miranda, Hanyang Zhao, Mohammad Rifqi Farhansyah, Garry Kuwanto, Derry Wijaya, and Genta Indra Winata

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:20.551043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:02f4a5bd1ae6d0e5cbed24b839081239827466885534126dbb620c51c5b1d39c

Observation 808a10a8-30b8-44f4-979b-5d8fad005d58 · outbound

This paper cites Distillation Scaling Laws.

A Primer in Post-Training Reasoning Data: What We Know About How It Works Distillation Scaling Laws

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:20.492970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:1d46e62644d0fae76c4c688cbbdaa5ded5a8072d730e0821c78d663d721022af

Observation 9f9e72e2-5555-44cf-8ce9-38574fba3712 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

A Primer in Post-Training Reasoning Data: What We Know About How It Works Evaluating Large Language Models Trained on Code

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:20.532810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:aedcdd4f4a6ff05d19a71c164d3b9f612047eae8c19869405060aa5a9c8e0685

Observation dd66c078-fe3a-462f-8955-428831c3e560 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

A Primer in Post-Training Reasoning Data: What We Know About How It Works Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:20.537818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:e33f6dc13d40babba05ab55d56fe7d5222d63ea6454341379802f2421ec3df73

Observation 3f39b1b4-ded4-4070-8a1b-a2fb37d013ba · outbound

This paper cites Mind2Web: Towards a Generalist Agent for the Web.

A Primer in Post-Training Reasoning Data: What We Know About How It Works Mind2Web: Towards a Generalist Agent for the Web

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:20.541844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:de799af8b00e999aea4b1f0dde1de74630863a552a0b3904e6960a0698bcbcf9

Observation 6b6c7bd1-eb96-4875-9c45-fd7ca5f451fc · outbound

This paper cites WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?.

A Primer in Post-Training Reasoning Data: What We Know About How It Works WorkArena: How Capable Are Web Agents at Solving Common Knowledge Work Tasks?

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:20.567771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:9b7d1af0fab5fd592e89d29ee17eacd3b10a7a3b652651ceac92eb4069c50f38

Observation 82041640-7447-4d81-9384-28a5337c6b9f · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

A Primer in Post-Training Reasoning Data: What We Know About How It Works Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:20.512563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:16f8d8fa4a59bc2528154c22a7b16b89d4d8b0700cf86c8c20f8e55ac3cade49

Observation 56d37dab-4fe5-41a2-8e94-82c19423c538 · outbound

This paper cites Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis.

A Primer in Post-Training Reasoning Data: What We Know About How It Works Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:18:41.255312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:ecbb221b0a6e356487d760c9d1f85314bf893a46102ceabded123890eea01bdf

Observation 67e9b3c8-9058-4b56-9138-3a6d0a996b31 · outbound

This paper cites Let's Verify Step by Step.

A Primer in Post-Training Reasoning Data: What We Know About How It Works Let's Verify Step by Step

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:20.563872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:d221d41376d7e3ce23300323cae06a8a5ae1d170609aa6cbd1dd445f676196dc

Observation 7d2a08dc-3f14-428a-b890-da61d4791cf8 · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

A Primer in Post-Training Reasoning Data: What We Know About How It Works GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:20.502805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:64456c1fdc456bcdb08275d3d95cba244b1a5926e64014234126b509d45a62cb

Observation 1b858ce7-4e44-4d77-a08b-8ca5bbb14aa8 · outbound

This paper cites Magistral.

A Primer in Post-Training Reasoning Data: What We Know About How It Works Magistral

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:20.508069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:b9d245e41f6c3768105ecb5089b218f30dcbcb8ce5b2e65536cb46cbf405a361

Observation 3ef11069-d661-469e-9df3-f117a04d2957 · outbound

This paper cites AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents.

A Primer in Post-Training Reasoning Data: What We Know About How It Works AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:20.546412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:f27f89f300490fff4213298084925ad91e9f82c363039abb0dfa65c0e7125fa1

Observation fa177d23-21bd-4573-af6e-34494a0d7843 · outbound

This paper cites AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents.

A Primer in Post-Training Reasoning Data: What We Know About How It Works AppWorld: A Controllable World of Apps and People for Benchmarking Interactive Coding Agents

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:20.559817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:6e731a1df703f791a4b6a13ff912def48201d2eac16b3c605d67a0a17aaf56d9

Observation b4e51fb7-0546-453b-bd12-eb062a79c0af · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

A Primer in Post-Training Reasoning Data: What We Know About How It Works Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T23:06:20.523341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:bd29f2014eb5e900ba137b228820a1bcc8190e6249e2c242bb53f73845e68f4f

Observation 182ff33b-f889-44ee-819c-54eae366d611 · outbound

This paper cites Naturalreasoning: Reasoning in the wild with 2.8 m challenging questions.

A Primer in Post-Training Reasoning Data: What We Know About How It Works Naturalreasoning: Reasoning in the wild with 2.8 m challenging questions

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:20.498100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:ab8b5d02866db7b500de6891e9c52f15ae9dcc4be55c466cf371de47088e29b6

Observation 59bda563-1572-4571-b99d-1c27eae965db · outbound

This paper cites InThe Tenth In- ternational Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022.

A Primer in Post-Training Reasoning Data: What We Know About How It Works InThe Tenth In- ternational Conference on Learning Representations, ICLR 2022, Virtual Event, April 25-29, 2022

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T14:40:21.583101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:69e1ee6c4f68ca22747caea9cd81625a243f7a0819fa530828f57d85a8944968

Observation 117f0cc5-3636-4ff7-aef3-fa8454790d5f · outbound

This paper cites TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance.

A Primer in Post-Training Reasoning Data: What We Know About How It Works TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T23:06:20.518308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:1228f72680aa03d12fa6cd42dce3ae1c93bfaef0126ecda9137b1a89446902ff

Observation 335b65e9-9f07-4abe-9224-25d17d13e7e9 · outbound

This paper cites JudgeLM: Fine-tuned Large Language Models are Scalable Judges.

A Primer in Post-Training Reasoning Data: What We Know About How It Works JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:06:20.555946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T14:40:21.583101Z digest=sha256:af80358020c3e3c4fcc1314d5267b3d18a5da9e7e8b65cd9da881ac5c8516759

Pith citing papers

Observation 7c4c5b44-b8a2-4d3f-9901-babc99db77f9 · inbound

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions cites this paper.

RealClawBench: Live OpenClaw Benchmarks from Real Developer-Agent Sessions A Primer in Post-Training Reasoning Data: What We Know About How It Works

Reference 35

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T10:01:52.590703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-28T09:56:36.860369Z digest=sha256:1975d5e8b1962808c96435366cb5f9207877a4238dec70b328ecc384550d2459