Pith. sign in

Paper Citation Record · LEDGER

Self-Refine Instruction-Tuning for Aligning Reasoning in Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2405.00402.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.00402 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:26.110877Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T14:13:08.051352Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 08e12084-5f36-479a-b28e-38976a4a530b · inbound

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark cites this paper.

ARB: A Comprehensive Arabic Multimodal Reasoning Benchmark Self-Refine Instruction-Tuning for Aligning Reasoning in Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:26.110877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:26.110877Z digest=sha256:a1532cf2e4d3ef02e19bcf309621ee76cdadb50e82d133845f173105e9682b4f

Observation 8ce904d1-c95e-4ebd-aec5-6cf269284008 · inbound

Self-Training Large Language Models with Confident Reasoning cites this paper.

Self-Training Large Language Models with Confident Reasoning Self-Refine Instruction-Tuning for Aligning Reasoning in Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:15.279244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:51:15.279244Z digest=sha256:155250e861cbad0fcac080702af20f0799707674871b76e91e680b83c3679930

Observation 12b19ab3-0ece-4e12-a6df-12fa7d1a6704 · inbound

Aligning VLM Assistants with Personalized Situated Cognition cites this paper.

Aligning VLM Assistants with Personalized Situated Cognition Self-Refine Instruction-Tuning for Aligning Reasoning in Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:40.144459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:40.144459Z digest=sha256:534ddf02173024d4edaf0950e2afc6bb1cdd83639141d41ad5aa64b997202f21

Observation ab441365-4ef7-47a3-b4fb-0541dcf8e75a · inbound

Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives cites this paper.

Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives Self-Refine Instruction-Tuning for Aligning Reasoning in Language Models

Reference 127

Resolution
unresolved
no resolver link, observed 2026-08-07T04:46:04.497694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:46:04.497694Z digest=sha256:c988ae9168940248d1fe18770a671e727920ba26d227769b32f77343839ce590

Observation dab931b3-230b-4ed7-a9b9-54a967789bb1 · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Self-Refine Instruction-Tuning for Aligning Reasoning in Language Models

Reference 289

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:13:08.054384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:13:07.314848Z digest=sha256:09df2dc2e6413e62bcc12f49410d6f8c2a81214fc3978497d8845147c450d8a0

Observation ae5bcf23-54a1-4ab0-9bea-cd0f8205249e · inbound

Future-KL Regularized GRPO: Process-Level Credit Assignment from $f$-Divergence Regularization cites this paper.

Future-KL Regularized GRPO: Process-Level Credit Assignment from $f$-Divergence Regularization Self-Refine Instruction-Tuning for Aligning Reasoning in Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T10:27:24.197217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:27:24.197217Z digest=sha256:026c6286787512fb72841dae204238ba219f60731796245eba734e7052923b50