Pith. sign in

Paper Citation Record · LEDGER

Language Model Self-improvement by Reinforcement Learning Contemplation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2305.14483.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2305.14483 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:57:30.462207Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:56:13.461679Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 28bd4512-e841-440e-9f47-4f4374a6446d · inbound

ORPO: Monolithic Preference Optimization without Reference Model cites this paper.

ORPO: Monolithic Preference Optimization without Reference Model Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:34:04.809974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T09:34:04.394588Z digest=sha256:5a1ebc90117d5496095d9dec91e5c03af5e24f9b3efa02a708b93d0f993ff270

Observation db0cbb67-4e25-44f8-8137-86ad2df344de · inbound

ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection cites this paper.

ReflectEvo: Improving Meta Introspection of Small LLMs by Learning Self-Reflection Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T15:04:40.606818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:04:40.606818Z digest=sha256:e6c783aea689037df2c52356d5c3bfeee22814df73e30e203e638886d6632a34

Observation 146c8736-2eec-4e14-a68a-1608ba13afdb · inbound

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning cites this paper.

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:25.233251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T04:40:23.895430Z digest=sha256:d54049f082e01cb9a60120f6016439d451b7da0d6fa4790caab5ad07f300a567

Observation 935ec621-b806-41bb-8664-74fac4b71b8c · inbound

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning cites this paper.

Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:49:10.282952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T22:48:49.322744Z digest=sha256:3f93d951967d9b7396dfe96bd05b7c9766779c0af6c7eb786c6602a906e4ff10

Observation bb69859f-194a-438b-a684-327ef2e2384c · inbound

SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation cites this paper.

SOLAR: A Self-Optimizing Open-Ended Autonomous Agent for Lifelong Learning and Continual Adaptation Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T11:24:08.664516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T11:21:30.867480Z digest=sha256:951024e9fee485d39050e7eef98da010a439dcd15ea44ef9db35fd60256edeb1

Observation 5b85f5bf-8ad4-41b0-8a03-48eef99ed6c6 · inbound

Improving Collaborative Storytelling with a Multi-Agent Framework Based on Large Language Models cites this paper.

Improving Collaborative Storytelling with a Multi-Agent Framework Based on Large Language Models Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T07:43:14.199980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T07:36:19.786202Z digest=sha256:cf98541d3189478e0e307f05045b0c9ac3c21924986af8a5e898129eeca40fa2

Observation 6911b7cb-e2a2-4332-af9f-7fed2b06b107 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.463123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:3f537d7d4704178eff7a7b007595eb61ebfb14f8a9e47370242e5a2630a6afb4

Observation a1838133-bd51-4c1f-bf27-58ee53ce2e5c · inbound

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR cites this paper.

Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR Language Model Self-improvement by Reinforcement Learning Contemplation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T00:57:30.462207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T00:57:30.462207Z digest=sha256:7f63d596d637d2c1837d99d8c9e96cc34d303548e142fcce5c35c94778e35d70