Pith. sign in

Paper Citation Record · LEDGER

A Survey on Post-training of Large Language Models

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2503.06072.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.06072 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:12.164371Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4d8180bf-c9a7-4835-ab8d-1e1323796074 · inbound

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment cites this paper.

A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment A Survey on Post-training of Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:12.164371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:12.164371Z digest=sha256:7f5a17d550599638b12546a5d0b86d8bef8469fe29654f5a06c2cd01f08bedb4

Observation 5c41d007-d776-4801-ab67-e3cb6d7bddcc · inbound

SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning cites this paper.

SpatialReasoner: Towards Explicit and Generalizable 3D Spatial Reasoning A Survey on Post-training of Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T05:41:31.003934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:41:31.003934Z digest=sha256:ec8fb8f90c3570649d63760dd1254baf711d22cba08e1bb4f7840500fa353c74

Observation 90576817-ff7f-49ce-ad8d-3f8dd5d8a6dc · inbound

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection cites this paper.

UFO-RL: Uncertainty-Focused Optimization for Efficient Reinforcement Learning Data Selection A Survey on Post-training of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T20:37:48.237252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:37:48.237252Z digest=sha256:62c1c7367eb6af1a378088e93ec61554b208b82c690a9d83383a4963fa0d3bca

Observation 2985e17c-62ae-4f72-b82f-ee393c0cc8da · inbound

Scaling Image and Video Generation via Test-Time Evolutionary Search cites this paper.

Scaling Image and Video Generation via Test-Time Evolutionary Search A Survey on Post-training of Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:49:48.412280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:49:48.412280Z digest=sha256:d8701d921a89503053b5174eafa18aafd42dfb22bf91d825de5f0f0c95113486

Observation 749fe4ba-84ee-499d-b40c-a4addda8c2c6 · inbound

Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents cites this paper.

Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents A Survey on Post-training of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:49:34.413285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:49:34.413285Z digest=sha256:1984c5075ea3e62b246742e7fe9e177827ad01d93e9c5911ca0ba45dc4ac4373

Observation bb41e4cd-3760-4d38-bda8-7dbc1c2cf741 · inbound

Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models cites this paper.

Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models A Survey on Post-training of Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:16.796706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:16.796706Z digest=sha256:b3c465f193594aee9365dd5abc16848627c9a1c60855f1610fbd72a1bcaa1aae

Observation 4e07f3b8-42ad-4c95-960f-c9b06133f978 · inbound

Evaluation of LLMs for mathematical problem solving cites this paper.

Evaluation of LLMs for mathematical problem solving A Survey on Post-training of Large Language Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:09.974387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:09.974387Z digest=sha256:efde8e1709b8ec628184d1a17ec31bbf5176087d2271366472a3cdb4f8ca251b

Observation 8f259394-b12b-49e3-80cc-aafd09da5c70 · inbound

Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning cites this paper.

Thinking in Character: Advancing Role-Playing Agents with Role-Aware Reasoning A Survey on Post-training of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:12.140220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:12.140220Z digest=sha256:5d26cb50e21e2fcbd9349c2650e1d19f983254ae5367da4067b59924bcf411fb

Observation fa508398-e9cd-4894-bbe0-055d2e765f73 · inbound

RACE-Align: Retrieval-Augmented and Chain-of-Thought Enhanced Preference Alignment for Large Language Models cites this paper.

RACE-Align: Retrieval-Augmented and Chain-of-Thought Enhanced Preference Alignment for Large Language Models A Survey on Post-training of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:21:40.383980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:21:40.383980Z digest=sha256:2971d3c185c3b4f9881b1e0494e5bff48079d98b18ab1c9689123de791d84ea1

Observation 5ba99d04-4eaa-4223-bd5f-f6e1e04022ab · inbound

StatsMerging: Statistics-Guided Model Merging via Task-Specific Teacher Distillation cites this paper.

StatsMerging: Statistics-Guided Model Merging via Task-Specific Teacher Distillation A Survey on Post-training of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:00.442702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:45:00.442702Z digest=sha256:64ebbaaf5cea39b8615ec89b0e331505738ed731ba71a76d1e0304cb29efb230

Observation 34e36404-aa2b-42b4-a444-4a6c8b13f947 · inbound

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning cites this paper.

Revisiting Test-Time Scaling: A Survey and a Diversity-Aware Method for Efficient Reasoning A Survey on Post-training of Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T10:42:41.050774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:42:41.050774Z digest=sha256:3ac7acc8576ece99a54d9231e2849c20d8966e1d6227255b381f02aa5ce9b5e8

Observation f034a09e-e2fa-4274-ab49-808463f0e4b1 · inbound

We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems cites this paper.

We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems A Survey on Post-training of Large Language Models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T00:32:36.669532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:32:36.669532Z digest=sha256:21f02f25727d7bac9f45a1b97b3b9529a7707899d392a1d6fe71180bf861a43d

Observation 344e74b7-6b6d-4ac5-b00e-f46ed77ed415 · inbound

Debating Truth: Debate-driven Claim Verification with Multiple Large Language Model Agents cites this paper.

Debating Truth: Debate-driven Claim Verification with Multiple Large Language Model Agents A Survey on Post-training of Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T02:52:56.578307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T02:52:18.207343Z digest=sha256:351594f559be6b1dedb92d6abff5bb1ac299806dedf696721eeecbc6e8913e5e

Observation 382a2fc7-5f2b-49d9-b335-a77c54885048 · inbound

UniDomain: Pretraining a Unified PDDL Domain from Real-World Demonstrations for Generalizable Robot Task Planning cites this paper.

UniDomain: Pretraining a Unified PDDL Domain from Real-World Demonstrations for Generalizable Robot Task Planning A Survey on Post-training of Large Language Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T03:02:00.195093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T03:00:28.619627Z digest=sha256:809931e06e96b271dfb94f73549d6a8722831d8dea7844bbc7eaee3758723438

Observation 0913d030-05ef-494d-9f47-54c58ddeba2e · inbound

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training cites this paper.

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training A Survey on Post-training of Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:45:26.381683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T06:40:51.046965Z digest=sha256:d656110608c7d469fd5f59baa0426e1c5081e6efb5c69e745244088c9d51fdf0

Observation 2200eae5-6411-4a75-924c-f094d4b5c763 · inbound

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training cites this paper.

CapTrack: Multifaceted Evaluation of Forgetting in LLM Post-Training A Survey on Post-training of Large Language Models

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-25T06:45:25.616198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T06:40:51.046965Z digest=sha256:9fa30a30e3cbed9a951d917201c5562f4c05dfa769e9adb109d3a80731e01efc

Observation f8152e74-e4a1-4d65-b15a-50b72014997f · inbound

MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU cites this paper.

MegaTrain: Full Precision Training of 100B+ Parameter Large Language Models on a Single GPU A Survey on Post-training of Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:40:51.686523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:57:25.256574Z digest=sha256:12604ecc5a934689da707ef9f3fb39efaade5c44798151034be10e25b21a53ea

Observation 1a496589-1f1e-4d66-ba05-cf69e8e7ef29 · inbound

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning cites this paper.

Large Language Model Post-Training: A Unified View of Off-Policy and On-Policy Learning A Survey on Post-training of Large Language Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:30:53.473040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:28:58.515666Z digest=sha256:12474b5b6068ef40ca70bfcffabf9a5d17671afe2ed05a42e59809ff32734a06

Observation 9e770bc3-738e-4d8b-a5c6-5308cc17a77c · inbound

Generalization in LLM Problem Solving: The Case of the Shortest Path cites this paper.

Generalization in LLM Problem Solving: The Case of the Shortest Path A Survey on Post-training of Large Language Models

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:39:38.183056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T10:37:45.355872Z digest=sha256:1d544bc90c4ea818767eb8e39feda8fdd161f419706d95f27103e0e8c2951088

Observation 36efc82d-2d95-4097-8bdb-5b428fcc64a3 · inbound

Towards a Data-Parameter Correspondence for LLMs: A Preliminary Discussion cites this paper.

Towards a Data-Parameter Correspondence for LLMs: A Preliminary Discussion A Survey on Post-training of Large Language Models

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:11:20.329052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T06:08:29.034434Z digest=sha256:6e493a48160e2453246179946f24a0d63103058ad189c457965b65d10ae7e7a6

Observation 4c9a9789-2f93-4942-9757-da5d066e8402 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models A Survey on Post-training of Large Language Models

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:21:30.072630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-07T06:30:09.945371Z digest=sha256:8588da244637d4d19e23eb9691ea95a7dfa79b54c16192b7a728e8926e3ce2e0

Observation 3505cd60-a493-4cbd-a313-2372364b0cc9 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models A Survey on Post-training of Large Language Models

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:11:17.120852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T03:12:19.414358Z digest=sha256:7d0caa23b893ae6ccb4e74be549482a35c7636d551b13d00750d96f45ab4e7c6

Observation f1668062-bcbe-4694-8b1f-6088aee1eda1 · inbound

Rethinking Agentic Reinforcement Learning In Large Language Models cites this paper.

Rethinking Agentic Reinforcement Learning In Large Language Models A Survey on Post-training of Large Language Models

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:02:40.755092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T16:58:41.558250Z digest=sha256:ead32b1a49b33190d4fd40b6640351dc36a71e2e3f9ba97270dcc7429ec12970

Observation 9aa2dfce-41db-485e-b939-ab9ecc64a18b · inbound

Theoretical Limits of Language Model Alignment cites this paper.

Theoretical Limits of Language Model Alignment A Survey on Post-training of Large Language Models

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:56.932350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T01:18:37.614335Z digest=sha256:3ed2545f0b13faaabfb3e60f34bd611849c39c71ad4655a24e5f3d8ce42634a2

Observation 7eaf3570-1433-45cc-bb52-90f1dded0cd6 · inbound

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems cites this paper.

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems A Survey on Post-training of Large Language Models

Reference 198

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:51:39.504015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T01:47:40.772146Z digest=sha256:0e85f949b6c0843428fc5c0fb9fd0116a7669bb05ac955b6349adafc61820ed5

Observation baa8338e-4d8c-4e59-b001-0520fc461a9e · inbound

LoopUS: Recasting Pretrained LLMs into Looped Latent Refinement Models cites this paper.

LoopUS: Recasting Pretrained LLMs into Looped Latent Refinement Models A Survey on Post-training of Large Language Models

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:22:28.687150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T07:21:04.820743Z digest=sha256:0e88057d10b25efb763280bfa5e94cf51b3499e3cf7601b29c39dbca23db4d3f

Observation 9e42d6c9-8b0d-4a5d-8a9f-e162057c0f33 · inbound

Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation cites this paper.

Perceive-then-Plan: Layout-as-Policy for Monocular 3D Scene Layout Estimation A Survey on Post-training of Large Language Models

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T23:14:01.326606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T23:11:59.282627Z digest=sha256:27aaac4cdb97ff29dbcc3258d3fa4fddf15e5869b882dba5b98fc074aad33d55

Observation 4e1e8bb6-a282-49be-84c9-eccaaa97a374 · inbound

Towards Precise Intent-Aligned VLA Aerial Navigation via Expert-Guided GRPO cites this paper.

Towards Precise Intent-Aligned VLA Aerial Navigation via Expert-Guided GRPO A Survey on Post-training of Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-01T23:16:24.741003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T14:28:23.759920Z digest=sha256:13c547a1c81309e04135e859eecadf4bd393b6960441568060f451b20941d6f4

Observation 10242a74-1712-47ef-ba48-c1b4056cc28b · inbound

Beyond Uniform Forgetting: A Study of Sequential Direct Preference Optimization Across Preference Settings cites this paper.

Beyond Uniform Forgetting: A Study of Sequential Direct Preference Optimization Across Preference Settings A Survey on Post-training of Large Language Models

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:29:29.878076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T18:02:52.726302Z digest=sha256:77722c5b2545cda3212a20392c6b8a251e1e3cf3ed00d80964655bc6bb60711b

Observation de8bfa77-b442-4afc-b43c-f4255e2bad24 · inbound

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning cites this paper.

Escaping the Self-Confirmation Trap: An Execute-Distill-Verify Paradigm for Agentic Experience Learning A Survey on Post-training of Large Language Models

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T17:20:00.112152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-25T23:49:38.932474Z digest=sha256:488a60b0df527ce86fa113efe581b7d24e15afe4f02958686d9f0aea9eac28f1

Observation 66091e5b-96d4-4368-84cf-fc0c4e4e9527 · inbound

Joint Learning of Experiential Rules and Policies for Large Language Model Agents cites this paper.

Joint Learning of Experiential Rules and Policies for Large Language Model Agents A Survey on Post-training of Large Language Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T14:09:52.567243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T04:33:27.112576Z digest=sha256:e3e9505bf3a2e9e5389d6839855d8d65a80fd57c012ba22c7c0de46664c2fd80

Observation 9afe6952-51de-40dd-8f4c-db80729f7bce · inbound

Large Language Models Have Unreliable Understanding of Software Engineering Terminology cites this paper.

Large Language Models Have Unreliable Understanding of Software Engineering Terminology A Survey on Post-training of Large Language Models

Reference 37

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T20:25:37.271473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-08T20:22:54.983733Z digest=sha256:e542e673da4dd72ec87172424372985a98209feef5b454d4ff73c8da720c734d

Observation 6c04f95a-7dec-4c72-858d-3d7ac5584f24 · inbound

Sound Probabilistic Safety Bounds for Large Language Models cites this paper.

Sound Probabilistic Safety Bounds for Large Language Models A Survey on Post-training of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T10:22:41.523534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:22:41.523534Z digest=sha256:29ef0a104c3990899ce2082cb84d153557a8f37192279788497947a598e73f9b