Pith. sign in

Paper Citation Record · LEDGER

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning

As of 16 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2607.21971.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21971 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T06:17:30.923859Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 53922a82-2efd-40f5-bf5b-a932d4c9fb47 · outbound

This paper cites Tool-r0: Self-evolving llm agents for tool-learning from zero data.arXiv preprint arXiv:2602.21320,.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Tool-r0: Self-evolving llm agents for tool-learning from zero data.arXiv preprint arXiv:2602.21320,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:27.893513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:27.893513Z digest=sha256:471c93cf0698a942462bd0122f927f14ffb257b033f3b4f21576b4d8b54554ce

Observation e454a080-4ae4-404d-a318-2b9023df8b0f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:28.286504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:28.286504Z digest=sha256:37e4cfd05809ffd6d92d2d63f8519055cd9a507003d4e29f045c8c8d6fc89d54

Observation b6fa36e3-4de7-41e5-a24e-a19dce39c1f2 · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:28.461299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:28.461299Z digest=sha256:de75c9c5afe76286e1f0aa8df53655ca92b9a785cf9d004b11240fa5ca6f8641

Observation 000030d2-0b33-4d6f-aa88-676fcb5bbf4e · outbound

This paper cites TACO: Topics in Algorithmic COde generation dataset.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning TACO: Topics in Algorithmic COde generation dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:28.636771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:28.636771Z digest=sha256:24958b23d26d68210c0a15e9cc8c414576a537ac3a1f38d2c075d5d45ae52c7a

Observation 435da973-f07f-41fe-a144-ae5e61668802 · outbound

This paper cites s1: Simple test-time scaling.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning s1: Simple test-time scaling

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:28.698988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:28.698988Z digest=sha256:de944525322ab7965c22b5784307e78d710c84bacbeca9029d073fb5633899c7

Observation bfd4648f-a578-497b-910d-7a094c0f71ee · outbound

This paper cites AlphaEvolve: A coding agent for scientific and algorithmic discovery.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning AlphaEvolve: A coding agent for scientific and algorithmic discovery

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:28.779932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:28.779932Z digest=sha256:a36c08b740d1509cb1760931dae787448cf46d1c7596bd3e823d758b386befc5

Observation 8d736a13-c85b-4f74-8951-abb03b108e4e · outbound

This paper cites Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:28.952018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:28.952018Z digest=sha256:6eb78ff927d2396b9cadb80180d6599826c741e257b2254080f8b80427fea92f

Observation e1f3796a-0c3e-46af-b599-7674d5e75dbe · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:29.146347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:29.146347Z digest=sha256:103a2ebd242ef490caaec5b917f14a29e17cb54be16b619aef5b6997295a89d1

Observation 173508c7-8ccc-400f-8a0e-1c8cfbaf5784 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:29.227523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:29.227523Z digest=sha256:b882e933ec78d41efb7d4606ca9bd9a552a42359f45de15b07a2f7e623b32003

Observation 6538a143-6205-42b7-b0ea-e1734f4021e0 · outbound

This paper cites Can Language Models Solve Olympiad Programming?.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Can Language Models Solve Olympiad Programming?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:29.317498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:29.317498Z digest=sha256:aa17f3f21a2442c8e9fcbccb826efefb1baff1ab16afe8d479393a3ee44f1f12

Observation ecfe5ef8-f368-4708-a020-1d19cc1f19c8 · outbound

This paper cites Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:29.411352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:29.411352Z digest=sha256:9c4d9989d2cec2cf6b99465d06395d7f8f05c1f76af341321469aac5e0b77a3e

Observation e9b24a70-7188-4c75-9263-5862a13fbec2 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:29.493127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:29.493127Z digest=sha256:8494a758db8ae92514cfaf8fd69dadb6dded5c167085bc26fe82a357c0bb8193

Observation 4669a225-09d9-4622-864f-4cb0d61d9aec · outbound

This paper cites Large language model reasoning failures.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Large language model reasoning failures

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:29.584985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:29.584985Z digest=sha256:4430bbe7b8d940dd87fd4d854074e793282f1775fc72cd4d33e1855fdbc7c126

Observation 2b6e34b7-b0d0-4270-8d81-542ffb413ae6 · outbound

This paper cites Under review.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Under review

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:29.668687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:29.668687Z digest=sha256:7f19ca39075ef1b9110d8a35091521fe6552b252c9ca019df8aedef10a46a896

Observation d221d477-c10b-4927-91fe-6ba648b9939b · outbound

This paper cites ThetaEvolve: Test-time Learning on Open Problems.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning ThetaEvolve: Test-time Learning on Open Problems

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:29.754644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:29.754644Z digest=sha256:50c4c6f8499522d84aca497829acede83bb112bfdfbd0e9b2ee121ff57a0400c

Observation 63972755-8e1c-4bd2-b8cb-66c9de8a73bd · outbound

This paper cites LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:29.782230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:29.782230Z digest=sha256:378ae44782632c9347909cd3dd46f8f117983e79776ba462be1cfc419a47d245

Observation c443b250-a9f3-4583-95bd-86a5d31f7d72 · outbound

This paper cites Qwen3 Technical Report.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Qwen3 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:29.938488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:29.938488Z digest=sha256:0cb6d4b9fb9793bcfb94fb694357885b5f2c3b9ca9ce8e781008cc82ca2ff6eb

Observation 1d592169-b77b-4224-8570-132ba46210a4 · outbound

This paper cites ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:30.221262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:30.221262Z digest=sha256:cb457239c1f6a9d70ccf0f637a7935ca9509e0bcff73dddf8986a234184194bb

Observation fb439d78-5096-4625-ab82-618defb3f50a · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:30.431790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:30.431790Z digest=sha256:1a75e4647f9fb14dea2d8881e678ccc3e35a58df2fc99de3ea1a957406006950

Observation 25b3f886-63a0-4844-9618-e8ea366327c8 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Fine-Tuning Language Models from Human Preferences

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:30.601405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:30.601405Z digest=sha256:e176fb52815ce36f24ec81c4117806a173b476dc70cc207372b567e81ef14c1a

Observation 3b0629b3-dfb0-46ce-bba4-bd80c8a03f01 · outbound

This paper cites an unresolved cited work.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:30.923859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:30.923859Z digest=sha256:04910dbd90f43d9d04b50d867e8bf38e304def5ec5831fc4c4f415c69dc6e310

Observation f2543b24-09ef-4288-bffd-db2be8b8fe91 · outbound

This paper cites A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?

Reference 1989

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:30.106561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:30.106561Z digest=sha256:d1bf3e64a966c4a075dd5c52782336c0dbcdc1db4b87c37e287d98bdc33dfb9a

Observation 0690d2b4-9e15-4ce8-b129-4e2ec6d75196 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Process Reinforcement through Implicit Rewards

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:28.092385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:28.092385Z digest=sha256:40658c4b0aabc43e9fab2a41d53e3785d793e8c79e83cf82301bb7daa0edbc1c

Observation 21158904-90a0-4cd4-874d-a9b7fab6cf84 · outbound

This paper cites Under review.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Under review

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:30.720072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:30.720072Z digest=sha256:cc8191556294503c59c9bf373363f20ca593395c590f623e46412b9b4446867d

Observation 7628df11-b20d-406b-bd58-a5252f404447 · outbound

This paper cites CodeBLEU: a Method for Automatic Evaluation of Code Synthesis.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning CodeBLEU: a Method for Automatic Evaluation of Code Synthesis

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:29.063792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:29.063792Z digest=sha256:cf0f72e121eb0353e5977b331a332199ab7746a35dc3f83ab40628d16e41e339

Observation ceee4493-3caf-4ecd-bbb1-9534e9b2cfe0 · outbound

This paper cites Ale-bench: A benchmark for long-horizon objective-driven algorithm engineering.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Ale-bench: A benchmark for long-horizon objective-driven algorithm engineering

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:28.376050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:28.376050Z digest=sha256:71fab0e178fdd0950c535ad85b2be45ea773efd1145d6e331293ff97729a7fab

Observation bf6c3b84-5ab0-4514-9465-41eb81c205c5 · outbound

This paper cites Qwen3-Coder-Next Technical Report.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Qwen3-Coder-Next Technical Report

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:28.013414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:28.013414Z digest=sha256:b2c2e66277a9dff2661856ae792008b87a86a1a934a43fa17929f363b1e27c81

Observation 4a99526a-bf74-4095-8a51-4b4eb8ceb6b6 · outbound

This paper cites Deltaevolve: Accelerating scientific discovery through momentum-driven evolution.arXiv preprint arXiv:2602.02919,.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Deltaevolve: Accelerating scientific discovery through momentum-driven evolution.arXiv preprint arXiv:2602.02919,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:28.551317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:28.551317Z digest=sha256:90d7131266ef2bab3d3dd3c959ea57c5a759f167a9112aaef003e26a5c8c82aa

Observation c328d9e8-e157-4951-b19d-0384681ffe9e · outbound

This paper cites OpenAI o1 System Card.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning OpenAI o1 System Card

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:28.866319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:28.866319Z digest=sha256:e5d6cd956cfbad7f34ea4381cdbcc7223dba177ea75fd98f59f39159b0266744

Observation 4276ef4b-d744-4847-8398-4e57ee03d566 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:28.178616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:28.178616Z digest=sha256:d765a90b17f83ed3d773c72f88818d9abef43ba2ff7f2b2f9c97850db9c95b13

Observation de564314-a21e-4d84-82d8-c59ab98209fc · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Constitutional AI: Harmlessness from AI Feedback

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T06:17:27.935566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T06:17:27.935566Z digest=sha256:672e814d04128b5c94a98fe0b2d5594404504dfdde1c8b9ee680cc5b45023c9e

Pith citing papers

No inbound Pith citation observations are available.