Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T06:17:30.923859Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2607.21971.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T06:17:30.923859Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 53922a82-2efd-40f5-bf5b-a932d4c9fb47 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Tool-r0: Self-evolving llm agents for tool-learning from zero data.arXiv preprint arXiv:2602.21320,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e454a080-4ae4-404d-a318-2b9023df8b0f · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6fa36e3-4de7-41e5-a24e-a19dce39c1f2 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 000030d2-0b33-4d6f-aa88-676fcb5bbf4e · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning TACO: Topics in Algorithmic COde generation dataset
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 435da973-f07f-41fe-a144-ae5e61668802 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning s1: Simple test-time scaling
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfd4648f-a578-497b-910d-7a094c0f71ee · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning AlphaEvolve: A coding agent for scientific and algorithmic discovery
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d736a13-c85b-4f74-8951-abb03b108e4e · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Rafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D Manning, Stefano Ermon, and Chelsea Finn
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1f3796a-0c3e-46af-b599-7674d5e75dbe · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 173508c7-8ccc-400f-8a0e-1c8cfbaf5784 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6538a143-6205-42b7-b0ea-e1734f4021e0 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Can Language Models Solve Olympiad Programming?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecfe5ef8-f368-4708-a020-1d19cc1f19c8 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9b24a70-7188-4c75-9263-5862a13fbec2 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4669a225-09d9-4622-864f-4cb0d61d9aec · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Large language model reasoning failures
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b6e34b7-b0d0-4270-8d81-542ffb413ae6 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Under review
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d221d477-c10b-4927-91fe-6ba648b9939b · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning ThetaEvolve: Test-time Learning on Open Problems
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63972755-8e1c-4bd2-b8cb-66c9de8a73bd · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning LeetCodeDataset: A Temporal Dataset for Robust Evaluation and Efficient Training of Code LLMs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c443b250-a9f3-4583-95bd-86a5d31f7d72 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Qwen3 Technical Report
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d592169-b77b-4224-8570-132ba46210a4 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning ArCHer: Training Language Model Agents via Hierarchical Multi-Turn RL
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb439d78-5096-4625-ab82-618defb3f50a · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b3f886-63a0-4844-9618-e8ea366327c8 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Fine-Tuning Language Models from Human Preferences
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b0629b3-dfb0-46ce-bba4-bd80c8a03f01 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2543b24-09ef-4288-bffd-db2be8b8fe91 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
Reference 1989
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0690d2b4-9e15-4ce8-b129-4e2ec6d75196 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Process Reinforcement through Implicit Rewards
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21158904-90a0-4cd4-874d-a9b7fab6cf84 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Under review
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7628df11-b20d-406b-bd58-a5252f404447 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning CodeBLEU: a Method for Automatic Evaluation of Code Synthesis
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceee4493-3caf-4ecd-bbb1-9534e9b2cfe0 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Ale-bench: A benchmark for long-horizon objective-driven algorithm engineering
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf6c3b84-5ab0-4514-9465-41eb81c205c5 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Qwen3-Coder-Next Technical Report
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a99526a-bf74-4095-8a51-4b4eb8ceb6b6 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Deltaevolve: Accelerating scientific discovery through momentum-driven evolution.arXiv preprint arXiv:2602.02919,
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c328d9e8-e157-4951-b19d-0384681ffe9e · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning OpenAI o1 System Card
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4276ef4b-d744-4847-8398-4e57ee03d566 · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de564314-a21e-4d84-82d8-c59ab98209fc · outbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning Constitutional AI: Harmlessness from AI Feedback
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.