Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:50:30.746297Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2510.03259.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:50:30.746297Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-10T22:46:24.057572Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T22:47:37.001412Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 72f27256-062e-4977-8a49-dd816e3169d6 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb7d3139-9c8e-4afe-84cf-1f0f727edebb · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models This finding suggests that the predicted notions can serve as useful cues for problem solving and may enable further performance gains when leveraged during inference
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 81a03056-42e1-4202-8f9d-ea5cec19d62b · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Rational Metareasoning for Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e14117a-1d47-43aa-be14-4f068fd2bcac · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Enhancing Math Reasoning in Small-sized LLMs via Preview Difficulty-Aware Intervention
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8aa4e49-60f9-4cfd-8f75-9e34b490732c · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Meta-R1: Empowering Large Reasoning Models with Metacognition
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 060c60f2-3ee7-43e3-9add-38dba8d67927 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Thinkless: LLM Learns When to Think
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0210e4cd-170e-450d-85b7-e82d1d1bb23a · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in Large Reasoning Models via Decoupled Reasoning and Control
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 655d1a57-9d47-49ec-a10d-288d54480e52 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Good Learners Think Their Thinking: Generative PRM Makes Large Reasoning Model More Efficient Math Learner
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8537c46-3728-48f9-872d-01f3385e696c · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models URLhttps://arxiv.org/abs/2505.18822
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation efd4feab-0cce-4f40-be27-6652d92b3363 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d9a4f5-ae23-40c0-819c-d341314aaf43 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Length-Controlled Margin-Based Preference Optimization without Reference Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c465e968-7558-4cef-b3ec-86586aae680a · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Understanding R1-Zero-Like Training: A Critical Perspective
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28959a3d-2057-4412-8dfc-c17f8898ea18 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Ziyang Ma, Qingyue Yuan, Zhenglin Wang, and Deyu Zhou
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c53317d-d5a3-4a31-bbe8-5a5a6f2aef8e · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models MeLA: A Metacognitive LLM-Driven Architecture for Automatic Heuristic Design
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 272acc1f-67a2-4e1c-a94c-5280c1bd8e7d · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da16f71c-ba51-44e4-8f6f-d5f87a377629 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dd391c5-f046-48c3-88f3-c235b574892c · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Efficient Reinforcement Finetuning via Adaptive Curriculum Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b158b652-4c66-4a5c-8741-98c348e1f969 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc43aa8e-3d76-4735-940d-3b19e1ec9db3 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Proofwriter: Generating implications, proofs, and abductive statements over natural language
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48a62f97-aebc-442b-8e85-748f3395d87a · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models LLaMA: Open and Efficient Foundation Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f0fab65-4878-4005-8546-e25e6efbac5f · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63b88547-fa44-4002-a585-e99974574f45 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8ce3474-982d-46c6-8bf3-d0ee46a5aba6 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Adaptive Deep Reasoning: Triggering Deep Thinking When Needed
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84067768-e4ab-44df-a498-07d6b4bca3d6 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e0888dc-3986-4834-ad6e-5912eb61255a · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f66f512b-2ada-40d0-9e3f-5a845488f422 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models AdaptThink: Reasoning Models Can Learn When to Think
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de7451e4-54b1-45b0-8633-0a96f4754abd · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models URLhttps://arxiv.org/abs/2504.09696
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1c03559e-9e82-4393-ade9-41941fd42281 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Analytical reasoning of text
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4e85181f-eb2e-4b2f-99c2-46df01202708 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82d7913f-c030-4e38-91c5-54a6e7205dd5 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29094c5c-a1a3-405c-9506-78fcd2869bf9 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Logic-lm: Empowering large lan- guage models with symbolic solvers for faithful logical reasoning
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f3b4834e-9805-4b7f-bf7c-74ee1bd3fecf · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Extreme Reasoning Efficiency in Large Language Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d795cdaf-31c7-4e31-b172-fad1c89479b0 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2408906d-ff44-4368-87b2-87ed76bb0f58 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models doi: 10.18653/v1/2022.findings-naacl.177
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1fb6ac6-5973-4bf2-a6e3-a0054b9b547e · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Concise: Confidence-guided compression in step-by-step efficient reasoning
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a5d3ff6-f692-44ed-8222-d88fffd8cbdd · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c00beebe-3387-48e6-a33a-6a1fae0f7fa0 · outbound
Verifying Meta-Awareness via Predictive Rewards in Reasoning Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32c50b34-3160-4960-80d1-44575255273c · inbound
When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Verifying Meta-Awareness via Predictive Rewards in Reasoning Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.