Pith. sign in

Paper Citation Record · LEDGER

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models

As of 20 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2510.03259.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.03259 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:50:30.746297Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-10T22:46:24.057572Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T22:47:37.001412Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact3
  • verified fuzzy3
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 72f27256-062e-4977-8a49-dd816e3169d6 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.486633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.486633Z digest=sha256:5e964c8e9fecba082802932ca767822d2488cca26e6e7846e6b38894f891d918

Observation bb7d3139-9c8e-4afe-84cf-1f0f727edebb · outbound

This paper cites This finding suggests that the predicted notions can serve as useful cues for problem solving and may enable further performance gains when leveraged during inference.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models This finding suggests that the predicted notions can serve as useful cues for problem solving and may enable further performance gains when leveraged during inference

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:50:31.922281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:50:30.746297Z digest=sha256:1b27a8a4497e9f690922e246d4671ec3a594a41b320a6c82673503a3b321351c

Observation 81a03056-42e1-4202-8f9d-ea5cec19d62b · outbound

This paper cites Rational Metareasoning for Large Language Models.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Rational Metareasoning for Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.511531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.511531Z digest=sha256:7344d2facd448e133717ac3f1698a7dcf253ebcff29d19cbc3b8e5cff8c93ced

Observation 0e14117a-1d47-43aa-be14-4f068fd2bcac · outbound

This paper cites Enhancing Math Reasoning in Small-sized LLMs via Preview Difficulty-Aware Intervention.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Enhancing Math Reasoning in Small-sized LLMs via Preview Difficulty-Aware Intervention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.520995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.520995Z digest=sha256:b03ea49769e95c7bfffadca21996c228e4586696588de8962c6d45952a2754c4

Observation a8aa4e49-60f9-4cfd-8f75-9e34b490732c · outbound

This paper cites Meta-R1: Empowering Large Reasoning Models with Metacognition.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Meta-R1: Empowering Large Reasoning Models with Metacognition

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:50:31.819104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:50:30.527738Z digest=sha256:f1a150619d876a546953c9636b6641fc4660084baa8559483c1c64a55dc4e10c

Observation 060c60f2-3ee7-43e3-9add-38dba8d67927 · outbound

This paper cites Thinkless: LLM Learns When to Think.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Thinkless: LLM Learns When to Think

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.535082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.535082Z digest=sha256:df3f2b963aed8a34e82061ace0361bdd93085552dfc07a6e9873eca9d0e4abbc

Observation 0210e4cd-170e-450d-85b7-e82d1d1bb23a · outbound

This paper cites From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in Large Reasoning Models via Decoupled Reasoning and Control.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models From "Aha Moments" to Controllable Thinking: Toward Meta-Cognitive Reasoning in Large Reasoning Models via Decoupled Reasoning and Control

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T15:50:31.156183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:50:30.550768Z digest=sha256:b4e01171b92b65f64738953994a299f64f39a5b9c2f76b61c18a92c3f20c3a8d

Observation 655d1a57-9d47-49ec-a10d-288d54480e52 · outbound

This paper cites Good Learners Think Their Thinking: Generative PRM Makes Large Reasoning Model More Efficient Math Learner.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Good Learners Think Their Thinking: Generative PRM Makes Large Reasoning Model More Efficient Math Learner

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.556542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.556542Z digest=sha256:c35233b345343c6476bb0e0c41adb0008d5f9e7e5469d2c146d372ceff2f0bff

Observation d8537c46-3728-48f9-872d-01f3385e696c · outbound

This paper cites URLhttps://arxiv.org/abs/2505.18822.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models URLhttps://arxiv.org/abs/2505.18822

Reference 13

Resolution
verified exact
doi, observed 2026-08-15T15:50:31.126155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:50:30.563064Z digest=sha256:5de170ec5c631a528bdd7e266755432570537b2b76e34cb105b0543f804f5d2d

Observation efd4feab-0cce-4f40-be27-6652d92b3363 · outbound

This paper cites How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models How Difficulty-Aware Staged Reinforcement Learning Enhances LLMs' Reasoning Capabilities: A Preliminary Experimental Study

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.575386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.575386Z digest=sha256:2d2b1b7483a16fbb359a3b930fe9b60a7c04f40b56d0af021278860d83f8e138

Observation c9d9a4f5-ae23-40c0-819c-d341314aaf43 · outbound

This paper cites Length-Controlled Margin-Based Preference Optimization without Reference Model.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Length-Controlled Margin-Based Preference Optimization without Reference Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.584550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.584550Z digest=sha256:93320b87091b0edeeb1386954b75945d6df48898fbb8560a3c5cef963d4e77a0

Observation c465e968-7558-4cef-b3ec-86586aae680a · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Understanding R1-Zero-Like Training: A Critical Perspective

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.590430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.590430Z digest=sha256:f850e9053e17f1e1e91a56f167f8a0e5e7a8f1200f18e27c7b27dd1889380776

Observation 28959a3d-2057-4412-8dfc-c17f8898ea18 · outbound

This paper cites Ziyang Ma, Qingyue Yuan, Zhenglin Wang, and Deyu Zhou.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Ziyang Ma, Qingyue Yuan, Zhenglin Wang, and Deyu Zhou

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.597921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.597921Z digest=sha256:2c8f342cae2c2e39ccd1ddc2e15013399f75b88e54742d36e75466a9dcbc9bf3

Observation 1c53317d-d5a3-4a31-bbe8-5a5a6f2aef8e · outbound

This paper cites MeLA: A Metacognitive LLM-Driven Architecture for Automatic Heuristic Design.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models MeLA: A Metacognitive LLM-Driven Architecture for Automatic Heuristic Design

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.626738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.626738Z digest=sha256:da78ac968536fc288573d817b7e13d69eda4747072700ccefa62f2c11592504c

Observation 272acc1f-67a2-4e1c-a94c-5280c1bd8e7d · outbound

This paper cites Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.633542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.633542Z digest=sha256:e24d219e2a3f7aee4df4f091a6c9f220829b56cfb48384fd9b20329a37d47c48

Observation da16f71c-ba51-44e4-8f6f-d5f87a377629 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.639257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.639257Z digest=sha256:576380f170fd00795154066d5c7591e4dde6f06981b08e0e14df1ecc66b026ef

Observation 8dd391c5-f046-48c3-88f3-c235b574892c · outbound

This paper cites Efficient Reinforcement Finetuning via Adaptive Curriculum Learning.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.645256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.645256Z digest=sha256:903af4e1f6bb4f9542973fcf6ac30710002422c611d8dce4041c16e8564154b7

Observation b158b652-4c66-4a5c-8741-98c348e1f969 · outbound

This paper cites Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Mastering Chess and Shogi by Self-Play with a General Reinforcement Learning Algorithm

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.662398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.662398Z digest=sha256:6684301c42b0a5ddfbe25a7d49ca233e189f075d2951d18a0bcfe3e556620369

Observation fc43aa8e-3d76-4735-940d-3b19e1ec9db3 · outbound

This paper cites Proofwriter: Generating implications, proofs, and abductive statements over natural language.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Proofwriter: Generating implications, proofs, and abductive statements over natural language

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.673152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.673152Z digest=sha256:1063d02aeed8cb86bee57a8402e55dceeadea7031be2781771547b9060e1c722

Observation 48a62f97-aebc-442b-8e85-748f3395d87a · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models LLaMA: Open and Efficient Foundation Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.681257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.681257Z digest=sha256:a9b0bb0931b81501d08f832b3049afcdc8eacdffd3717e4d7122ac4009e9790f

Observation 4f0fab65-4878-4005-8546-e25e6efbac5f · outbound

This paper cites Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Learning when to think: Shaping adaptive reasoning in r1-style models via multi-stage rl

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.688457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.688457Z digest=sha256:d54ce867efdbfd127ac8cf647bb02f8fb0c5499e10fc3e84b01e1ef890734444

Observation 63b88547-fa44-4002-a585-e99974574f45 · outbound

This paper cites ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models ReMA: Learning to Meta-think for LLMs with Multi-Agent Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.694124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.694124Z digest=sha256:afcbf9fd1d39550a2191162ac4157ddcc39771a4c504dbe959b77668c1a4ed9d

Observation d8ce3474-982d-46c6-8bf3-d0ee46a5aba6 · outbound

This paper cites Adaptive Deep Reasoning: Triggering Deep Thinking When Needed.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Adaptive Deep Reasoning: Triggering Deep Thinking When Needed

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.701259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.701259Z digest=sha256:af936992274f78e380a8472b8856920b21e741cc12920d66203a1d78e33f4b7a

Observation 84067768-e4ab-44df-a498-07d6b4bca3d6 · outbound

This paper cites Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Just Enough Thinking: Efficient Reasoning with Adaptive Length Penalties Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.707660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.707660Z digest=sha256:20d1dd5625de3b9dcab17ca94a0c4cd7f4173d197d6c1e3fb0cbc34adff65be2

Observation 4e0888dc-3986-4834-ad6e-5912eb61255a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.713753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.713753Z digest=sha256:b9dbaa5d9a58b5fe8f0cca8ce3915c5db03f4a4065f1c5c7a1fb990a78d68078

Observation f66f512b-2ada-40d0-9e3f-5a845488f422 · outbound

This paper cites AdaptThink: Reasoning Models Can Learn When to Think.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models AdaptThink: Reasoning Models Can Learn When to Think

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.720466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.720466Z digest=sha256:6b335fcfc94eb40af67109551f6ae28675a23281d3a9732d36c0520137a4c1c7

Observation de7451e4-54b1-45b0-8633-0a96f4754abd · outbound

This paper cites URLhttps://arxiv.org/abs/2504.09696.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models URLhttps://arxiv.org/abs/2504.09696

Reference 34

Resolution
verified exact
doi, observed 2026-08-15T15:50:30.911216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:50:30.725943Z digest=sha256:ddaff9a77115d1fc73f9126489e352b712bff23f163e1d596ea02ecd99267e33

Observation 1c03559e-9e82-4393-ade9-41941fd42281 · outbound

This paper cites Analytical reasoning of text.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Analytical reasoning of text

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:50:31.941002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:50:30.731741Z digest=sha256:a2ec5854eb2211c6d141ddaea5867e647deda5f924a05ee74fea9a5a6ea48c6a

Observation 4e85181f-eb2e-4b2f-99c2-46df01202708 · outbound

This paper cites Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Meta-Reasoner: Dynamic Guidance for Optimized Inference-time Reasoning in Large Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.667850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.667850Z digest=sha256:f6898103c80422d59a62065724f35ce029acf9d99258e82cff8a0900d4cb07f4

Observation 82d7913f-c030-4e38-91c5-54a6e7205dd5 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.503057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.503057Z digest=sha256:fc55021d0b67dc0c8d8f36e97ee69f365b22d3a2a6c33ad08450edb225d761bf

Observation 29094c5c-a1a3-405c-9506-78fcd2869bf9 · outbound

This paper cites Logic-lm: Empowering large lan- guage models with symbolic solvers for faithful logical reasoning.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Logic-lm: Empowering large lan- guage models with symbolic solvers for faithful logical reasoning

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:50:31.983802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-15T15:50:30.614165Z digest=sha256:e02af275c1456fc6660b850193702c7a35b7c6b1a69e138c01a62b825b743ef3

Observation f3b4834e-9805-4b7f-bf7c-74ee1bd3fecf · outbound

This paper cites Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Extreme Reasoning Efficiency in Large Language Models.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Aware First, Think Less: Dynamic Boundary Self-Awareness Drives Extreme Reasoning Efficiency in Large Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.494962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.494962Z digest=sha256:305cf461b4dae93f72d464a4bae7136daf6411209a08f56c155c82ae6da26e1f

Observation d795cdaf-31c7-4e31-b172-fad1c89479b0 · outbound

This paper cites Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Meta-Thinking in LLMs via Multi-Agent Reinforcement Learning: A Survey

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.476760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.476760Z digest=sha256:adb04cf9d56e91832cf69f1a4b9bba5013467786447ff629207c0e5c897d4b3e

Observation 2408906d-ff44-4368-87b2-87ed76bb0f58 · outbound

This paper cites doi: 10.18653/v1/2022.findings-naacl.177.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models doi: 10.18653/v1/2022.findings-naacl.177

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.741094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.741094Z digest=sha256:41053b6bfabeac98384df055de78b049b4f2b755c1fdcc9a17f9f1a8bfe8fd78

Observation b1fb6ac6-5973-4bf2-a6e3-a0054b9b547e · outbound

This paper cites Concise: Confidence-guided compression in step-by-step efficient reasoning.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models Concise: Confidence-guided compression in step-by-step efficient reasoning

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.620237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.620237Z digest=sha256:c1b1ff2c433ac6e7715098aa497b80625bf59ef58665a28f3fe6134be61e6f55

Observation 0a5d3ff6-f692-44ed-8222-d88fffd8cbdd · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.542674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.542674Z digest=sha256:740ba94f0050baaaac9b4777890140f5689fd3c296b4068670062ef9b282c8b1

Observation c00beebe-3387-48e6-a33a-6a1fae0f7fa0 · outbound

This paper cites L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning.

Verifying Meta-Awareness via Predictive Rewards in Reasoning Models L1: Controlling How Long A Reasoning Model Thinks With Reinforcement Learning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T15:50:30.465339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:50:30.465339Z digest=sha256:431e0b230a076435a03fbf85685cb13eb4328d95b0db7a602abcbe5d9d0d6966

Pith citing papers

Observation 32c50b34-3160-4960-80d1-44575255273c · inbound

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning cites this paper.

When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning Verifying Meta-Awareness via Predictive Rewards in Reasoning Models

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T22:47:37.015452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-07-10T22:46:24.057572Z digest=sha256:588815c3768b86405e1e5e0e8300fccfb0a433d5e5129be59fb3c4a29c789eb9