Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T09:48:56.990745Z
Paper Citation Record · LEDGER
As of 3 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 4 inbound Pith citation observations for arXiv:2506.13351.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T09:48:56.990745Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T10:23:23.168854Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T23:07:27.269519Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1c1b743d-b5eb-483e-b48c-637472cb1aaf · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks The pitfalls of next-token prediction
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 179665dc-1fcf-454b-9c3c-6bf43aa56ab8 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Language Models are Hidden Reasoners: Unlocking Latent Reasoning Capabilities via Self-Rewarding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6d886974-da84-403d-9a03-a5b3cba3403c · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Enhancing uncertainty modeling with semantic graph for hallucination detection
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b71a2602-5e4e-42db-bc26-2cf26d7a6d0f · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Think, Prune, Train, Improve: Scaling Reasoning without Scaling Models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 752909ee-b529-4540-8312-0ee5af3e777a · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks KTO: Model Alignment as Prospect Theoretic Optimization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ad3a6eb2-fe66-45df-899e-6c975039dc0e · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks A Survey on LLM-as-a-Judge
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 73fb2433-34c9-418f-b41d-ea294dc85efb · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 76a1baff-bdc8-4954-9157-54deaae169d3 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Language Model Cascades: Token-level uncertainty and beyond
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation cfe89eaf-3165-4c07-abf4-1bc669127b95 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 27fc0b52-f5ff-4bdb-a168-adbcd38d5fa2 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Self-Evolved Reward Learning for LLMs
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 474cd497-e9cc-4eb2-aaf9-b8898aceabd2 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Towards Reasoning in Large Language Models: A Survey
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7202e720-1d9f-4df2-b612-4e6f98e094d0 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks OpenAI o1 System Card
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a8b57cdf-c553-4123-9ce6-413cc63ce350 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks s3: You don’t need that much data to train a search agent via rl
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 05bd0ce7-3127-4dea-9c8a-0f5d40bcb033 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks CASIMIR: A Corpus of Scientific Articles enhanced with Multiple Author-Integrated Revisions
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0cb9471e-feb2-4b79-9286-0d95976ef168 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks ParaRev: Building a dataset for Scientific Paragraph Revision annotated with revision instruction
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 99011ba1-ab7a-4c81-8572-1e6628cf102d · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Log Probabilities Are a Reliable Estimate of Semantic Plausibility in Base and Instruction-Tuned Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 103b764c-d556-4b2b-9bc6-2338fcdc6286 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Gemini 2.5: Our most intelligent AI model
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 707dcb62-1026-45f1-b86b-11462a399b4e · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2786c23c-680d-489f-9c4f-6d2f8954ec77 · outbound
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4aec18e5-fffe-4e03-8174-e28aba8cd4d8 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Understanding R1-Zero-Like Training: A Critical Perspective
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e0278dc4-8130-4281-990a-5a781151c114 · outbound
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 09a83d25-af58-4e89-a882-da1e0dc94b45 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks s1: Simple test-time scaling
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ce90225b-c8f8-479b-bf1c-8504c24fbff9 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Proximal Policy Optimization Algorithms
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7a06cae3-a6b7-4b61-87ee-2ee1a1e5f28d · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c0442bea-760a-4676-97e7-2fc13f8b4d84 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7745fa08-d56a-4cc0-aa9c-f06996f5d17d · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ed179084-8be9-4cb2-abeb-95217cf7a310 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Beyond Verifiable Rewards: Scaling Reinforcement Learning for Language Models to Unverifiable Data
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4e99e11f-e55e-4aa3-80d5-5f162e0ee8da · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks A Stitch in Time Saves Nine: Detecting and Mitigating Hallucinations of LLMs by Validating Low-Confidence Generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 415cfe12-9a22-4b4a-8b22-8b65b53e2caf · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Genius: A Generalizable and Purely Unsupervised Self-Training Framework For Advanced Reasoning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 95300ed3-3406-401a-a6fc-5af3c0cdc81e · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks LIMO: Less is More for Reasoning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c92f8f91-3a7a-4a5d-94f8-4032ea539520 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fa98ac6d-2f5b-4c26-8992-8370e5586ff0 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e9b1f88b-5d58-4037-921a-64db6211d813 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Scaling of Search and Learning: A Roadmap to Reproduce o1 from Reinforcement Learning Perspective
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 74be46e2-347f-41b3-9412-e3c4e10abec1 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Automatic Chain of Thought Prompting in Large Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7200d265-0d91-40f9-b422-a5fac556602d · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Absolute Zero: Reinforced Self-play Reasoning with Zero Data
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 09f1d75b-3d1f-411a-81d3-dc06e696881d · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks Reinforcing General Reasoning without Verifiers
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation cb697b38-80ee-4fda-8a36-4d881b70c9f2 · outbound
Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks TTRL: Test-Time Reinforcement Learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3984fa88-075b-41d5-be28-a20913e35529 · inbound
Likelihood scoring for continuations of mathematical text: a self-supervised benchmark with tests for shortcut vulnerabilities Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8ee4c914-181f-481e-bcdc-8889895a1c0f · inbound
Trust Region On-Policy Distillation Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks
Reference 169
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8d635831-158b-4d2e-9f7c-e682eb7be96e · inbound
Momentum for Reasoning: Dense Intrinsic Signals in Policy Optimization Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f2777606-e344-481d-91d4-74dea9cbc06a · inbound
Open Security Benchmark: Towards Autonomous Enterprise Cyber Defense Direct Reasoning Optimization: Token-Level Reasoning Reflectivity Meets Rubric Gates for Unverifiable Tasks
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.