Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T10:40:50.239122Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2510.09278.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T10:40:50.239122Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 361a8eb8-0174-4642-b389-c59f74f9f910 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77f1045f-9409-4096-9730-639a51ae7717 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afba2411-c0db-4330-bd1c-911667d6a48b · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82004794-885e-4aa5-a803-aaf580e011e2 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts HuatuoGPT-o1, Towards Medical Complex Reasoning with LLMs
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 399715de-4411-4b6e-a0b8-7346bfde06b6 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Reasoning Models Don't Always Say What They Think
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ff1e785-3ad6-44f2-a00f-3418d956fcfa · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Towards Consistent Natural-Language Explanations via Explanation-Consistency Finetuning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bba21fd-6a2b-42b2-9553-09838920feee · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Bias-Augmented Consistency Training Reduces Biased Reasoning in Chain-of-Thought
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 083f581b-80e7-4a24-93a0-99e2548948f7 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts DeepSeek-V3 Technical Report
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19582271-b0b6-431f-8b71-61e74e59e90f · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62ef4912-82a5-48f2-89a5-b2c5682cfa19 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts ReCode: Reinforcing Code Generation with Reasoning-Process Rewards
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1dc259b-d0ea-4212-8337-94bb634b95ba · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16b77f73-248e-4669-95fe-35d2d5d1f121 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Truthful or Fabricated? Using Causal Attribution to Mitigate Reward Hacking in Explanations
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f66c8e89-2702-4cfc-9bc1-c620cbc15149 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99c30de1-b493-4faf-b77e-26080463a2a8 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts A Survey on LLM-as-a-Judge
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b543e0d-9588-4705-b2ba-d72e6409fbe9 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f742fe26-bddd-4058-9644-5250871c6264 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1610e9e-0199-4a63-8962-402cc25e0f91 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f131002-14e0-4128-b7f8-fbd9143e9613 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts RAG-RL: Advancing Retrieval-Augmented Generation via RL and Curriculum Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc19292c-a333-448f-a6d9-6b92b54af040 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 633c5962-5bbe-4d59-af80-61a5e622d631 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65080f60-5aa7-457e-b13f-db4814363f61 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts What Disease does this Patient Have? A Large-scale Open Domain Question Answering Dataset from Medical Exams
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22b37d82-c8b9-4cda-9e10-2b91e8a3397e · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts PubMedQA: A Dataset for Biomedical Research Question Answering
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c63b186b-05fb-47b6-a7b8-365e099d1d7e · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd420bf7-a3e9-4cb5-bfca-956a24b62bcf · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Prover-Verifier Games improve legibility of LLM outputs
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84786868-3129-4d84-9da9-749e7413b8ec · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Gonzalez, Hao Zhang, and Ion Stoica
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6a6d043-c3c8-4645-8703-761607e6d954 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Reasoning-Table: Exploring Reinforcement Learning for Table Reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7948169f-0219-41a3-b275-fb64d3c67ad5 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3560fc99-bac1-426f-a0dc-af3069e0e8f4 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts LIMR: Less is More for RL Scaling
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ef278a6-c580-4aa6-8739-30e00e1c6b21 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbffb5cc-a76d-4be1-ab53-2bab677d2646 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Frontier Models are Capable of In-context Scheming
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc1b9482-67ca-459a-a80c-163e8712ba4f · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts GPT-4o System Card
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b628e38-ffc3-41ed-882a-cb25f9cd7aed · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts OpenAI o1 System Card
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26cea3b5-6e5c-4162-9666-895705db7df8 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51da2358-127a-4fff-a8c1-6b007bace80c · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a003e3a2-a1a3-4431-917c-c05e411e3e94 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Making Reasoning Matter: Measuring and Improving Faithfulness of Chain-of-Thought Reasoning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a26a8150-b357-4ef7-9f27-7a9b98c5a117 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e32dbfc1-15da-42af-b8bf-71eda8124d6c · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Qwen2.5 Technical Report
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 436f5eb2-4f65-4e35-ba16-5a7cf8f6c83b · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Proximal Policy Optimization Algorithms
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bb43ed4-80ab-46b8-97aa-4f70436b58aa · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 299ae545-5864-445e-9222-3709189eb5a1 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54e137fd-72b3-4241-95c9-53f001ea1304 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5678ad83-14b1-4b82-b552-9a4d16097830 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Crossing the Reward Bridge: Expanding RL with Verifiable Rewards Across Diverse Domains
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 580f43ad-e190-4243-a5c5-cc26060f8c41 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts CodeBoost: Boosting Code LLMs by Squeezing Knowledge from Code Snippets with RL
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36edb7ea-7158-4493-82aa-accde33d1d79 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Chain-of-Probe: Examining the Necessity and Accuracy of CoT Step-by-Step
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b6e4891-99c2-4707-8f6c-8c110fcd1ba1 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91b278bf-e186-4ec8-855e-1a2b944eafe5 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6a232ff-50c6-4fae-bf07-a1a1895c37c8 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts ReasonFlux: Hierarchical LLM Reasoning via Scaling Thought Templates
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43b5f8b9-8e61-4b7c-8c6c-198a7f4189aa · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a78efc93-19c0-4077-8f49-858bcd059b8e · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts From Graph Retrieval to Schema Realization: Counterfactual Validation for Text-to-SPARQL over Heterogeneous Knowledge Graphs
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 297cb829-c30f-433d-a7ca-7164f4b980ad · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Unresolved cited work
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ce2f5b1-d381-4916-9404-e26dce96ef3f · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts Does Learning Mathematical Problem-Solving Generalize to Broader Reasoning?
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 248c655b-998c-4f70-b69e-13c16eb45eb3 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts online" 'onlinestring :=
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd20c3ad-e6b4-43f5-acf4-8e89e68fccb7 · outbound
CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts write newline
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.