Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:22:55.544927Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 39 inbound Pith citation observations for arXiv:2508.12790.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T17:22:55.544927Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:05.606844Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
43 of 43 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 39a79538-4557-480e-82b3-c0a67eda5aab · outbound
Reinforcement Learning with Rubric Anchors write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfcc1da7-b2b6-4590-90d9-aac1e127e120 · outbound
Reinforcement Learning with Rubric Anchors HealthBench: Evaluating Large Language Models Towards Improved Human Health
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56b95a8e-0acf-4eaf-98c2-86988a458fdb · outbound
Reinforcement Learning with Rubric Anchors Constitutional AI: Harmlessness from AI Feedback
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33dd6bd7-fb08-497e-b1dc-f728faa53080 · outbound
Reinforcement Learning with Rubric Anchors Tombench: Benchmarking theory of mind in large language models, 2024
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d7a7ed89-1313-40c7-b47a-303bc73f3ed2 · outbound
Reinforcement Learning with Rubric Anchors Gemini models, 2025
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d9febb96-9ebc-45eb-a6e1-4030d16fbd70 · outbound
Reinforcement Learning with Rubric Anchors DeepSeek-V3 Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00e6d183-f73c-4242-bcad-ac5f48f1fa04 · outbound
Reinforcement Learning with Rubric Anchors AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c708330f-c051-4878-a747-a64b4c76c3c0 · outbound
Reinforcement Learning with Rubric Anchors Deliberative Alignment: Reasoning Enables Safer Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b6277b4-526c-4de3-88d5-67dac0393577 · outbound
Reinforcement Learning with Rubric Anchors Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43d70c7b-a10d-464a-b549-a71c0710435d · outbound
Reinforcement Learning with Rubric Anchors DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 233ec089-8316-49eb-b9cf-280e084359db · outbound
Reinforcement Learning with Rubric Anchors Measuring Massive Multitask Language Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69026bb4-5363-4548-bd3e-219769e5e705 · outbound
Reinforcement Learning with Rubric Anchors Measuring Mathematical Problem Solving With the MATH Dataset
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e53c6a11-816d-4e1f-bd6d-e29450ec4153 · outbound
Reinforcement Learning with Rubric Anchors Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a46c880-32ad-4fee-9c00-cb67ae67964a · outbound
Reinforcement Learning with Rubric Anchors LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07ddeaf5-2bc5-4939-a046-627a98985364 · outbound
Reinforcement Learning with Rubric Anchors How Many Instructions Can LLMs Follow at Once?
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0d81586-06cb-4d39-adce-cef428eb8b83 · outbound
Reinforcement Learning with Rubric Anchors Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2550a10-00ec-47ba-85c9-31b0de1f457a · outbound
Reinforcement Learning with Rubric Anchors Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 060eb61e-5f5e-4e77-9b81-04d8dd8e9471 · outbound
Reinforcement Learning with Rubric Anchors Omni-think: Scaling cross-domain generalization in llms via multi-task rl with hybrid rewards
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a48aa7df-3b35-4581-af6c-40f5f09c88d0 · outbound
Reinforcement Learning with Rubric Anchors Deepcoder: A fully open-source 14b coder at o3-mini level, 2025 a
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9515a252-8ddc-4115-8839-b6d01915192b · outbound
Reinforcement Learning with Rubric Anchors Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9519b6ff-a7ad-4619-984b-af40aca38658 · outbound
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dc61ebf7-5bdc-4c5c-ab5d-af87efbe20a4 · outbound
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e1347977-d5ef-4634-a2c0-d3068ce5e254 · outbound
Reinforcement Learning with Rubric Anchors LLM Critics Help Catch LLM Bugs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d0fd8fa-4d9b-47be-8774-13def922938a · outbound
Reinforcement Learning with Rubric Anchors Rule based rewards for language model safety
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 39c4c190-44fc-4bc3-b0a2-7ae0148ac381 · outbound
Reinforcement Learning with Rubric Anchors Learning to reason with llms, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e94348b-0999-4453-a83b-ee51f5279948 · outbound
Reinforcement Learning with Rubric Anchors Introducing openai o3 and o4-mini, 2025
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 20e50f6c-f468-41c6-bc90-a888b4054189 · outbound
Reinforcement Learning with Rubric Anchors EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation add607c4-1b27-4eda-a868-cc8e81c7cc84 · outbound
Reinforcement Learning with Rubric Anchors CoQA: A Conversational Question Answering Challenge
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78cd5978-417d-4a9b-9d09-af04dfd7a264 · outbound
Reinforcement Learning with Rubric Anchors GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea9bb1af-918a-4428-bffa-c97eef429118 · outbound
Reinforcement Learning with Rubric Anchors SocialIQA: Commonsense Reasoning about Social Interactions
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae255f6-9a31-4457-91a9-ecefdb2e7460 · outbound
Reinforcement Learning with Rubric Anchors Self-critiquing models for assisting human evaluators
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45bace34-c9a9-4c44-9866-975703c5cb55 · outbound
Reinforcement Learning with Rubric Anchors A Simple and Effective Approach to the Story Cloze Test
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81c21057-2345-4e1a-89cc-6eb9278261ac · outbound
Reinforcement Learning with Rubric Anchors Salmon: Self-alignment with instructable reward models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e43ff105-eee4-4971-bbe6-e99d87fb61cd · outbound
Reinforcement Learning with Rubric Anchors GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15a57863-b5aa-493a-bb07-fe8b52374975 · outbound
Reinforcement Learning with Rubric Anchors Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36f69961-1b8d-4950-808f-010774c0c55f · outbound
Reinforcement Learning with Rubric Anchors Checklists are better than reward models for aligning language models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be3ff487-5f58-4715-8462-abebdf9164d1 · outbound
Reinforcement Learning with Rubric Anchors Safety reasoning with guidelines
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 68c2858d-672c-48e9-89b2-208d25e43fb8 · outbound
Reinforcement Learning with Rubric Anchors Writingbench: A comprehensive benchmark for generative writing
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97d6c44e-0dda-467e-a7b9-26593fbbe8f1 · outbound
Reinforcement Learning with Rubric Anchors Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e72979b-b199-4153-a6cd-8a8f76563f1a · outbound
Reinforcement Learning with Rubric Anchors Qwen3 Technical Report
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05dbad6d-3e2e-45c7-93ca-b827ff3ec3af · outbound
Reinforcement Learning with Rubric Anchors COLLIE: Systematic Construction of Constrained Text Generation Tasks
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcbd5eb4-cd14-4cc7-94a3-1a28c43a3bb3 · outbound
Reinforcement Learning with Rubric Anchors HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41e378f6-6c00-4e67-8c5f-ef872e0ad9d1 · outbound
Reinforcement Learning with Rubric Anchors Instruction-Following Evaluation for Large Language Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6326af1-e0f9-4191-a4e2-04bd455e6ecc · inbound
Baichuan-M2: Scaling Medical Capability with Large Verifier System Reinforcement Learning with Rubric Anchors
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 018b4547-fad7-4e0d-8e7b-7e451f2ebff9 · inbound
A Survey of Reinforcement Learning for Large Reasoning Models Reinforcement Learning with Rubric Anchors
Reference 215
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9de7eea5-c94b-4642-b31b-8ed74e67d626 · inbound
DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents Reinforcement Learning with Rubric Anchors
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4746e6f1-f1f0-4529-ab0b-84912315f37c · inbound
RubricRL: Simple Generalizable Rewards for Text-to-Image Generation Reinforcement Learning with Rubric Anchors
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6488fe73-7c94-4a9a-8407-0c4765c06b47 · inbound
Verbalizing LLM's Higher-order Uncertainty via Imprecise Probabilities Reinforcement Learning with Rubric Anchors
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e85495b3-c760-4a36-93ab-334b15bc8a38 · inbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Reinforcement Learning with Rubric Anchors
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a7e05a27-7df6-4208-9dac-f1a9d3ef5e80 · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Reinforcement Learning with Rubric Anchors
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 65507374-b0d4-4cc0-b6b6-3bfae866d82b · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Reinforcement Learning with Rubric Anchors
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0674eb1b-f6aa-4b87-ac43-a269823d2265 · inbound
Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Reinforcement Learning with Rubric Anchors
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38f57d8b-89ed-46d4-ba32-3d9c730b232f · inbound
Visual Preference Optimization with Rubric Rewards Reinforcement Learning with Rubric Anchors
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e8d9f12-3a28-4b71-b19a-11c4ed0d7be1 · inbound
C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences Reinforcement Learning with Rubric Anchors
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 16d76370-e6f3-4204-bb10-96dbd490a6b7 · inbound
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence Reinforcement Learning with Rubric Anchors
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation df2ee1d2-1693-4e1f-85b2-8b40ec090896 · inbound
Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Reinforcement Learning with Rubric Anchors
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 88c56139-5a71-4347-88e6-e8ab5c921788 · inbound
SCPRM: A Schema-aware Cumulative Process Reward Model for Knowledge Graph Question Answering Reinforcement Learning with Rubric Anchors
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c3918a16-8ba2-42a6-8562-cd5a1e8e0d23 · inbound
SHARP: A Self-Evolving Human-Auditable Rubric Policy for Financial Trading Agents Reinforcement Learning with Rubric Anchors
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d3e350a5-deb2-4130-8035-99bf15500bf3 · inbound
Rubric-based On-policy Distillation Reinforcement Learning with Rubric Anchors
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 42bde171-b5d2-459e-89e8-bb95268ecd7d · inbound
Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning Reinforcement Learning with Rubric Anchors
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 51564e78-31fb-4dc0-a3a0-559410604284 · inbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Reinforcement Learning with Rubric Anchors
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 9d8d874a-d852-4e7a-ae23-d1f197eeb25b · inbound
CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics Reinforcement Learning with Rubric Anchors
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 42c593ff-22f8-4dfc-8afb-e8052569f067 · inbound
Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants Reinforcement Learning with Rubric Anchors
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 81591429-6b5e-4b23-85e2-969e3ab1a6d9 · inbound
Reward Hacking in Rubric-Based Reinforcement Learning Reinforcement Learning with Rubric Anchors
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 021d35ac-1a5e-407d-9ce2-1bfe14bf133d · inbound
AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment Reinforcement Learning with Rubric Anchors
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 48721ea1-1005-40b4-91ed-325c4ec024ad · inbound
AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment Reinforcement Learning with Rubric Anchors
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a5ff79c8-0193-4555-926d-0264d28a58ea · inbound
Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models Reinforcement Learning with Rubric Anchors
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0cd010db-204d-4801-9824-627fd227dd98 · inbound
Prompt-Level Reward Specifications for Open-Ended Post-Training Reinforcement Learning with Rubric Anchors
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c8279687-10cd-4ea4-98f6-a15847dc1e38 · inbound
Reinforcement Learning with Robust Rubric Rewards Reinforcement Learning with Rubric Anchors
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3941bab8-d0b9-40f1-bfed-1402aadd0e4f · inbound
Deep Research as Rubric for Reinforcement Learning Reinforcement Learning with Rubric Anchors
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 96b9cdbc-9019-4bc3-9646-98fd34835cc5 · inbound
Trust Region On-Policy Distillation Reinforcement Learning with Rubric Anchors
Reference 184
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 07f74a91-5a3c-4680-8b02-058efe9fb726 · inbound
BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents Reinforcement Learning with Rubric Anchors
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 4f8e58e2-58a4-4a08-950f-7dca9e664cd2 · inbound
QUBRIC: Co-Designing Queries and Rubrics for RL Beyond Verifiable Rewards Reinforcement Learning with Rubric Anchors
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation dfd3c8c5-b0dc-411a-a051-e7bb251a8bea · inbound
RUBAS: Rubric-Based Reinforcement Learning for Agent Safety Reinforcement Learning with Rubric Anchors
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5e4e8840-30f9-4b58-a5c2-a0bbfe3639b0 · inbound
Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Reinforcement Learning with Rubric Anchors
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 20f00e88-edfb-4634-be64-74dcfead656b · inbound
SAW: Stage-Aware Dynamic Weighting for Multi-Objective Reinforcement Learning in Large Language Models Reinforcement Learning with Rubric Anchors
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 13feac50-cbab-421b-9a7d-61a4161e476b · inbound
Baichuan-M4: A Clinical-Grade Medical Agent System for Continuous Care Reinforcement Learning with Rubric Anchors
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e63d44ae-0109-4635-a1ee-d950a41873e6 · inbound
ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents Reinforcement Learning with Rubric Anchors
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1fdf1afe-73ec-4da6-a950-c733a942c6f2 · inbound
ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents Reinforcement Learning with Rubric Anchors
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 985a6654-89ee-4175-a9ba-f9daad0930ae · inbound
Teach-to-Reason: Competition-Guided Reasoning with a Self-Improving Teacher Reinforcement Learning with Rubric Anchors
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5ef81299-f809-4b3c-8eaf-dc99e5af92e2 · inbound
LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation Reinforcement Learning with Rubric Anchors
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4831f5a2-441e-4dab-8509-26f24275999b · inbound
Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL Reinforcement Learning with Rubric Anchors
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.