Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T13:30:36.529139Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 4 inbound Pith citation observations for arXiv:2605.17291.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-20T13:30:36.529139Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T01:47:25.697650Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-02T02:06:27.221887Z
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4330624b-45a4-409d-94c0-2510dc66117a · outbound
Step-wise Rubric Rewards for LLM Reasoning Reward and guidance through rubrics: Promoting exploration to improve multi-domain reasoning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5851d4f6-769c-4b6f-92e9-0d896e34dc6a · outbound
Step-wise Rubric Rewards for LLM Reasoning BabyVision: Visual Reasoning Beyond Language
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 658ac151-cdb4-4ef2-85af-75943987ac57 · outbound
Step-wise Rubric Rewards for LLM Reasoning DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning.Nature, 645: 633–638
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8bf5e72e-4470-4818-a7ab-bd1385747faa · outbound
Step-wise Rubric Rewards for LLM Reasoning Interleaved latent visual reasoning with selective perceptual modeling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5efb6a2e-21fc-4d18-818a-62aa3faac16c · outbound
Step-wise Rubric Rewards for LLM Reasoning Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 719a3b04-0f42-47f0-a0ac-0ff6396aa84f · outbound
Step-wise Rubric Rewards for LLM Reasoning LLM-Rubric: A multidi- mensional, calibrated approach to automated evaluation of natural language texts
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e4e1609c-2cea-49f4-951a-18b01022a77a · outbound
Step-wise Rubric Rewards for LLM Reasoning OlympiadBench: A challenging benchmark for promoting AGI with olympiad- level bilingual multimodal scientific problems
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9e89cd13-aaa9-4a83-a482-c9330a3f3409 · outbound
Step-wise Rubric Rewards for LLM Reasoning Measuring mathematical problem solving with the MATH dataset.Advancesin Neural Information Processing Systems, 34:7294–7305
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 910cc2d2-cada-4c09-a841-2dc0bdd52f99 · outbound
Step-wise Rubric Rewards for LLM Reasoning Prometheus 2: An open source language model specialized in evaluating other language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 88527283-f309-4595-897a-6b934506ff0d · outbound
Step-wise Rubric Rewards for LLM Reasoning Solving quantitative reasoning problems with language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5bd3a20a-4af3-4f08-a1dd-d340701256a6 · outbound
Step-wise Rubric Rewards for LLM Reasoning Let’s verify step by step
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b5fd02f7-d1b5-4b53-a3b7-2f8d8221ec52 · outbound
Step-wise Rubric Rewards for LLM Reasoning Improve Mathematical Reasoning in Language Models by Automated Process Supervision
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6de5dff7-98e3-470e-929c-34a50bf767d0 · outbound
Step-wise Rubric Rewards for LLM Reasoning Learning to reason with LLMs.https://openai.com/index/learning-to-reason-with-llms/
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ee8cf371-d122-43e1-9e8c-205ee8378af7 · outbound
Step-wise Rubric Rewards for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 81ec88ef-28d3-429a-bcb3-6cd1398fe962 · outbound
Step-wise Rubric Rewards for LLM Reasoning From Context to Skills: Can Language Models Learn from Context Skillfully?
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f68524b7-ad15-43ff-9729-cd6186729a94 · outbound
Step-wise Rubric Rewards for LLM Reasoning Longcat-next: Lexicalizing modalities as discrete tokens
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8755cc8f-40e4-40c8-951a-0dde1d36f226 · outbound
Step-wise Rubric Rewards for LLM Reasoning GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 51f037e9-46c3-4fd7-bc1f-d69331ee3aa7 · outbound
Step-wise Rubric Rewards for LLM Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 60c55ce6-0a68-4359-924b-d5c17da0b21c · outbound
Step-wise Rubric Rewards for LLM Reasoning Self-consistency improves chain of thought reasoning in language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d55849a0-be20-4661-94b8-deae9ac5eacb · outbound
Step-wise Rubric Rewards for LLM Reasoning Chain-of-thought prompting elicits reasoning in large language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4f273bfa-a752-401a-868d-710693e26ba2 · outbound
Step-wise Rubric Rewards for LLM Reasoning Simple statistical gradient-following algorithms for connectionist reinforcement learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 41de338a-be42-471d-9703-ae9c83707a3f · outbound
Step-wise Rubric Rewards for LLM Reasoning Grouter: Decoupling Routing from Representation for Accelerated MoE Training
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7ab392d1-4bd7-4711-8f8a-b6e1ba7467da · outbound
Step-wise Rubric Rewards for LLM Reasoning Qwen3 Technical Report
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f41140d8-d977-4be8-95ee-05938c66645e · outbound
Step-wise Rubric Rewards for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation db987f36-1fe8-4fd6-b475-5ceacfa11467 · outbound
Step-wise Rubric Rewards for LLM Reasoning MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3a794f11-ef10-439f-8406-0b7510fb5881 · outbound
Step-wise Rubric Rewards for LLM Reasoning ### Step N
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b8f17991-d620-441a-934c-8944b476935f · outbound
Step-wise Rubric Rewards for LLM Reasoning Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 08dd61c5-18b2-4f2a-8a00-a2cced80459a · outbound
Step-wise Rubric Rewards for LLM Reasoning Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 135e2db5-1a30-4880-826e-773baf10d53f · outbound
Step-wise Rubric Rewards for LLM Reasoning CORRECT” if the step is logically and mathematically correct •“INCORRECT
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 08d8ef37-71af-4ccd-b066-4b8279dd96bb · outbound
Step-wise Rubric Rewards for LLM Reasoning Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cb819e68-bec2-48f3-b03b-1b4f0c59b891 · outbound
Step-wise Rubric Rewards for LLM Reasoning Unresolved cited work
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 750feb92-579e-4e7b-9f18-d465435178f0 · outbound
Step-wise Rubric Rewards for LLM Reasoning valid": true or false
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 36bac34e-5768-497a-b0b8-b2fc2bb8e02a · outbound
Step-wise Rubric Rewards for LLM Reasoning Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 29a93bac-9cb3-410d-83d7-73695b7b3446 · outbound
Step-wise Rubric Rewards for LLM Reasoning Proof.Follows directly from Propositions O.1 and O.2
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation da969cc1-711a-4bd4-a8d6-6d1960a1a9ff · inbound
Grouter: Decoupling Routing from Representation for Accelerated MoE Training Step-wise Rubric Rewards for LLM Reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe052b4f-c69f-4a88-bf87-f8e3ce9a11ff · inbound
Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Step-wise Rubric Rewards for LLM Reasoning
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6086d151-76e3-40ac-9188-d1c93b6227bd · inbound
SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning Step-wise Rubric Rewards for LLM Reasoning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acb8ae21-3960-487b-9a56-fbcdf4db8c5b · inbound
SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning Step-wise Rubric Rewards for LLM Reasoning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.