Pith. sign in

Paper Citation Record · LEDGER

Step-wise Rubric Rewards for LLM Reasoning

As of 4 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 4 inbound Pith citation observations for arXiv:2605.17291.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.17291 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T13:30:36.529139Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T01:47:25.697650Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T02:06:27.221887Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact15
  • verified fuzzy14
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4330624b-45a4-409d-94c0-2510dc66117a · outbound

This paper cites Reward and guidance through rubrics: Promoting exploration to improve multi-domain reasoning.

Step-wise Rubric Rewards for LLM Reasoning Reward and guidance through rubrics: Promoting exploration to improve multi-domain reasoning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:33:19.249714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:49e5725f40f2169f98f82700fc6a471d784ff7bcbeae8744ab10d2337881cd96

Observation 5851d4f6-769c-4b6f-92e9-0d896e34dc6a · outbound

This paper cites BabyVision: Visual Reasoning Beyond Language.

Step-wise Rubric Rewards for LLM Reasoning BabyVision: Visual Reasoning Beyond Language

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:4497c50daed27f160e6fc4f6945920a9f13728ff5b8fc06d7055aec42b9e2551

Observation 658ac151-cdb4-4ef2-85af-75943987ac57 · outbound

This paper cites DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning.Nature, 645: 633–638.

Step-wise Rubric Rewards for LLM Reasoning DeepSeek-R1: Incentivizing reasoning capability in LLMs via reinforcement learning.Nature, 645: 633–638

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T13:33:29.986902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:c993031b28939d7358c1f2fbc8ead370070de651b8d46af1dc9b2bfe0d01ce9b

Observation 8bf5e72e-4470-4818-a7ab-bd1385747faa · outbound

This paper cites Interleaved latent visual reasoning with selective perceptual modeling.

Step-wise Rubric Rewards for LLM Reasoning Interleaved latent visual reasoning with selective perceptual modeling

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:33:19.231448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:88c15b24d5ad2f468b078ea83370b6f2ba0f764f11bb075e6bda91af3f349986

Observation 5efb6a2e-21fc-4d18-818a-62aa3faac16c · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

Step-wise Rubric Rewards for LLM Reasoning Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:33:19.243229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:48aaa451d9e816027da7e9ad5eef3238a7e9558dc46e91d08e69c60689eb84d1

Observation 719a3b04-0f42-47f0-a0ac-0ff6396aa84f · outbound

This paper cites LLM-Rubric: A multidi- mensional, calibrated approach to automated evaluation of natural language texts.

Step-wise Rubric Rewards for LLM Reasoning LLM-Rubric: A multidi- mensional, calibrated approach to automated evaluation of natural language texts

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T13:33:30.012270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:01d9ab6c343eccd27a63df1af4d3bf11f1b66f219f034d84c68885faf7c7045a

Observation e4e1609c-2cea-49f4-951a-18b01022a77a · outbound

This paper cites OlympiadBench: A challenging benchmark for promoting AGI with olympiad- level bilingual multimodal scientific problems.

Step-wise Rubric Rewards for LLM Reasoning OlympiadBench: A challenging benchmark for promoting AGI with olympiad- level bilingual multimodal scientific problems

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T13:33:30.010078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:3a72bd6172e1947e42800615b844c0bbb88c82695bf6c8e9f6c341d1fd742057

Observation 9e89cd13-aaa9-4a83-a482-c9330a3f3409 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.Advancesin Neural Information Processing Systems, 34:7294–7305.

Step-wise Rubric Rewards for LLM Reasoning Measuring mathematical problem solving with the MATH dataset.Advancesin Neural Information Processing Systems, 34:7294–7305

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T13:33:29.998141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:a4a11b5b9d974e44820548891afb00036bb4677ea5967be71d54e1675653928a

Observation 910cc2d2-cada-4c09-a841-2dc0bdd52f99 · outbound

This paper cites Prometheus 2: An open source language model specialized in evaluating other language models.

Step-wise Rubric Rewards for LLM Reasoning Prometheus 2: An open source language model specialized in evaluating other language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T13:33:30.000034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:07e17a9e64bc108021ba8ed2143b520476d69b0135ded8dc1b7da5de9b977f95

Observation 88527283-f309-4595-897a-6b934506ff0d · outbound

This paper cites Solving quantitative reasoning problems with language models.

Step-wise Rubric Rewards for LLM Reasoning Solving quantitative reasoning problems with language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T13:33:30.001956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:b52875a1c24b4fff28f89eb5ad442a5e1e4a2d68bd8d6cf823c8643314c2d58a

Observation 5bd3a20a-4af3-4f08-a1dd-d340701256a6 · outbound

This paper cites Let’s verify step by step.

Step-wise Rubric Rewards for LLM Reasoning Let’s verify step by step

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T13:33:30.008122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:05ec4bda585f9632062929d02e516b7aed55903346f990bcea3b0a3d359fce6d

Observation b5fd02f7-d1b5-4b53-a3b7-2f8d8221ec52 · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

Step-wise Rubric Rewards for LLM Reasoning Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:33:19.266993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:fba183cb90f788a925aed5f87ff8161b7de39a229c805845ec76a23fcb8d47c6

Observation 6de5dff7-98e3-470e-929c-34a50bf767d0 · outbound

This paper cites Learning to reason with LLMs.https://openai.com/index/learning-to-reason-with-llms/.

Step-wise Rubric Rewards for LLM Reasoning Learning to reason with LLMs.https://openai.com/index/learning-to-reason-with-llms/

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T13:33:29.981099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:2c037fa4fe3ffc21bf75b438fcab5ddbbafd72327117b39c44de50cb5c8e3411

Observation ee8cf371-d122-43e1-9e8c-205ee8378af7 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Step-wise Rubric Rewards for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:33:19.245999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:6316a614d5bb2158ef44323476b84bf2a9951d40826dfde2a08f81af8521e35b

Observation 81ec88ef-28d3-429a-bcb3-6cd1398fe962 · outbound

This paper cites From Context to Skills: Can Language Models Learn from Context Skillfully?.

Step-wise Rubric Rewards for LLM Reasoning From Context to Skills: Can Language Models Learn from Context Skillfully?

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:33:19.234344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:dfb4fc46d29c10a60a2b0e9c65df1db55440e4961415e812ac437c7b5671a87c

Observation f68524b7-ad15-43ff-9729-cd6186729a94 · outbound

This paper cites Longcat-next: Lexicalizing modalities as discrete tokens.

Step-wise Rubric Rewards for LLM Reasoning Longcat-next: Lexicalizing modalities as discrete tokens

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:33:19.270466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:b9a4f980b49e61302c469b6e3c28dc95b3e8726a3e21fc2dc8c01bc2d8c17fea

Observation 8755cc8f-40e4-40c8-951a-0dde1d36f226 · outbound

This paper cites GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning.

Step-wise Rubric Rewards for LLM Reasoning GRPO-VPS: Enhancing Group Relative Policy Optimization with Verifiable Process Supervision for Effective Reasoning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:33:19.228280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:3a0ca043b0b6869a117c4b622b3a997ab3d83e9334e9e761067399e3a3a4c69e

Observation 51f037e9-46c3-4fd7-bc1f-d69331ee3aa7 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Step-wise Rubric Rewards for LLM Reasoning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:33:19.253393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:3849f403f3ef8ea43ab29f7fdde43a980eb600ab9eca27c8381813c253e27800

Observation 60c55ce6-0a68-4359-924b-d5c17da0b21c · outbound

This paper cites Self-consistency improves chain of thought reasoning in language models.

Step-wise Rubric Rewards for LLM Reasoning Self-consistency improves chain of thought reasoning in language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T13:33:29.984784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:53d49fef0ccc8f0ed9bb31ddfc15b63ccdb95c8dc38227cb6ec8510abde0cba7

Observation d55849a0-be20-4661-94b8-deae9ac5eacb · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Step-wise Rubric Rewards for LLM Reasoning Chain-of-thought prompting elicits reasoning in large language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T13:33:29.996149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:a5d32d09eef358f307df3346fd0dc7d51b8e2cc88fcfce86cfec61ce2c91395f

Observation 4f273bfa-a752-401a-868d-710693e26ba2 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.

Step-wise Rubric Rewards for LLM Reasoning Simple statistical gradient-following algorithms for connectionist reinforcement learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T13:33:29.979214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:9ef1f0152124562de2850e396acd516cb5b453285fe5c3729b69b50a4b690596

Observation 41de338a-be42-471d-9703-ae9c83707a3f · outbound

This paper cites Grouter: Decoupling Routing from Representation for Accelerated MoE Training.

Step-wise Rubric Rewards for LLM Reasoning Grouter: Decoupling Routing from Representation for Accelerated MoE Training

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-26T02:03:10.990663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:241d5bd2dc3457bfe843740ae765314a4dde7e947036b6ddd0d7129f9f2dcbfa

Observation 7ab392d1-4bd7-4711-8f8a-b6e1ba7467da · outbound

This paper cites Qwen3 Technical Report.

Step-wise Rubric Rewards for LLM Reasoning Qwen3 Technical Report

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:33:19.259761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:62f1bba9d2aa5ce07d39052246827c718f22d0a6b34b6e1accfbdda8ae7eaec0

Observation f41140d8-d977-4be8-95ee-05938c66645e · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Step-wise Rubric Rewards for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-20T13:33:19.273497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:8e31d8763f3d2e8cfcb382d812cf27583c89ee98b320dc09e6036705430bb0d6

Observation db987f36-1fe8-4fd6-b475-5ceacfa11467 · outbound

This paper cites MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models.

Step-wise Rubric Rewards for LLM Reasoning MENTOR: Efficient Multimodal-Conditioned Tuning for Autoregressive Vision Generation Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:54.697622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:c015ab9dab2ae5a66a860b2fcb24992d34c72f58cbeacb4cd64b16fddfa97b64

Observation 3a794f11-ef10-439f-8406-0b7510fb5881 · outbound

This paper cites ### Step N.

Step-wise Rubric Rewards for LLM Reasoning ### Step N

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T13:33:29.982894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:4ded3c2e6259835525c36f35df875cebe16af6ec8865731b4c80eb93c88b530e

Observation b8f17991-d620-441a-934c-8944b476935f · outbound

This paper cites an unresolved cited work.

Step-wise Rubric Rewards for LLM Reasoning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-05-20T13:33:29.977125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:5bf9ea13888f1142217ad703cb57512d55770b484fb1e4f75d7d603903c7429e

Observation 08dd61c5-18b2-4f2a-8a00-a2cced80459a · outbound

This paper cites an unresolved cited work.

Step-wise Rubric Rewards for LLM Reasoning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-20T13:33:29.975001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:eece54a507a45bc60f67b5138fb9299c85f3a6fb153eb7f761b201c56f0ef794

Observation 135e2db5-1a30-4880-826e-773baf10d53f · outbound

This paper cites CORRECT” if the step is logically and mathematically correct •“INCORRECT.

Step-wise Rubric Rewards for LLM Reasoning CORRECT” if the step is logically and mathematically correct •“INCORRECT

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T13:33:29.993730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:abd7f95c3353f82a18bc41cf64769af7d18071072bbc867c9a2557e3834e900e

Observation 08d8ef37-71af-4ccd-b066-4b8279dd96bb · outbound

This paper cites an unresolved cited work.

Step-wise Rubric Rewards for LLM Reasoning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-20T13:33:29.989204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:1299dee61f145d77dbd5f7d77bc13ae2d09686d871e8a6943298c75e6eff7d99

Observation cb819e68-bec2-48f3-b03b-1b4f0c59b891 · outbound

This paper cites an unresolved cited work.

Step-wise Rubric Rewards for LLM Reasoning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-20T13:33:30.004003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:e67736db3af1f5062fd521b2b6c69238093d85d6e0b504459d93957631f84d9c

Observation 750feb92-579e-4e7b-9f18-d465435178f0 · outbound

This paper cites valid": true or false.

Step-wise Rubric Rewards for LLM Reasoning valid": true or false

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:33:19.256550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:aa156be5e551c5ac965990f6463431bb3a07d2af69bab5c3dfe2967517c2ef42

Observation 36bac34e-5768-497a-b0b8-b2fc2bb8e02a · outbound

This paper cites an unresolved cited work.

Step-wise Rubric Rewards for LLM Reasoning Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-20T13:33:29.991310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:35d07ee73fe4a61e1535392668ea876e0490a8406664f80e1dbc2e5cdbc69d46

Observation 29a93bac-9cb3-410d-83d7-73695b7b3446 · outbound

This paper cites Proof.Follows directly from Propositions O.1 and O.2.

Step-wise Rubric Rewards for LLM Reasoning Proof.Follows directly from Propositions O.1 and O.2

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T13:33:30.006293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:30:36.529139Z digest=sha256:a3dc33ed16141fc52751e0c79356022a33d9006ff626d51d9ee631f65cc20a92

Pith citing papers

Observation da969cc1-711a-4bd4-a8d6-6d1960a1a9ff · inbound

Grouter: Decoupling Routing from Representation for Accelerated MoE Training cites this paper.

Grouter: Decoupling Routing from Representation for Accelerated MoE Training Step-wise Rubric Rewards for LLM Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T21:50:59.519853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:50:59.519853Z digest=sha256:bb1e30fd31c19373e527615338222e6b8cc48425ace1e4eee5012592406e9ee8

Observation fe052b4f-c69f-4a88-bf87-f8e3ce9a11ff · inbound

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning cites this paper.

Right Makes Might: Aligning Verified Hidden States Empowers RL Reasoning Step-wise Rubric Rewards for LLM Reasoning

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:06:27.223751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:13:21.565082Z digest=sha256:c96e48361c0b4a2a866120981b4840874ecb0151e983b54e5315a0bde59158a2

Observation 6086d151-76e3-40ac-9188-d1c93b6227bd · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning Step-wise Rubric Rewards for LLM Reasoning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-30T18:33:28.644437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T18:33:28.644437Z digest=sha256:668b6871d875a493e7cd74e7988c4525a1310785986accd7b51107fa708a2af9

Observation acb8ae21-3960-487b-9a56-fbcdf4db8c5b · inbound

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning cites this paper.

SERPO: Self-Evolving Rubric Policy Optimization for Open-Ended Test-Time Reinforcement Learning Step-wise Rubric Rewards for LLM Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T01:47:25.697650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:47:25.697650Z digest=sha256:e9d372537997b7bcf16c918b2b5cc8fb259ac6e5455b375784e51907f6eee87a