Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Rubric Anchors

As of 19 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 39 inbound Pith citation observations for arXiv:2508.12790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12790 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:22:55.544927Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 39 of 39 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:05.606844Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 39a79538-4557-480e-82b3-c0a67eda5aab · outbound

This paper cites write newline.

Reinforcement Learning with Rubric Anchors write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.336572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.336572Z digest=sha256:3cdafc81febde9f977e6293bd76294dd6cc818b34b053a91690892d60c7b11d8

Observation dfcc1da7-b2b6-4590-90d9-aac1e127e120 · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

Reinforcement Learning with Rubric Anchors HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.343159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.343159Z digest=sha256:33833412c62ef86b2c3a9bb081f267e2779bde3704f51bce70edfae97e6865df

Observation 56b95a8e-0acf-4eaf-98c2-86988a458fdb · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Reinforcement Learning with Rubric Anchors Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.350400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.350400Z digest=sha256:024c854c5a9b36a28d148ab734a651c5d5153686b849ed3071bd41fce3d9c21b

Observation 33dd6bd7-fb08-497e-b1dc-f728faa53080 · outbound

This paper cites Tombench: Benchmarking theory of mind in large language models, 2024.

Reinforcement Learning with Rubric Anchors Tombench: Benchmarking theory of mind in large language models, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:22:56.432877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:22:55.355350Z digest=sha256:a1a0f92fde4bbd0cd5ddae5cb00d4b6e3bd283f5334a6cb6040061a730fc9c96

Observation d7a7ed89-1313-40c7-b47a-303bc73f3ed2 · outbound

This paper cites Gemini models, 2025.

Reinforcement Learning with Rubric Anchors Gemini models, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:22:56.416452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:22:55.359946Z digest=sha256:77708a99dce5f2a9eb89e81c69707e45de9e21ead02fbb6d72b8e52401d48535

Observation d9febb96-9ebc-45eb-a6e1-4030d16fbd70 · outbound

This paper cites DeepSeek-V3 Technical Report.

Reinforcement Learning with Rubric Anchors DeepSeek-V3 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.365391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.365391Z digest=sha256:36e00454716b75f032d71c77192aefb5e9c30c9fc546180a83db102bcd1dc130

Observation 00e6d183-f73c-4242-bcad-ac5f48f1fa04 · outbound

This paper cites AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning.

Reinforcement Learning with Rubric Anchors AReaL: A Large-Scale Asynchronous Reinforcement Learning System for Language Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.370681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.370681Z digest=sha256:88365e37dcb69dcf1891b8638e5b8f333a82e13ca4a302d4428c1402f68b226f

Observation c708330f-c051-4878-a747-a64b4c76c3c0 · outbound

This paper cites Deliberative Alignment: Reasoning Enables Safer Language Models.

Reinforcement Learning with Rubric Anchors Deliberative Alignment: Reasoning Enables Safer Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.376142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.376142Z digest=sha256:a23077c67a7a112a98353ef4ecc530946e54a057e5655d847f4f6cf133d3c858

Observation 8b6277b4-526c-4de3-88d5-67dac0393577 · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

Reinforcement Learning with Rubric Anchors Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.380630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.380630Z digest=sha256:d3a0fc078adb615242bf460385034336bab78a5d9e189f8c62c48eacc1538a56

Observation 43d70c7b-a10d-464a-b549-a71c0710435d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Reinforcement Learning with Rubric Anchors DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.385201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.385201Z digest=sha256:f4ef29367ad1c7e7b72bee8f3c8d25e90c9d05e0f78a2151f05011aa5ad65e1d

Observation 233ec089-8316-49eb-b9cf-280e084359db · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Reinforcement Learning with Rubric Anchors Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.390625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.390625Z digest=sha256:deba63acf23a726c45f08bc6076b2f2ddc4f24a8052b33305d1841ec036b1593

Observation 69026bb4-5363-4548-bd3e-219769e5e705 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Reinforcement Learning with Rubric Anchors Measuring Mathematical Problem Solving With the MATH Dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.395475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.395475Z digest=sha256:24f4080815aafe96e7cb420d6c4c3b4ee738401bd42ab7226e64d41497cc80fa

Observation e53c6a11-816d-4e1f-bd6d-e29450ec4153 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

Reinforcement Learning with Rubric Anchors Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.400964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.400964Z digest=sha256:114c7c98077d03faff320c4d7416bb39cf3a75b551b4d99f46b145a8df66a012

Observation 8a46c880-32ad-4fee-9c00-cb67ae67964a · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

Reinforcement Learning with Rubric Anchors LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.405357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.405357Z digest=sha256:512b739b5319c16bbcf2576cefb5817943d6534a8df36f7f2eb999a8b05fd2f9

Observation 07ddeaf5-2bc5-4939-a046-627a98985364 · outbound

This paper cites How Many Instructions Can LLMs Follow at Once?.

Reinforcement Learning with Rubric Anchors How Many Instructions Can LLMs Follow at Once?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.411192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.411192Z digest=sha256:e9b4df9ef339120e8f91fc821b8478f28a5b5420a5518b0c2e9ea7f6418a3727

Observation d0d81586-06cb-4d39-adce-cef428eb8b83 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

Reinforcement Learning with Rubric Anchors Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.416169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.416169Z digest=sha256:aa05921201c732507a60ab266dddcba903538509c476229de821ff0578962ce8

Observation f2550a10-00ec-47ba-85c9-31b0de1f457a · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Reinforcement Learning with Rubric Anchors Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.421350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.421350Z digest=sha256:bb63b2b42facbde7a46ca71671b2082b08c265b14b11490d991701be35146b26

Observation 060eb61e-5f5e-4e77-9b81-04d8dd8e9471 · outbound

This paper cites Omni-think: Scaling cross-domain generalization in llms via multi-task rl with hybrid rewards.

Reinforcement Learning with Rubric Anchors Omni-think: Scaling cross-domain generalization in llms via multi-task rl with hybrid rewards

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.426205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.426205Z digest=sha256:52bba47bb05446c6014ef59b5da54eb4dadc2e5f1f7b45a9012258b1bec3826a

Observation a48aa7df-3b35-4581-af6c-40f5f09c88d0 · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level, 2025 a.

Reinforcement Learning with Rubric Anchors Deepcoder: A fully open-source 14b coder at o3-mini level, 2025 a

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:22:56.389978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:22:55.431184Z digest=sha256:65ce55a9e40e2b1b3d74a1de6ef3af2ede01dd390e97b05da29cbe27393d4b9a

Observation 9515a252-8ddc-4115-8839-b6d01915192b · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

Reinforcement Learning with Rubric Anchors Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:22:56.372314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:22:55.435886Z digest=sha256:eb551654baeaf46e4cb2f38e4330bea1c36763c5850e88586439b78ef85f48c5

Observation 9519b6ff-a7ad-4619-984b-af40aca38658 · outbound

This paper cites Aime 2024.

Reinforcement Learning with Rubric Anchors Aime 2024

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:22:56.356308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:22:55.440494Z digest=sha256:8419eab649443021f263ab67b8c0c846b96b46184b35f2d5c743a03cd72833db

Observation dc61ebf7-5bdc-4c5c-ab5d-af87efbe20a4 · outbound

This paper cites Aime 2025.

Reinforcement Learning with Rubric Anchors Aime 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:22:56.339508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:22:55.444921Z digest=sha256:4f4322299842c82adcdf0dce0aaa5936eb957ee266e92af89f53b63fd4e05e14

Observation e1347977-d5ef-4634-a2c0-d3068ce5e254 · outbound

This paper cites LLM Critics Help Catch LLM Bugs.

Reinforcement Learning with Rubric Anchors LLM Critics Help Catch LLM Bugs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.449553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.449553Z digest=sha256:21e705ff978ab8f432d60e83f34d746e518a86086f47222c34a8fe6036a0765c

Observation 9d0fd8fa-4d9b-47be-8774-13def922938a · outbound

This paper cites Rule based rewards for language model safety.

Reinforcement Learning with Rubric Anchors Rule based rewards for language model safety

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:22:56.323716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:22:55.454605Z digest=sha256:de40daa4e6bd283f20bf70117ab39b3b989f056848bb629d15e7613349633425

Observation 39c4c190-44fc-4bc3-b0a2-7ae0148ac381 · outbound

This paper cites Learning to reason with llms, 2024.

Reinforcement Learning with Rubric Anchors Learning to reason with llms, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.459222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.459222Z digest=sha256:876bdadb1e69db44b0b89dffd5012687a3d6a1b40cf41238a988258daf122068

Observation 9e94348b-0999-4453-a83b-ee51f5279948 · outbound

This paper cites Introducing openai o3 and o4-mini, 2025.

Reinforcement Learning with Rubric Anchors Introducing openai o3 and o4-mini, 2025

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:22:56.298387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:22:55.463807Z digest=sha256:a2ad22239fc49cfc7d2d555f5168229abcb9f2f2dc726919fb0eaf3e5b169124

Observation 20e50f6c-f468-41c6-bc90-a888b4054189 · outbound

This paper cites EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models.

Reinforcement Learning with Rubric Anchors EQ-Bench: An Emotional Intelligence Benchmark for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.468458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.468458Z digest=sha256:e72980037154adef5cd1b948d895d229074702091ba553c885475eff142ee428

Observation add607c4-1b27-4eda-a868-cc8e81c7cc84 · outbound

This paper cites CoQA: A Conversational Question Answering Challenge.

Reinforcement Learning with Rubric Anchors CoQA: A Conversational Question Answering Challenge

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.473270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.473270Z digest=sha256:db9788030475d59ff200cb7a5e8528969482b0a14bcdfd77fffbbfde104086b2

Observation 78cd5978-417d-4a9b-9d09-af04dfd7a264 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Reinforcement Learning with Rubric Anchors GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.477827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.477827Z digest=sha256:e4e1df527310f4c2f776a62f50c5b2b3921090cbb4256d5c0c205c1e6d7350a3

Observation ea9bb1af-918a-4428-bffa-c97eef429118 · outbound

This paper cites SocialIQA: Commonsense Reasoning about Social Interactions.

Reinforcement Learning with Rubric Anchors SocialIQA: Commonsense Reasoning about Social Interactions

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.482483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.482483Z digest=sha256:ad031b37a176cb51fed4c17ee3de384b92a38756bef2ac0fe986c50fb2dba7e6

Observation 9ae255f6-9a31-4457-91a9-ecefdb2e7460 · outbound

This paper cites Self-critiquing models for assisting human evaluators.

Reinforcement Learning with Rubric Anchors Self-critiquing models for assisting human evaluators

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.487329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.487329Z digest=sha256:9d926364948e44e05742bad6c01ab48334b225fb3821d6b4db53969f4c1ef96e

Observation 45bace34-c9a9-4c44-9866-975703c5cb55 · outbound

This paper cites A Simple and Effective Approach to the Story Cloze Test.

Reinforcement Learning with Rubric Anchors A Simple and Effective Approach to the Story Cloze Test

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.492196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.492196Z digest=sha256:12a2f5ff238d196c630987d4a186e81f2d76c48faf6a8cc567c93bf5c43176cd

Observation 81c21057-2345-4e1a-89cc-6eb9278261ac · outbound

This paper cites Salmon: Self-alignment with instructable reward models.

Reinforcement Learning with Rubric Anchors Salmon: Self-alignment with instructable reward models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:22:56.283635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:22:55.497019Z digest=sha256:e5cb545b727a91a76732bd6bff6113fa06e5c2904ef84a57729c21ced4f3d2ad

Observation e43ff105-eee4-4971-bbe6-e99d87fb61cd · outbound

This paper cites GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models.

Reinforcement Learning with Rubric Anchors GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.501479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.501479Z digest=sha256:eab16a08edeee5c3aeb1f98dd1321e7e1bf7f3cf88a8717137611c7306b26ba0

Observation 15a57863-b5aa-493a-bb07-fe8b52374975 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Reinforcement Learning with Rubric Anchors Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.506644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.506644Z digest=sha256:f4065010b26c63a476c675fa4af1ddf468859b82b7a97b176c4609ac278395a6

Observation 36f69961-1b8d-4950-808f-010774c0c55f · outbound

This paper cites Checklists are better than reward models for aligning language models.

Reinforcement Learning with Rubric Anchors Checklists are better than reward models for aligning language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.511371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.511371Z digest=sha256:3905623952610cac80929c362725f77ea34237a1e1d48d5256b6058bd7b17e40

Observation be3ff487-5f58-4715-8462-abebdf9164d1 · outbound

This paper cites Safety reasoning with guidelines.

Reinforcement Learning with Rubric Anchors Safety reasoning with guidelines

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:22:56.268139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-15T17:22:55.515811Z digest=sha256:cad57115f493f8effa73ee603ce6486ccf3d8c8ee1eecb9c87b8bf51693a4495

Observation 68c2858d-672c-48e9-89b2-208d25e43fb8 · outbound

This paper cites Writingbench: A comprehensive benchmark for generative writing.

Reinforcement Learning with Rubric Anchors Writingbench: A comprehensive benchmark for generative writing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.520571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.520571Z digest=sha256:08868fcb3a2e20c0c270c921e649053642a32d94460f4d7fa3e81264fb3c479e

Observation 97d6c44e-0dda-467e-a7b9-26593fbbe8f1 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

Reinforcement Learning with Rubric Anchors Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.525007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.525007Z digest=sha256:953cd75dc9c3c1816d2c29456fab71b6b734fe42daa4aa159165f3c853225124

Observation 1e72979b-b199-4153-a6cd-8a8f76563f1a · outbound

This paper cites Qwen3 Technical Report.

Reinforcement Learning with Rubric Anchors Qwen3 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.530083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.530083Z digest=sha256:c8644470c5d4eb097a9aa284be9db95950b448b76fa8f7b9d283882cf5b01463

Observation 05dbad6d-3e2e-45c7-93ca-b827ff3ec3af · outbound

This paper cites COLLIE: Systematic Construction of Constrained Text Generation Tasks.

Reinforcement Learning with Rubric Anchors COLLIE: Systematic Construction of Constrained Text Generation Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.534857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.534857Z digest=sha256:1fb5d27fafa8c7726a0cddbe94a1d3f969f24978eb1de7b668e724c53e5171ce

Observation dcbd5eb4-cd14-4cc7-94a3-1a28c43a3bb3 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Reinforcement Learning with Rubric Anchors HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.539628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.539628Z digest=sha256:cb116561dddb6256dc2d74310f26f5e5073e93d0331e28ffc2dca95b9c5cfa45

Observation 41e378f6-6c00-4e67-8c5f-ef872e0ad9d1 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

Reinforcement Learning with Rubric Anchors Instruction-Following Evaluation for Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T17:22:55.544927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:22:55.544927Z digest=sha256:f4d6786cf0db9916b882efe3e1d632a5343bbfaacbe288953991b755daba7f6f

Pith citing papers

Observation b6326af1-e0f9-4191-a4e2-04bd455e6ecc · inbound

Baichuan-M2: Scaling Medical Capability with Large Verifier System cites this paper.

Baichuan-M2: Scaling Medical Capability with Large Verifier System Reinforcement Learning with Rubric Anchors

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T11:50:15.856760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:50:15.856760Z digest=sha256:84391fc16f438104a9a0ed7baa780d871ec360f67f3c05bca851b7d03f4efa2e

Observation 018b4547-fad7-4e0d-8e7b-7e451f2ebff9 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models Reinforcement Learning with Rubric Anchors

Reference 215

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:05:31.565658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:961ebb2a999d5c9d014e33c1cf5ea3eb2272fb8f6084aef9de7229f2ba81f539

Observation 9de7eea5-c94b-4642-b31b-8ed74e67d626 · inbound

DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents cites this paper.

DeepTravel: An End-to-End Agentic Reinforcement Learning Framework for Autonomous Travel Planning Agents Reinforcement Learning with Rubric Anchors

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:49:15.974500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:49:15.974500Z digest=sha256:d33f1b2756bf6147aafbdd56b9fb58d259f05c3537bd0512427940c7eea6c56d

Observation 4746e6f1-f1f0-4529-ab0b-84912315f37c · inbound

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation cites this paper.

RubricRL: Simple Generalizable Rewards for Text-to-Image Generation Reinforcement Learning with Rubric Anchors

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T20:15:28.323243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:15:28.323243Z digest=sha256:47b60d37012e839845c339a0097d9279375f016c2bce6e53b766b5b0a4906479

Observation 6488fe73-7c94-4a9a-8407-0c4765c06b47 · inbound

Verbalizing LLM's Higher-order Uncertainty via Imprecise Probabilities cites this paper.

Verbalizing LLM's Higher-order Uncertainty via Imprecise Probabilities Reinforcement Learning with Rubric Anchors

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T18:31:25.099015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:31:25.099015Z digest=sha256:8547104eee5b2f1c5817a1984ae99e3c4ac0a6d3dc261bf4ea2c971a2cf773d4

Observation e85495b3-c760-4a36-93ab-334b15bc8a38 · inbound

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy cites this paper.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Reinforcement Learning with Rubric Anchors

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.378397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:64ac08c1ab7e244e335e716f45913e91e450a75379fc07d836586d2b4dff99d9

Observation a7e05a27-7df6-4208-9dac-f1a9d3ef5e80 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Reinforcement Learning with Rubric Anchors

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:50.499815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T18:18:19.955943Z digest=sha256:47511d924d9b2739a5bc07657d7989a45a380a7eb6651c66daccae34dbb1d1a1

Observation 65507374-b0d4-4cc0-b6b6-3bfae866d82b · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Reinforcement Learning with Rubric Anchors

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T16:40:55.419909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:40:55.419909Z digest=sha256:84a55382fe1c7f01d9416b0fb3be635e251c865c0ab9bb030d04355c0d94826f

Observation 0674eb1b-f6aa-4b87-ac43-a269823d2265 · inbound

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models cites this paper.

Self-Preference Bias in Rubric-Based Evaluation of Large Language Models Reinforcement Learning with Rubric Anchors

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T05:36:43.322784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:36:43.322784Z digest=sha256:bf997ada37971d8967fdfd828ac44a2d46cdf103beb8c4b222eaf61312d13cab

Observation 38f57d8b-89ed-46d4-ba32-3d9c730b232f · inbound

Visual Preference Optimization with Rubric Rewards cites this paper.

Visual Preference Optimization with Rubric Rewards Reinforcement Learning with Rubric Anchors

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:56:00.726563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:45:52.980881Z digest=sha256:553a0fc59ff4354d06ebf115bb37f0c261964d9337521a09810700911f84265f

Observation 8e8d9f12-3a28-4b71-b19a-11c4ed0d7be1 · inbound

C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences cites this paper.

C2: Scalable Rubric-Augmented Reward Modeling from Binary Preferences Reinforcement Learning with Rubric Anchors

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:40:26.979653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T13:38:37.580448Z digest=sha256:65fd3a35b8b33cfda5c046e542caf3dacd6dedd928521c442d730a3eb5db4744

Observation 16d76370-e6f3-4204-bb10-96dbd490a6b7 · inbound

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence cites this paper.

Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence Reinforcement Learning with Rubric Anchors

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:25:54.256422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T05:24:00.503836Z digest=sha256:1acbd13dc4ff1ef55e9fdf1d712a54b68af33cdfb84b88fe9c62d0eea37de6b2

Observation df2ee1d2-1693-4e1f-85b2-8b40ec090896 · inbound

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text cites this paper.

Bootstrapping Post-training Signals for Open-ended Tasks via Rubric-based Self-play on Pre-training Text Reinforcement Learning with Rubric Anchors

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:21:15.766370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T01:51:00.913166Z digest=sha256:49748bd0779d44e42336c5d1d1400407a064bd18481a5cbfdd03693176429c42

Observation 88c56139-5a71-4347-88e6-e8ab5c921788 · inbound

SCPRM: A Schema-aware Cumulative Process Reward Model for Knowledge Graph Question Answering cites this paper.

SCPRM: A Schema-aware Cumulative Process Reward Model for Knowledge Graph Question Answering Reinforcement Learning with Rubric Anchors

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:55:43.638055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-08T17:57:18.130401Z digest=sha256:6f698f9c0c9e52d4cb976dcd84ebaf4ea8f4c32f6d90737de2d30b1cdac1b419

Observation c3918a16-8ba2-42a6-8562-cd5a1e8e0d23 · inbound

SHARP: A Self-Evolving Human-Auditable Rubric Policy for Financial Trading Agents cites this paper.

SHARP: A Self-Evolving Human-Auditable Rubric Policy for Financial Trading Agents Reinforcement Learning with Rubric Anchors

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:06:04.001409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T00:51:53.252358Z digest=sha256:bae1ed3664422d08a80de265d0ad0da2ec29732ff485cb43b5d82ac265abed48

Observation d3e350a5-deb2-4130-8035-99bf15500bf3 · inbound

Rubric-based On-policy Distillation cites this paper.

Rubric-based On-policy Distillation Reinforcement Learning with Rubric Anchors

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:50:57.003709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T02:11:48.615352Z digest=sha256:a8417835b932fc16566f3c8651c4db67ec2f055cd017472633b2427252186766

Observation 42bde171-b5d2-459e-89e8-bb95268ecd7d · inbound

Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning cites this paper.

Rubric-Grounded RL: Structured Judge Rewards for Generalizable Reasoning Reinforcement Learning with Rubric Anchors

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:58.166891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-11T02:18:33.057569Z digest=sha256:21d172c562407d2ede462e875b1e027fd7b624a4d3d287f91c0c7093ef4d986e

Observation 51564e78-31fb-4dc0-a3a0-559410604284 · inbound

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification cites this paper.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Reinforcement Learning with Rubric Anchors

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.455347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:00e228035cc2ef6daf4388abb930cd15f7fca564cdcc3a13eb0ead4d1e4781f4

Observation 9d8d874a-d852-4e7a-ae23-d1f197eeb25b · inbound

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics cites this paper.

CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics Reinforcement Learning with Rubric Anchors

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:11:23.637013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-12T04:32:16.930291Z digest=sha256:25d46fa519b62d286556c677dea4d4be0f80b6ba8518b3a054bcab46c1225793

Observation 42c593ff-22f8-4dfc-8afb-e8052569f067 · inbound

Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants cites this paper.

Quantifying the Utility of User Simulators for Building Collaborative LLM Assistants Reinforcement Learning with Rubric Anchors

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:36:47.868358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T02:28:13.317630Z digest=sha256:19c78ecc510a25de93c407435715cb27c9d258654db5ddd5298c3dfb334ab460

Observation 81591429-6b5e-4b23-85e2-969e3ab1a6d9 · inbound

Reward Hacking in Rubric-Based Reinforcement Learning cites this paper.

Reward Hacking in Rubric-Based Reinforcement Learning Reinforcement Learning with Rubric Anchors

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:12:14.026353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T04:08:51.578772Z digest=sha256:3f0f03d166d7c664dc7b2f14d9135eb41244a47781a0808dc0409430f32d4de5

Observation 021d35ac-1a5e-407d-9ce2-1bfe14bf133d · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment Reinforcement Learning with Rubric Anchors

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.068131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T12:11:23.775843Z digest=sha256:42adb7d5c8fcfe6404e1d11188c12a2b99824e9c8a10ce59bb3cc4755966f1a2

Observation 48721ea1-1005-40b4-91ed-325c4ec024ad · inbound

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment cites this paper.

AutoRubric-T2I: Robust Rule-Based Reward Model for Text-to-Image Alignment Reinforcement Learning with Rubric Anchors

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:21:21.490340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T09:19:39.848194Z digest=sha256:036e7202276cd1aa07681c64a608bf11db922ed0a377f77760310cf5877609dd

Observation a5ff79c8-0193-4555-926d-0264d28a58ea · inbound

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models cites this paper.

Lens: Rethinking Training Efficiency for Foundational Text-to-Image Models Reinforcement Learning with Rubric Anchors

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:34:46.719351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T09:34:14.596976Z digest=sha256:84af908d3666e519ec8b367590431339171cf2ead77fffa70c4281b4d40b5eef

Observation 0cd010db-204d-4801-9824-627fd227dd98 · inbound

Prompt-Level Reward Specifications for Open-Ended Post-Training cites this paper.

Prompt-Level Reward Specifications for Open-Ended Post-Training Reinforcement Learning with Rubric Anchors

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T08:13:15.885637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T08:04:49.358966Z digest=sha256:2d67d8e335b069c6a854022c8513cd71c70bd4068294a1f227fbd748b9cf1033

Observation c8279687-10cd-4ea4-98f6-a15847dc1e38 · inbound

Reinforcement Learning with Robust Rubric Rewards cites this paper.

Reinforcement Learning with Robust Rubric Rewards Reinforcement Learning with Rubric Anchors

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:43:13.790780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T07:39:21.677389Z digest=sha256:aa94303abb5c428bc1748282dedddc0480572cd2c5d1ac66bde8eeba5619156e

Observation 3941bab8-d0b9-40f1-bfed-1402aadd0e4f · inbound

Deep Research as Rubric for Reinforcement Learning cites this paper.

Deep Research as Rubric for Reinforcement Learning Reinforcement Learning with Rubric Anchors

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-28T17:22:25.224950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T17:16:16.243967Z digest=sha256:f4f3d45bccfd9cc4d5bc1ab85e686937f99b3257a612f84a41d93065a5525363

Observation 96b9cdbc-9019-4bc3-9646-98fd34835cc5 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Reinforcement Learning with Rubric Anchors

Reference 184

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.632639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:fb0c26d2416e7a0f580cf7039930667db7d5f461b5e5a9a30f0da3376c9f51ba

Observation 07f74a91-5a3c-4680-8b02-058efe9fb726 · inbound

BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents cites this paper.

BigFinanceBench: A Workflow-Grounded Benchmark for Financial-Research Agents Reinforcement Learning with Rubric Anchors

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:56:34.686496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T09:35:28.694416Z digest=sha256:071ced1a03a8a6e5020bdc586804b103eb93921806017f60f366ac08803372f4

Observation 4f8e58e2-58a4-4a08-950f-7dca9e664cd2 · inbound

QUBRIC: Co-Designing Queries and Rubrics for RL Beyond Verifiable Rewards cites this paper.

QUBRIC: Co-Designing Queries and Rubrics for RL Beyond Verifiable Rewards Reinforcement Learning with Rubric Anchors

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:56:29.409752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T10:28:11.576647Z digest=sha256:de2827c791a5a6be8c485c82c4c25fe877a0055c983cf32ef0fa23deb0e8bdf5

Observation dfd3c8c5-b0dc-411a-a051-e7bb251a8bea · inbound

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety cites this paper.

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety Reinforcement Learning with Rubric Anchors

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:16:26.549345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T11:07:47.814115Z digest=sha256:ddcd6b9985a30a258c8440789addc057030192e6a976c5b1b6b78d7b7b98586a

Observation 5e4e8840-30f9-4b58-a5c2-a0bbfe3639b0 · inbound

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning cites this paper.

Reproducing, Analyzing, and Detecting Reward Hacking in Rubric-Based Reinforcement Learning Reinforcement Learning with Rubric Anchors

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T06:16:43.765955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T07:28:09.142557Z digest=sha256:01184e9b0288ae645479087c3c3b3d3489af44410a4342201558e2cf124a5856

Observation 20f00e88-edfb-4634-be64-74dcfead656b · inbound

SAW: Stage-Aware Dynamic Weighting for Multi-Objective Reinforcement Learning in Large Language Models cites this paper.

SAW: Stage-Aware Dynamic Weighting for Multi-Objective Reinforcement Learning in Large Language Models Reinforcement Learning with Rubric Anchors

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:27:09.446012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T22:39:36.238242Z digest=sha256:7d3fe1a74a2836bae57fe1f1f157c067113a13c535b60fe1ea9b4549c39cf082

Observation 13feac50-cbab-421b-9a7d-61a4161e476b · inbound

Baichuan-M4: A Clinical-Grade Medical Agent System for Continuous Care cites this paper.

Baichuan-M4: A Clinical-Grade Medical Agent System for Continuous Care Reinforcement Learning with Rubric Anchors

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-27T17:01:07.401905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T16:58:54.293859Z digest=sha256:4feba91f1db0141b1232d1bec323b4813b0ebdd2f8839dcb33d8ef18aac5c772

Observation e63d44ae-0109-4635-a1ee-d950a41873e6 · inbound

ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents cites this paper.

ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents Reinforcement Learning with Rubric Anchors

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:19:38.643902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T14:33:50.123077Z digest=sha256:fb1574dfa91d4c5d22ca8373ab6d84d1616a9daa51c8a8ab9ac29d24ff4cf88d

Observation 1fdf1afe-73ec-4da6-a950-c733a942c6f2 · inbound

ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents cites this paper.

ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents Reinforcement Learning with Rubric Anchors

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T10:43:53.683485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:43:53.683485Z digest=sha256:21ad05d55e292cc86997039e1115ee9dd3171e48fc1410006371a1d6df974bf5

Observation 985a6654-89ee-4175-a9ba-f9daad0930ae · inbound

Teach-to-Reason: Competition-Guided Reasoning with a Self-Improving Teacher cites this paper.

Teach-to-Reason: Competition-Guided Reasoning with a Self-Improving Teacher Reinforcement Learning with Rubric Anchors

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:50:10.678512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-25T20:59:22.132943Z digest=sha256:81d4624633be1c27562ff5424f5e0a868201b225e19bbe8e058fbca7374de729

Observation 5ef81299-f809-4b3c-8eaf-dc99e5af92e2 · inbound

LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation cites this paper.

LEAP: Lean Environment-Feedback via Adaptive Pruning for Code RL in GPU Kernel Generation Reinforcement Learning with Rubric Anchors

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:38.445310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:38.445310Z digest=sha256:e7f49c958650356cc4b0de699b2a3f21ba99dce3eb0159e3d0d8c6205040d7da

Observation 4831f5a2-441e-4dab-8509-26f24275999b · inbound

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL cites this paper.

Rubric Dropout: A Simple Way to Mitigate Reward Hacking in Rubric-as-Reward RL Reinforcement Learning with Rubric Anchors

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:05.606844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:05.606844Z digest=sha256:03eda7df4d45ca1b5a9a700607db74250b243ced907fc036bb194bfbe9ec5c54