Pith. sign in

Paper Citation Record · LEDGER

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy

As of 3 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2603.15646.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.15646 v2

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T17:08:23.269275Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T15:03:42.266254Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-12T07:51:38.881425Z

Reference resolution

28 of 28 outbound references displayed

  • verified exact17
  • verified fuzzy6
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f7ed8e7f-0c5a-4c00-9a81-8acffdac8a78 · outbound

This paper cites H., Gendler, A., Baruch, E.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy H., Gendler, A., Baruch, E

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:10:11.210472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:e8c1a746cc87ac890d2d2c88a3478cb3bb2fa6ca4ad60cee776f312a099d3e0e

Observation bf8328d8-4c95-4557-a204-05683fd2b2ec · outbound

This paper cites HealthBench: Evaluating Large Language Models Towards Improved Human Health.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy HealthBench: Evaluating Large Language Models Towards Improved Human Health

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:10:10.350720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:d4066fd3112fa1f6f15796dc44c3e653df4e71ab59277b37f0dbc2f50fa7636e

Observation 02e6903c-9be5-42d8-bf6e-aef97cd6cd38 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Constitutional AI: Harmlessness from AI Feedback

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:10:10.332916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:a2632750e1195e0cb0b816c0fd107a951da99191710d69ae5ff107ac0eb375d6

Observation 1bacc138-fa68-4527-b4fe-2a77ad2b0d19 · outbound

This paper cites XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-26T03:04:06.160463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:4d754cdc8418acf6e22a5a2cc31b09fc285be65db481090521b757fdd9926780

Observation 67c5f086-2392-481f-9936-c73202b9170d · outbound

This paper cites Language models that think, chat better.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Language models that think, chat better

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.403000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:bd831d35221a12d46c0ca9a208614832eeb0febb7e24d663323bf665f3bc1d55

Observation 5d59b23f-9125-4152-91fd-b49ca58de998 · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:10:10.378588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:c8b1271074446f92aef74f7075f53c6694de90789e477e7f3df6fac89528ccf9

Observation babd34df-a8b7-480c-8daf-1b43780aa7f0 · outbound

This paper cites arXiv preprint arXiv:2511.10507 , year=.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy arXiv preprint arXiv:2511.10507 , year=

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.357083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:b643fa5693e46485b382863eeaf2ff78552fdaa5e3cba938bc67609c8faa28ee

Observation e85495b3-c760-4a36-93ab-334b15bc8a38 · outbound

This paper cites Reinforcement Learning with Rubric Anchors.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Reinforcement Learning with Rubric Anchors

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.378397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:a2d2d9f4e920ebc795616489141d9510a2d1ce8399c280080bac533b1dcd33c5

Observation 926a2f97-4b4c-4901-b841-6f6c88568dad · outbound

This paper cites MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T17:10:10.383743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:09270d0351d88848fdbfebfaf10914a4422d8501cc285e1a6c59e398171c20e6

Observation 743c52f8-3add-4133-91d6-b5d6be436fbe · outbound

This paper cites Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.403325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:c5735a517ee405e517ff7cf17b2e001bb8d4e83627ff40b0e38925435284cbba

Observation b8c6df27-566c-470d-9452-5dd60f4f12a1 · outbound

This paper cites Learning to optimize multi-objective alignment through dynamic reward weighting.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Learning to optimize multi-objective alignment through dynamic reward weighting

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.396976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:2b1a7ee17ec0a5a032a00490f1c4217f9b9b0a892cd2773ed44fdf834d7e55a3

Observation 64ad7a06-22b8-4fe7-8cd6-be8316a5cda4 · outbound

This paper cites Online rubrics elicitation from pairwise comparisons.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Online rubrics elicitation from pairwise comparisons

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.372183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:59a245984837427bb0f311686162ba5861e7245ed91693262ebd0486cf3b89ca

Observation 230ddb86-54b7-48d0-85af-f7e6af4abd7e · outbound

This paper cites Proximal Policy Optimization Algorithms.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Proximal Policy Optimization Algorithms

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:10:10.339593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:a5cbff1913b45e4c12f882d68b365023080b62c545e7ba8fc990ca3dea4d4a18

Observation 4d72c8f2-1d8b-486d-aed7-81956b721a65 · outbound

This paper cites DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T00:00:28.487879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:1d19666c9c26eb07cebcaa507611f9960413e904d9168c42f1a690aab56f890e

Observation b353b112-8ab4-4c92-b29c-0352a7f7fd9f · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:10:10.373798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:fc172a055af5cfa1806980019bfb3fc897c2f381daf250b0e337bb381525122f

Observation 98fc3c8c-3610-4314-9f0a-407b49f7958d · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy HybridFlow: A Flexible and Efficient RLHF Framework

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:10:10.390724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:bfff3eb39dbf667765f9aab68510f43c09621616e6658193dd960f7308e31582

Observation 1544c5fe-57fb-47d7-8d66-bcc4ca1ba7ba · outbound

This paper cites Nemotron-cascade: Scaling cascaded reinforcement learning for general-purpose reasoning models.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Nemotron-cascade: Scaling cascaded reinforcement learning for general-purpose reasoning models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.384722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:f2b9113bf9af06e071fdeea1aeae8481d2d117ce3f97bc1034cbd70427d8b406

Observation 58184558-409e-4387-8175-a31b123cae3c · outbound

This paper cites arXiv preprint arXiv:2602.01511 , year=.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy arXiv preprint arXiv:2602.01511 , year=

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.351260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:b61e0fdf4d5f824104c5a058ca4cdfa8a463f0278c24d3fbee526ad3479072f0

Observation 03f65816-e7db-4a58-b780-04d4effd86f9 · outbound

This paper cites Qwen3 Technical Report.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Qwen3 Technical Report

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:10:10.389425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:0ceaad358dc73a342f707a45dc86f79820b6066dbda0f58bcaec7de9427f69ab

Observation df6a2493-f2bf-462a-84cb-ac6ed9e751ae · outbound

This paper cites Does this image satisfy this rule?.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Does this image satisfy this rule?

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T17:10:10.409449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:4cd187b1ccaf0661402809483eddd15e5bc7f012e94c52a869ff626d34a1e910

Observation a8683238-ad7e-4b0a-aac5-b3bfd90353d4 · outbound

This paper cites Group Sequence Policy Optimization.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Group Sequence Policy Optimization

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-15T17:10:10.360727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:297412daac77aa31dc9d4ad4aac55d21f655ff50833286b6b1a0eee506132d69

Observation f49e56ad-0276-47f8-ac17-be9fbaa1550f · outbound

This paper cites Breaking the exploration bottleneck: Rubric-scaffolded reinforcement learning for general llm reasoning.arXiv preprint arXiv:2508.16949.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Breaking the exploration bottleneck: Rubric-scaffolded reinforcement learning for general llm reasoning.arXiv preprint arXiv:2508.16949

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:10:10.314742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:383c2a9c1f047386d8c64b2fb808046c3f86ea907a3cbd26ffc7e94cdaf56895

Observation d3a505f3-a108-4713-8047-b886cb783476 · outbound

This paper cites I’m an emergency medicine physician.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy I’m an emergency medicine physician

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:10:11.215521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:318ce5dbbdb812954827cbde747712c7d58f84c76b6fe9b05514162200842aca

Observation 86af8d85-5403-4cc1-8bdd-d5597f3e90b2 · outbound

This paper cites accuracy.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy accuracy

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:10:11.209097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:44c82b048abba53a3b220b9605e92a0bce420954caf3fa68a06c2570f9896bd8

Observation c880ebe9-ddf9-44eb-a386-db0d3aaeffc7 · outbound

This paper cites ‘list [ { “criterion.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy ‘list [ { “criterion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:10:11.208091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:a8d095b589bd6c8a3482bf326cdf4a35076b02890b2c5065e36dc698a2bbc6cb

Observation 858c46ae-06b5-4637-ae2b-1cad7bfae204 · outbound

This paper cites Llama-3.1-8B-Instruct starts at 0.34 and achieves 0.70, while Qwen3-8B starts at a higher score0.58and achieves0.76.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Llama-3.1-8B-Instruct starts at 0.34 and achieves 0.70, while Qwen3-8B starts at a higher score0.58and achieves0.76

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:10:11.213055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:fb050edb9e2be600509d82a72306151781735ef16c2d1e3e7538c58dc5635193

Observation 58b581e5-cd56-4160-a5e6-b9b8e55bdbfe · outbound

This paper cites In ARL, we use the fixed meta-class Order 0: [completeness, accuracy, instruction following, context awareness, communication quality].

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy In ARL, we use the fixed meta-class Order 0: [completeness, accuracy, instruction following, context awareness, communication quality]

Reference 27

Resolution
malformed identifier
raw_fallback, observed 2026-05-15T17:10:11.204202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:277b32e50e41f192e59cabc8028a9e839f6b0ddeb375f515346a6576a1197225

Observation 28e83434-7346-41a8-b8c9-87c6ed66c940 · outbound

This paper cites The maximum response length is set to 2048, and the temperature in LLM sampling is set to 1.0 in the training process.

Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy The maximum response length is set to 2048, and the temperature in LLM sampling is set to 1.0 in the training process

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T17:10:11.211669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T17:08:23.269275Z digest=sha256:79261b8ae4ffb3134b4895aeea3388af46d8bb5299475be7bf0098db51448c14

Pith citing papers

Observation 01682382-2014-48d5-872e-b6959ccacae6 · inbound

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems cites this paper.

Reinforcement Learning for Scalable and Trustworthy Intelligent Systems Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy

Reference 164

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:51:38.936160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T01:47:40.772146Z digest=sha256:61089b653b7cc4fe7965e956a3fc02476f33e82504736e63ec01017c7d764ede

Observation 623215f7-1583-430c-8a30-443982b723de · inbound

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions cites this paper.

Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T15:03:42.266254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:03:42.266254Z digest=sha256:21e3cd46cac57f4a4d780e63e78ab645d2db2d375780aef1dfc0ce8b0e47eb58