Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T17:08:23.269275Z
Paper Citation Record · LEDGER
As of 3 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 2 inbound Pith citation observations for arXiv:2603.15646.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T17:08:23.269275Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T15:03:42.266254Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-12T07:51:38.881425Z
28 of 28 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f7ed8e7f-0c5a-4c00-9a81-8acffdac8a78 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy H., Gendler, A., Baruch, E
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation bf8328d8-4c95-4557-a204-05683fd2b2ec · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy HealthBench: Evaluating Large Language Models Towards Improved Human Health
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 02e6903c-9be5-42d8-bf6e-aef97cd6cd38 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Constitutional AI: Harmlessness from AI Feedback
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1bacc138-fa68-4527-b4fe-2a77ad2b0d19 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy XRPO: Pushing the limits of GRPO with Targeted Exploration and Exploitation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 67c5f086-2392-481f-9936-c73202b9170d · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Language models that think, chat better
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 5d59b23f-9125-4152-91fd-b49ca58de998 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation babd34df-a8b7-480c-8daf-1b43780aa7f0 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy arXiv preprint arXiv:2511.10507 , year=
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation e85495b3-c760-4a36-93ab-334b15bc8a38 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Reinforcement Learning with Rubric Anchors
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 926a2f97-4b4c-4901-b841-6f6c88568dad · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy MaPPO: Maximum a Posteriori Preference Optimization with Prior Knowledge
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 743c52f8-3add-4133-91d6-b5d6be436fbe · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Gradient-Adaptive Policy Optimization: Towards Multi-Objective Alignment of Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b8c6df27-566c-470d-9452-5dd60f4f12a1 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Learning to optimize multi-objective alignment through dynamic reward weighting
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 64ad7a06-22b8-4fe7-8cd6-be8316a5cda4 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Online rubrics elicitation from pairwise comparisons
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 230ddb86-54b7-48d0-85af-f7e6af4abd7e · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Proximal Policy Optimization Algorithms
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 4d72c8f2-1d8b-486d-aed7-81956b721a65 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation b353b112-8ab4-4c92-b29c-0352a7f7fd9f · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 98fc3c8c-3610-4314-9f0a-407b49f7958d · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy HybridFlow: A Flexible and Efficient RLHF Framework
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 1544c5fe-57fb-47d7-8d66-bcc4ca1ba7ba · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Nemotron-cascade: Scaling cascaded reinforcement learning for general-purpose reasoning models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 58184558-409e-4387-8175-a31b123cae3c · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy arXiv preprint arXiv:2602.01511 , year=
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 03f65816-e7db-4a58-b780-04d4effd86f9 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Qwen3 Technical Report
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation df6a2493-f2bf-462a-84cb-ac6ed9e751ae · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Does this image satisfy this rule?
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation a8683238-ad7e-4b0a-aac5-b3bfd90353d4 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Group Sequence Policy Optimization
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation f49e56ad-0276-47f8-ac17-be9fbaa1550f · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Breaking the exploration bottleneck: Rubric-scaffolded reinforcement learning for general llm reasoning.arXiv preprint arXiv:2508.16949
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation d3a505f3-a108-4713-8047-b886cb783476 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy I’m an emergency medicine physician
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 86af8d85-5403-4cc1-8bdd-d5597f3e90b2 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy accuracy
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation c880ebe9-ddf9-44eb-a386-db0d3aaeffc7 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy ‘list [ { “criterion
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 858c46ae-06b5-4637-ae2b-1cad7bfae204 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy Llama-3.1-8B-Instruct starts at 0.34 and achieves 0.70, while Qwen3-8B starts at a higher score0.58and achieves0.76
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 58b581e5-cd56-4160-a5e6-b9b8e55bdbfe · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy In ARL, we use the fixed meta-class Order 0: [completeness, accuracy, instruction following, context awareness, communication quality]
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 28e83434-7346-41a8-b8c9-87c6ed66c940 · outbound
Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy The maximum response length is set to 2048, and the temperature in LLM sampling is set to 1.0 in the training process
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 01682382-2014-48d5-872e-b6959ccacae6 · inbound
Reinforcement Learning for Scalable and Trustworthy Intelligent Systems Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
Reference 164
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.
Observation 623215f7-1583-430c-8a30-443982b723de · inbound
Toward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions Alternating Reinforcement Learning with Contextual Rubric Rewards: Beyond the Scalarization Strategy
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.