Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T17:16:16.243967Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2606.01091.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-28T17:16:16.243967Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T15:12:55.560232Z
A source-named dated measurement, never combined with another source.
Source: cited_works
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b27f6f4b-0a86-4554-8f44-e5054f702b34 · outbound
Deep Research as Rubric for Reinforcement Learning Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2b1eccd0-265f-4be2-9dc1-07a55dba0580 · outbound
Deep Research as Rubric for Reinforcement Learning Does this image satisfy this rule?
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b05954bf-1324-41ad-89f0-0fc2b5e37e35 · outbound
Deep Research as Rubric for Reinforcement Learning Breaking the Exploration Bottleneck: Rubric-Scaffolded Reinforcement Learning for General LLM Reasoning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 746a89f2-8575-410a-b460-9f61e95aca4d · outbound
Deep Research as Rubric for Reinforcement Learning HealthBench: Evaluating Large Language Models Towards Improved Human Health
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 62da4def-c4b9-4d5f-b6eb-a2f5511e4da1 · outbound
Deep Research as Rubric for Reinforcement Learning arXiv preprint arXiv:2510.07743 , year=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5907b663-2a4f-4d98-acb8-805fa0ad37cc · outbound
Deep Research as Rubric for Reinforcement Learning Auto-rubric: Learning from implicit weights to explicit rubrics for reward modeling
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3941bab8-d0b9-40f1-bfed-1402aadd0e4f · outbound
Deep Research as Rubric for Reinforcement Learning Reinforcement Learning with Rubric Anchors
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bc13d997-a825-491b-865f-883841171186 · outbound
Deep Research as Rubric for Reinforcement Learning Ace-rl: Adaptive constraint-enhanced reward for long-form gen- eration reinforcement learning.arXiv:2509.04903
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 64f4f153-758b-4776-99f7-0dcf40ba721e · outbound
Deep Research as Rubric for Reinforcement Learning DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 232c5f56-bb53-4007-83d2-b362eaf15298 · outbound
Deep Research as Rubric for Reinforcement Learning Online rubrics elicitation from pairwise comparisons
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4117f923-51b9-4dc9-bbe8-0eb643545e28 · outbound
Deep Research as Rubric for Reinforcement Learning Constitutional AI: Harmlessness from AI Feedback
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6e8cb5c8-7a52-4cff-aab2-e8887021991f · outbound
Deep Research as Rubric for Reinforcement Learning Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3f85ff3b-6de0-4164-9839-a75efce56570 · outbound
Deep Research as Rubric for Reinforcement Learning Self-Rewarding Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cbb69757-6b27-44c5-853a-7470c156e2be · outbound
Deep Research as Rubric for Reinforcement Learning Yifei, Allen Chang, Chaitanya Malaviya, and Mark Yatskar
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5155c358-dafc-4720-a626-f8af2e1cefc5 · outbound
Deep Research as Rubric for Reinforcement Learning DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 848a95d9-be29-41db-bd86-63b17b26cdf8 · outbound
Deep Research as Rubric for Reinforcement Learning LocalSearchBench: Benchmarking Agentic Search in Real-World Local Life Services
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d87bd96b-f254-4167-ab32-88a1a02e5d6f · outbound
Deep Research as Rubric for Reinforcement Learning GPQA: A Graduate-Level Google-Proof Q&A Benchmark
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3cc9edea-ea60-4c30-9f8a-3b0291df48e2 · outbound
Deep Research as Rubric for Reinforcement Learning Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23dc7db0-17e9-4eb9-bf0a-36c7003507ba · outbound
Deep Research as Rubric for Reinforcement Learning Measuring massive multitask language understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bba97290-5146-478a-90af-d15bfaaab146 · outbound
Deep Research as Rubric for Reinforcement Learning Qwen3 Technical Report
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6687768f-6bdd-4fdf-8518-d9f7a85e1949 · outbound
Deep Research as Rubric for Reinforcement Learning Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9821d9a6-fad3-4048-a98a-743e8d7f257d · outbound
Deep Research as Rubric for Reinforcement Learning WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 88a9dbb6-e9b6-4cbe-a35a-8e1aaa2ab213 · outbound
Deep Research as Rubric for Reinforcement Learning OpenAI GPT-5 System Card
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6c6e216f-1c95-453f-aa55-18750f45483a · outbound
Deep Research as Rubric for Reinforcement Learning Gemini 3.1 pro model card
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54014562-b249-4d55-8e3b-2a83a15dda66 · outbound
Deep Research as Rubric for Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8cec7187-514e-4c1f-bddf-fea4ebe3799e · outbound
Deep Research as Rubric for Reinforcement Learning WebThinker: Empowering Large Reasoning Models with Deep Research Capability
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 44e285c9-dbe7-449c-b9e3-40da58f07dd3 · outbound
Deep Research as Rubric for Reinforcement Learning Tongyi DeepResearch Technical Report
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7dbca3a0-339b-4850-b327-830d3a2e4d52 · outbound
Deep Research as Rubric for Reinforcement Learning DeerFlow: Deep exploration and efficient research flow
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 738da4a9-0a84-4420-b07f-6d75c523112d · outbound
Deep Research as Rubric for Reinforcement Learning Narasimhan, and Yuan Cao
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8012d27c-7e90-4291-b2d3-ccf9558b6322 · outbound
Deep Research as Rubric for Reinforcement Learning Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9517785-0bf6-48e1-9394-07be6455558d · outbound
Deep Research as Rubric for Reinforcement Learning Ministral 3
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 06387279-218c-4fce-9484-efc3897278f5 · outbound
Deep Research as Rubric for Reinforcement Learning Mirothinker-1.7 & h1: Towards heavy-duty research agents via verification
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 39935e55-a2a6-41ca-831f-6b4beb816b50 · outbound
Deep Research as Rubric for Reinforcement Learning 15 System Prompt: Stage II (Rubric Synthesis) # Role Definition You are an expert in evaluation framework design for academic research
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c357a92e-209b-4ac7-b756-d09db642d14e · outbound
Deep Research as Rubric for Reinforcement Learning This report contains the necessary factual information (algorithms, parameters, benchmarks, etc.) that a high-quality responseshouldcontain
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eb1a042-58c3-480d-852c-c5df007ad39a · outbound
Deep Research as Rubric for Reinforcement Learning Prohibited Content:
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1d1d5c3-33f1-46d8-9339-2d75ec8ad8c3 · outbound
Deep Research as Rubric for Reinforcement Learning Unresolved cited work
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faeef9c8-f511-4659-940b-59175fcded02 · outbound
Deep Research as Rubric for Reinforcement Learning Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cdd257f-5e00-4301-a97e-7a0fa58140f3 · outbound
Deep Research as Rubric for Reinforcement Learning >=[X] specific algorithms included
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 530da140-3af1-47d5-9af4-bcbe9578a6d2 · outbound
Deep Research as Rubric for Reinforcement Learning #### Core Dimension:
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7642a616-1f32-43a3-b269-9e3947dd9f70 · outbound
Deep Research as Rubric for Reinforcement Learning Stop ifP n >max(0.15,2×P n−1)
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d4faf62-6194-4773-ad17-b23b647fac37 · outbound
Deep Research as Rubric for Reinforcement Learning Unresolved cited work
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a4d200d-0955-4ddd-94f3-dcd863013a83 · outbound
Deep Research as Rubric for Reinforcement Learning Retrospective validation on both scales: this rule correctly selects BS-2 and terminates at BS-3, avoiding BS-4/BS-5 and saving 40–60% of total bootstrap compute
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52111dd3-7d53-42c2-8a8e-8c4deb7fa28d · inbound
Fetch-then-Explore: Decoupling Selection from Extraction over a Persistent Workspace for Search Agents Deep Research as Rubric for Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.