Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T04:41:44.833354Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2605.09269.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T04:41:44.833354Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
60 of 60 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5879cfb6-b351-4b6a-a0d1-f3a6e905a593 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0c62d594-07ab-4aa6-b87a-64e137cbb3ef · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation beff6b32-0c4e-4e49-913d-7aa713cb5529 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Mllm-as-a-judge: Assessing multimodal llm-as- a-judge with vision-language benchmark
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 721fd5d2-f459-4a16-bb02-45c142a95eb7 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9d32ec0f-a3ba-43e5-bbd1-5de9823bc9c3 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e76484cd-742b-4f30-89b0-7804273f59ea · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification NVLM: Open Frontier-Class Multimodal LLMs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 63e8f2bc-c183-470b-a1c6-d7a81ff910b5 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Deepseek-v4: Towards highly efficient million-token context intelligence
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4a7e8e15-c74a-41b7-9da5-85144ed0bd88 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Molmo and pixmo: Open weights and open data for state-of-the-art vision-language models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e557804f-5c5e-4380-88d4-25009bd235b7 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Arm-thinker: Reinforcing multimodal generative reward models with agentic tool use and visual reasoning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0c44c035-3c16-4f7a-b0ea-0d993eedbdd5 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cf011423-04e0-46e1-98c8-978acf86f2a8 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification The Llama 3 Herd of Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b3e0c731-125b-405c-851e-0013630e2485 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d5b69159-a3cd-4b4d-9d6c-41c5f44e20e6 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 523982d0-2968-4c72-8130-c3fd81a068b7 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 51564e78-31fb-4dc0-a3a0-559410604284 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Reinforcement Learning with Rubric Anchors
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 19145b6f-a13f-44a0-93dc-15ecfaa5f144 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b220222e-417c-4acb-87c9-4d0357f2f929 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Prometheus: Inducing fine-grained evaluation capability in language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fd478e17-5cc3-4e11-a8ca-58fdda61b457 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Reinforcement Learning from Human Feedback
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f597b9fb-6654-4b1b-8dec-2bf762075d70 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aeddf4f5-023a-4ad1-84c9-aed1bc01d76e · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Rewardbench: Evaluating reward models for language modeling
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c46d7b98-661c-4377-80a2-e04da4941c49 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Vl-rewardbench: a challenging benchmark for vision-language generative reward models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ca736119-30ee-4cae-9c82-f376cb704189 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Self-Rewarding Vision-Language Model via Reasoning Decomposition
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 725f7bef-73d7-4d9a-b52f-75d9f9a78d84 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Stable and efficient single-rollout rl for multimodal reasoning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 32e8f812-a10f-45f5-9af5-50317690d48e · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification V ogue: Guiding exploration with visual uncertainty improves multimodal reasoning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aef104c3-1531-41d6-8438-dfae74c1807b · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification arXiv preprint arXiv:2510.07743 , year=
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b6efb9dc-d12f-432a-9d13-b20f05622242 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Decoupled weight decay regularization
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2603bb94-aa15-46d5-abb0-9e2a6b308547 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 21ac7b83-9bd5-4d5b-8aef-b8f8d329fa74 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Direct preference optimization: Your language model is secretly a reward model
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ef7f680b-328d-47bf-a816-051e1af342cc · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Self-critiquing models for assisting human evaluators
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e5ff6404-54a1-40ab-b09e-c11afc607563 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d852f360-6bb7-4d8f-a8a5-d40d39f39964 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 932d95d9-3453-4e4a-85e8-5af2d9cf89d8 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification arXiv preprint arXiv:2602.10885 , year=
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation da880bd9-6413-4078-b427-eb13109afe13 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification A Long Way to Go: Investigating Length Correlations in RLHF
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ff946a35-3462-4f25-afcb-daa9f650cf43 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Aligning large multimodal models with factually augmented rlhf
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3517bce8-c0c5-4222-a544-0919d999f018 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification ReFT: Reasoning with reinforced fine-tuning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 74e056fc-f5d1-4b6a-9dc0-272562497052 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification R e FT : Reasoning with reinforced fine-tuning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 813ab63e-583d-4ddf-9d75-735ff016222d · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Chenglong Wang, Yang Gan, Yifu Huo, Yongyu Mu, Qiaozhi He, Murun Yang, Bei Li, Tong Xiao, Chunliang Zhang, Tongran Liu, and Jingbo Zhu
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 12b6094e-4cc1-4e5a-ae09-2b5122137a63 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Msrl: Scaling generative multimodal reward modeling via multi-stage reinforcement learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation af2cc219-e8a8-4cc8-9b79-1180bdc3fa53 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Unified multimodal chain-of-thought reward model through reinforcement fine-tuning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4d4cfac3-a999-46fb-9a3f-a68183138aa6 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Chain-of-thought prompting elicits reasoning in large language models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 48325d23-b0ba-4b51-82cb-733e2e0cf397 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Llava-critic: Learning to evaluate multimodal models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 880f99d9-7df2-4c46-89cd-d2a57caebb1a · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification arXiv preprint arXiv:2602.01511 , year=
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5a2bc0ff-d920-4e03-8077-660d03fc6295 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 90fb2227-de76-4f6c-9864-882575feae60 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0a3fb8f6-bc8a-4139-a563-448799a4da7b · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bc0a6206-20c3-414f-b6fd-b9c53bd5925a · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Rlaif-v: Open-source ai feedback leads to super gpt-4v trustworthiness
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 26517c6e-f81a-43aa-95a6-ff49946f9574 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Benchmarking Large Multimodal Models against Common Corruptions
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a04a95b7-c373-4b04-8db4-ace8ef5aa026 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6868191b-257a-47b9-9050-a49334d09a96 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification MM-RLHF: The Next Step Forward in Multimodal LLM Alignment
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 787fec71-615b-40f6-ac3e-5b0818cb4e29 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Basereward: A strong baseline for multimodal reward model
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2cbe5227-9cba-44b3-a2c2-257944ce121e · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cca0ac9b-fc25-41ba-a2b9-ef19002ac348 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Parallel-R1: Towards Parallel Thinking via Reinforcement Learning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1e6090ce-cd54-4a0b-a880-4601e42475f1 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Easyr1: An efficient, scalable, multi-modality rl training framework
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 35a22826-9bed-4c54-ad0b-4b26450c99e0 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ec9d4ab1-8589-4aa9-b5f2-0a0bbd25eeec · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d5b12c00-4a40-4616-a19e-63b611fcaced · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e78a21ec-1ec1-43fb-90d8-845a7129140c · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c6a62fd2-3ccb-4bd2-89ba-9697b5d693f2 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2e9dfcfa-35b5-45ec-b8c7-8c690445f80e · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Unresolved cited work
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1dab923f-89f5-464c-a8db-e70758b72d13 · outbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification DeltaRubric Evaluation Prompt You are a fair judge
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
No inbound Pith citation observations are available.