Pith. sign in

Paper Citation Record · LEDGER

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification

As of 4 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2605.09269.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.09269 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T04:41:44.833354Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact32
  • verified fuzzy21
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5879cfb6-b351-4b6a-a0d1-f3a6e905a593 · outbound

This paper cites Qwen3-VL Technical Report.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:24.295028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:3bcca07ab6d96b1506e71a7093b5b696c0f47c6ba13654efcd8f09fefa071435

Observation 0c62d594-07ab-4aa6-b87a-64e137cbb3ef · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:24.603346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:b0717b613ab552895e59fdc9487cca54c0d8fdb20f303fd563b6f072928ccf77

Observation beff6b32-0c4e-4e49-913d-7aa713cb5529 · outbound

This paper cites Mllm-as-a-judge: Assessing multimodal llm-as- a-judge with vision-language benchmark.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Mllm-as-a-judge: Assessing multimodal llm-as- a-judge with vision-language benchmark

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.703241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:e61a1897901cf990908291cd32c0fd36bf3d0da1c11bbdca1aeec9b740da9ff7

Observation 721fd5d2-f459-4a16-bb02-45c142a95eb7 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.684737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:51b10c7503b7508df9a873355d89fdf4c0330040274e0191ebf4fe19de604766

Observation 9d32ec0f-a3ba-43e5-bbd1-5de9823bc9c3 · outbound

This paper cites CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification CDE: Curiosity-Driven Exploration for Efficient Reinforcement Learning in Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.281809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:fdc7c5e8a79494d70116a4434de87038f63494d6c02e6ebab84bdd014d4f8d66

Observation e76484cd-742b-4f30-89b0-7804273f59ea · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification NVLM: Open Frontier-Class Multimodal LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.444129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:96e9eaff826edb98f327a60c8afc7ccff76921b3d4b81a1480a69d1e60caa259

Observation 63e8f2bc-c183-470b-a1c6-d7a81ff910b5 · outbound

This paper cites Deepseek-v4: Towards highly efficient million-token context intelligence.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Deepseek-v4: Towards highly efficient million-token context intelligence

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.698451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:f00043e8510d7d7afb03a551161574d83e5899b5df787a0184f99d336be860d9

Observation 4a7e8e15-c74a-41b7-9da5-85144ed0bd88 · outbound

This paper cites Molmo and pixmo: Open weights and open data for state-of-the-art vision-language models.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Molmo and pixmo: Open weights and open data for state-of-the-art vision-language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.692834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:bc5c4bcaccc4c75288dabdbf22707c0b8e3656bf561337caf9eebbe7ba61743d

Observation e557804f-5c5e-4380-88d4-25009bd235b7 · outbound

This paper cites Arm-thinker: Reinforcing multimodal generative reward models with agentic tool use and visual reasoning.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Arm-thinker: Reinforcing multimodal generative reward models with agentic tool use and visual reasoning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.429052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:999aff7564ad7ebef4adfd493d85833dbbe33406fa7159ef0db84fccbc94d3bb

Observation 0c44c035-3c16-4f7a-b0ea-0d993eedbdd5 · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.955297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:884e6452b782e08b885387936dbb347a8f850fede89880d8a889d95c1b1ea4a5

Observation cf011423-04e0-46e1-98c8-978acf86f2a8 · outbound

This paper cites The Llama 3 Herd of Models.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification The Llama 3 Herd of Models

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:24.506738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:7d54b8417ee40ec36d92d32313d4f02c2fe7e02fc65601df3f574a5a244c02c7

Observation b3e0c731-125b-405c-851e-0013630e2485 · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:56.884714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:2ac912d17b78f2daea0c761ab4dbfefff39f615f1ad92db01dd8b62813cd8564

Observation d5b69159-a3cd-4b4d-9d6c-41c5f44e20e6 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:24.470450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:9290f7ddd663512338964f6477f478f4113bb0b81e654299d20cbf084a26ba3f

Observation 523982d0-2968-4c72-8130-c3fd81a068b7 · outbound

This paper cites Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.734575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:3a78e577edd9c3dab6e622e0318fd7f78bd4fd635f86d3bb93cdda51d7a393c0

Observation 51564e78-31fb-4dc0-a3a0-559410604284 · outbound

This paper cites Reinforcement Learning with Rubric Anchors.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Reinforcement Learning with Rubric Anchors

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.455347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:5dbacbf33a27d7461c794cf077d73450ea560165b36075bcd84c961dc5e1f582

Observation 19145b6f-a13f-44a0-93dc-15ecfaa5f144 · outbound

This paper cites AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification AutoRubric: Rubric-Based Generative Rewards for Faithful Multimodal Reasoning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:24.590349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:16eecc739748e72f829ecb0eb2545eeba562cf58b6c5a6f811fc5b045b43ced3

Observation b220222e-417c-4acb-87c9-4d0357f2f929 · outbound

This paper cites Prometheus: Inducing fine-grained evaluation capability in language models.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Prometheus: Inducing fine-grained evaluation capability in language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.756042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:bd835c3ed60f721df132533bcda2193a071ef2c4011d73c97ffa514f46264bf5

Observation fd478e17-5cc3-4e11-a8ca-58fdda61b457 · outbound

This paper cites Reinforcement Learning from Human Feedback.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Reinforcement Learning from Human Feedback

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:24.420238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:640ba8d4eb05796fed33436af760b314886b4f4d6d95611852350cfc5c40d656

Observation f597b9fb-6654-4b1b-8dec-2bf762075d70 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:24.260624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:1341c682d650ce757f73433aaaec817085fdbadeb9a2b03b49a60578848c25ac

Observation aeddf4f5-023a-4ad1-84c9-aed1bc01d76e · outbound

This paper cites Rewardbench: Evaluating reward models for language modeling.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Rewardbench: Evaluating reward models for language modeling

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.769381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:44acf1105a5fc3c8aca419a969b70b563885424bd3a49117801f41f269d36a20

Observation c46d7b98-661c-4377-80a2-e04da4941c49 · outbound

This paper cites Vl-rewardbench: a challenging benchmark for vision-language generative reward models.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Vl-rewardbench: a challenging benchmark for vision-language generative reward models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.751702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:c9b93050a79aae00360b06fb8420044e816cead22f5fd910af1c7ede3343b576

Observation ca736119-30ee-4cae-9c82-f376cb704189 · outbound

This paper cites Self-Rewarding Vision-Language Model via Reasoning Decomposition.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Self-Rewarding Vision-Language Model via Reasoning Decomposition

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:24.251497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:9e3c5e2908527e378e5aac234b1cf0b116eefc5983b67496d248893b5cef2cbd

Observation 725f7bef-73d7-4d9a-b52f-75d9f9a78d84 · outbound

This paper cites Stable and efficient single-rollout rl for multimodal reasoning.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Stable and efficient single-rollout rl for multimodal reasoning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.216769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:33cbdb342e90048e1432732fa32630e8613f216be4ded7a37ee11fc562ee4c45

Observation 32e8f812-a10f-45f5-9af5-50317690d48e · outbound

This paper cites V ogue: Guiding exploration with visual uncertainty improves multimodal reasoning.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification V ogue: Guiding exploration with visual uncertainty improves multimodal reasoning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.388524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:f09b9fdabb33b01e992674f721a236969342dbf82a035e48b56455f1e05491ec

Observation aef104c3-1531-41d6-8438-dfae74c1807b · outbound

This paper cites arXiv preprint arXiv:2510.07743 , year=.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification arXiv preprint arXiv:2510.07743 , year=

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.242097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:a66a4abccc6016fe14d8701c760d4eab3174facf246b618bf8c33a799d2707b4

Observation b6efb9dc-d12f-432a-9d13-b20f05622242 · outbound

This paper cites Decoupled weight decay regularization.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Decoupled weight decay regularization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.765087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:5e9ac61f016dc2de848d870414b1530b359568c8ecf3f8b9fb617ba1d7ddcba8

Observation 2603bb94-aa15-46d5-abb0-9e2a6b308547 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.743511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:fe8033e283eb7511be2b5113db9b64ca2391dca8c2b4aee52968eaa60e5181db

Observation 21ac7b83-9bd5-4d5b-8aef-b8f8d329fa74 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Direct preference optimization: Your language model is secretly a reward model

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.717473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:3f9c424979866f7fa3228b08c6cbe88dab3a833885ee84a58e1247bc0ed40548

Observation ef7f680b-328d-47bf-a816-051e1af342cc · outbound

This paper cites Self-critiquing models for assisting human evaluators.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Self-critiquing models for assisting human evaluators

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:25:41.981797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:b004b518ecd55663363f1c3d67ed33f239a2f7fb9b568b0fb789933d103872dc

Observation e5ff6404-54a1-40ab-b09e-c11afc607563 · outbound

This paper cites DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification DR Tulu: Reinforcement Learning with Evolving Rubrics for Deep Research

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:00:28.487879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:91b65a9e11589c0e851be3a6477502f1682ada9cf88c063318e0b9150a51ef15

Observation d852f360-6bb7-4d8f-a8a5-d40d39f39964 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:24.366769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:101bee91512fad3a9aa1d37745cfc3458304e1f35350b51715d513bd539493a8

Observation 932d95d9-3453-4e4a-85e8-5af2d9cf89d8 · outbound

This paper cites arXiv preprint arXiv:2602.10885 , year=.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification arXiv preprint arXiv:2602.10885 , year=

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.271966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:bddd51e6139e146ab3f9dad184c99970b0ffdc6e411e9f645541a6950ede707e

Observation da880bd9-6413-4078-b427-eb13109afe13 · outbound

This paper cites A Long Way to Go: Investigating Length Correlations in RLHF.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification A Long Way to Go: Investigating Length Correlations in RLHF

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.374651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:211f7967f314fafe242ca02c0892212cc8cff7360a40d05c72b366908bc21ef1

Observation ff946a35-3462-4f25-afcb-daa9f650cf43 · outbound

This paper cites Aligning large multimodal models with factually augmented rlhf.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Aligning large multimodal models with factually augmented rlhf

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.707894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:6218f37164b296a708062fdcd6b5e5c83110b75fc06a7512f3b4b56da57ffddc

Observation 3517bce8-c0c5-4222-a544-0919d999f018 · outbound

This paper cites ReFT: Reasoning with reinforced fine-tuning.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification ReFT: Reasoning with reinforced fine-tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.712808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:13937cc0b218cdb48dc7973215e37742acf85cc8217aafc044d4b5d121b57820

Observation 74e056fc-f5d1-4b6a-9dc0-272562497052 · outbound

This paper cites R e FT : Reasoning with reinforced fine-tuning.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification R e FT : Reasoning with reinforced fine-tuning

Reference 36

Resolution
metadata mismatch
doi, observed 2026-05-12T05:16:23.874409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:392e57ed4bc621560d5ea19c62aa329ce0064950d52071f42da0102ddd598681

Observation 813ab63e-583d-4ddf-9d75-735ff016222d · outbound

This paper cites Chenglong Wang, Yang Gan, Yifu Huo, Yongyu Mu, Qiaozhi He, Murun Yang, Bei Li, Tong Xiao, Chunliang Zhang, Tongran Liu, and Jingbo Zhu.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Chenglong Wang, Yang Gan, Yifu Huo, Yongyu Mu, Qiaozhi He, Murun Yang, Bei Li, Tong Xiao, Chunliang Zhang, Tongran Liu, and Jingbo Zhu

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:01:24.401201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:904279925af031259ac7acd611d145f5ef9d6190a3b5a7c7eabdf34a4b528e57

Observation 12b6094e-4cc1-4e5a-ae09-2b5122137a63 · outbound

This paper cites Msrl: Scaling generative multimodal reward modeling via multi-stage reinforcement learning.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Msrl: Scaling generative multimodal reward modeling via multi-stage reinforcement learning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.484234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:21fafc29072a2c0c970f79fc0140fc1af162359f63c6ce755ab5113e7243edf0

Observation af2cc219-e8a8-4cc8-9b79-1180bdc3fa53 · outbound

This paper cites Unified multimodal chain-of-thought reward model through reinforcement fine-tuning.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Unified multimodal chain-of-thought reward model through reinforcement fine-tuning

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.410905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:4dfa45b02e797ce448303544dcfad0752d4436c68de91b111570e18b86d8fdd2

Observation 4d4cfac3-a999-46fb-9a3f-a68183138aa6 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Chain-of-thought prompting elicits reasoning in large language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.760409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:da491711b8c6c896134f159274363f9228fcc481fe4a5c46bd3eac471fa97966

Observation 48325d23-b0ba-4b51-82cb-733e2e0cf397 · outbound

This paper cites Llava-critic: Learning to evaluate multimodal models.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Llava-critic: Learning to evaluate multimodal models

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.747646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:36cbb3e6148946fb08e2c28023f6d66bc3a1d7006ef3c3b4011bd36217437ee3

Observation 880f99d9-7df2-4c46-89cd-d2a57caebb1a · outbound

This paper cites arXiv preprint arXiv:2602.01511 , year=.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification arXiv preprint arXiv:2602.01511 , year=

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.573347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:26f386e60c5bac5d1113c1b96add700c2b15ed395d56f522709d6b5c24b0c50f

Observation 5a2bc0ff-d920-4e03-8077-660d03fc6295 · outbound

This paper cites Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Multimodal RewardBench: Holistic Evaluation of Reward Models for Vision Language Models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.318015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:4bc9dde095814e63f6a4c536f29bfcee5b0281c5aa1d453630c133e9cef14cf2

Observation 90fb2227-de76-4f6c-9864-882575feae60 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:01:24.303618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:d07e7b6e3c4693173d3f1a599da58effaa1170b59d99d544901ef4bae13f2008

Observation 0a3fb8f6-bc8a-4139-a563-448799a4da7b · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.730161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:6568910cc6e88182fa78f9d3d96001782db255402476239e713a369f6cb475f4

Observation bc0a6206-20c3-414f-b6fd-b9c53bd5925a · outbound

This paper cites Rlaif-v: Open-source ai feedback leads to super gpt-4v trustworthiness.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Rlaif-v: Open-source ai feedback leads to super gpt-4v trustworthiness

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.721727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:0aa9ea5c06d3481bec8a661027628d421e9900906d9d347ea65ab1b977256184

Observation 26517c6e-f81a-43aa-95a6-ff49946f9574 · outbound

This paper cites Benchmarking Large Multimodal Models against Common Corruptions.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Benchmarking Large Multimodal Models against Common Corruptions

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.267389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:070d58ae529f0d97bca8d6c3c8bc2770c0ed09c64ca6070bb2f880648f13c2c0

Observation a04a95b7-c373-4b04-8db4-ace8ef5aa026 · outbound

This paper cites R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.224742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:5b5dfd7c59906fbd980e5c70505080489968c2b132ae284e93316feb6c5b5468

Observation 6868191b-257a-47b9-9050-a49334d09a96 · outbound

This paper cites MM-RLHF: The Next Step Forward in Multimodal LLM Alignment.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.331700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:107d6e43d0c17c8f883d8f3c115e8abfd3759decbe0555c55bd41bd8302a4d9a

Observation 787fec71-615b-40f6-ac3e-5b0818cb4e29 · outbound

This paper cites Basereward: A strong baseline for multimodal reward model.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Basereward: A strong baseline for multimodal reward model

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.773460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:03e6ed147dab743535f5671cc14bf326ba03b9f5b60806dc4a0ba552e120e197

Observation 2cbe5227-9cba-44b3-a2c2-257944ce121e · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Judging llm-as-a-judge with mt-bench and chatbot arena.Advances in neural information processing systems, 36:46595–46623

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.726088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:3bbf851d5b6afddd7ca18f21f86748006070585f668c845510e9e627f4d52905

Observation cca0ac9b-fc25-41ba-a2b9-ef19002ac348 · outbound

This paper cites Parallel-R1: Towards Parallel Thinking via Reinforcement Learning.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.337093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:cefc9809a6172a54dd9d482c830f8c9a44ffb512f26f7a436d9aea57fa822786

Observation 1e6090ce-cd54-4a0b-a880-4601e42475f1 · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Easyr1: An efficient, scalable, multi-modality rl training framework

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:56:36.738458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:e5f48cff855ba8ec57d382a73aa26b11438e6fff1cb3c0d4453ab654b667f1a7

Observation 35a22826-9bed-4c54-ad0b-4b26450c99e0 · outbound

This paper cites A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification A Comprehensive Survey of Reward Models: Taxonomy, Applications, Challenges, and Future

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.230952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:5b64e81e87c8e3ba1bf6ca85a4c0003c851c46e71c20170f66070c6d27710149

Observation ec9d4ab1-8589-4aa9-b5f2-0a0bbd25eeec · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 55

Resolution
malformed identifier
local_arxiv, observed 2026-05-12T06:01:24.379421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:57ee01870f37a4044d4e0c8dfeafaf027793531ddd559b7ff28db3cb4361cf8c

Observation d5b12c00-4a40-4616-a19e-63b611fcaced · outbound

This paper cites an unresolved cited work.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-12T13:51:38.004569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:175d02614a2b35576dbf29b474e3528cde0a89eaee4a657ddc8116a07289a868

Observation e78a21ec-1ec1-43fb-90d8-845a7129140c · outbound

This paper cites an unresolved cited work.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-12T13:51:37.989158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:16986c4610361739760a38e35df8a38bb0ddcd5ee44c8f9cabe2bd08495dc128

Observation c6a62fd2-3ccb-4bd2-89ba-9697b5d693f2 · outbound

This paper cites an unresolved cited work.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-05-12T13:51:37.994662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:6efe1b6a09456d903e809d2b5c8e8412b688571e8f2ecad684adf0355bbe3b5b

Observation 2e9dfcfa-35b5-45ec-b8c7-8c690445f80e · outbound

This paper cites an unresolved cited work.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-12T13:51:38.000376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:aa2576ac1525d5cf318ff15719433162426c55af770a17a8f2ca81375cce4cc0

Observation 1dab923f-89f5-464c-a8db-e70758b72d13 · outbound

This paper cites DeltaRubric Evaluation Prompt You are a fair judge.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification DeltaRubric Evaluation Prompt You are a fair judge

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T13:51:37.984268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:b2520946f534dbb3b4116f04e1b2890ac99ddabea240bde4d4b6024069bb951a

Pith citing papers

No inbound Pith citation observations are available.