Pith. sign in

Paper Citation Record · LEDGER

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

As of 21 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 23 inbound Pith citation observations for arXiv:2505.02835.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02835 v2

Coverage vector

measured 78 of 78 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:45:55.083191Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:04.359664Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T23:57:29.062496Z

Reference resolution

78 of 78 outbound references displayed

  • verified exact0
  • verified fuzzy69
  • unresolved8
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb3da6f0-176d-4687-8ed3-ea513e5b55e9 · outbound

This paper cites Pixtral 12b.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Pixtral 12b.arXiv, 2024

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.625774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.588843Z digest=sha256:b923e746de6d014640736c0f45914b4d9028a41a1d51433cb9f8f1d52fe05f3f

Observation 5159fb98-5c63-4e79-b22a-264fd0f65345 · outbound

This paper cites Qwen2.5-vl technical report.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Qwen2.5-vl technical report

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.606516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.596141Z digest=sha256:2219c79040b07e37b2550f59ab26de6330e7e4f41725aaaf7e95cf230759bd07

Observation 5b7fb8ec-7d7d-4531-92af-255041bfeb12 · outbound

This paper cites Mllm-as-a-judge: Assessing multimodal llm-as- a-judge with vision-language benchmark.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Mllm-as-a-judge: Assessing multimodal llm-as- a-judge with vision-language benchmark

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-16T00:45:55.286878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.602057Z digest=sha256:4da6b64e4192ba8ea8cae7374cc3c94d06467eeb3a9620ce4fed250a8e1122fa

Observation d0847718-5782-4db8-81f8-7a6fe733c833 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.arXiv, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.593108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.608238Z digest=sha256:1ed1495143d7048cd9765fb09a86edae69e6fb200df16888383465fa1c123bbd

Observation fb81d020-4d5d-42e8-90a3-e340f87c56a3 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.arXiv, 2023.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.arXiv, 2023

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.579108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.614865Z digest=sha256:003e023f9697d6ece9648a6c60eb569c6e71adca9a4542fed07288051d24abcf

Observation 4b712200-ce9d-4d5c-ab84-7d5071faebce · outbound

This paper cites Sft memorizes, rl generalizes: A comparative study of foundation model post-training.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Sft memorizes, rl generalizes: A comparative study of foundation model post-training.arXiv, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.564393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.620089Z digest=sha256:fe1a048ed51196463367544d51acb284b5c32cb6c5934ee34e1e469fabbb7300

Observation 144d6781-ff14-453b-ae5c-76e5569b233d · outbound

This paper cites Process reinforcement through implicit rewards.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Process reinforcement through implicit rewards.arXiv, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.551719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.625581Z digest=sha256:f31a61bb85593b5716094a0e6c92838133b8482dcbcc50c5f426551f451d4698

Observation 5f7225c0-7e6d-4c92-b7b8-6ddfbaf51743 · outbound

This paper cites Nvlm: Open frontier-class multimodal llms.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Nvlm: Open frontier-class multimodal llms.arXiv, 2024

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.539203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.630612Z digest=sha256:de2f5fc6515ca6f8e665d9ae9c0b1e69a07618ef9c8c43efec7d69f27390bd41

Observation b45b43a4-1f50-4f86-9368-af0f72100062 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv, 2025

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.514826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.635385Z digest=sha256:abe4e0f7496d43224486f8f6272cc3b2738dfebc369a8412e67e7dbd3e009ac7

Observation e3b754c6-3549-447f-9814-4726b341c733 · outbound

This paper cites Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models.arXiv, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.495121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.642279Z digest=sha256:27a8bacab4f338cf9e8d0f65992a506104739ece9dc4e0f2d26854801396a665

Observation 06880698-8e6f-4e08-831d-1c0a50aa2d42 · outbound

This paper cites Video-r1: Reinforcing video reasoning in mllms.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Video-r1: Reinforcing video reasoning in mllms.arXiv, 2025

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.478824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.648565Z digest=sha256:999c4bc4630155918f9374d521c0e07075e3cbd0cf28383c01bd727bd8e244b6

Observation 044552c2-92d9-409b-9fff-637ec2c0412c · outbound

This paper cites Vita: Towards open-source interactive omni multimodal llm.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Vita: Towards open-source interactive omni multimodal llm.arXiv, 2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.463839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.653233Z digest=sha256:8960157a17fe3f287222f832435bce0e8e70bd4bd2442552e3226f3f484c2908

Observation 5e1422f8-75f4-4421-9ba1-bebec8f9aac1 · outbound

This paper cites Vita-1.5: Towards gpt-4o level real-time vision and speech interaction.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Vita-1.5: Towards gpt-4o level real-time vision and speech interaction.arXiv, 2025

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.446471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.659157Z digest=sha256:14fc76f399b709fbc3a846bdb86587e5b6c70ea19c86d124ef2bfc96d0fba854

Observation 94e1e321-9c7f-4ae4-bf68-276b742b1f0c · outbound

This paper cites Mme-survey: A comprehensive survey on evaluation of multimodal llms.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Mme-survey: A comprehensive survey on evaluation of multimodal llms.arXiv, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.433285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.662788Z digest=sha256:68935e14506a3dd17e042be5b1fa51c384750bfecc5c1241b669bf018afdc5b1

Observation f8504d21-e52b-4733-9ac5-ffdf864ceb6c · outbound

This paper cites Trips to the zoo last year.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Trips to the zoo last year

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.420023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.666616Z digest=sha256:79606742dbcf545199196e05978e408187b13ba044648b94888be0892f5e4dac

Observation bf6781c0-2064-4d0d-b120-a3bf4e541ef6 · outbound

This paper cites The question asks for the number of members who went to the zoo fewer than 2 times.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning The question asks for the number of members who went to the zoo fewer than 2 times

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.406618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.673870Z digest=sha256:95fc098b3021c4b901c78f3fb4b6c98ad5ef3204262a24f105adc63db94e7319

Observation fe3fca72-1f83-4d48-8cd2-7d375d2ee741 · outbound

This paper cites Trips to the zoo last year.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Trips to the zoo last year

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.391366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.678575Z digest=sha256:c3c95922efef74fe2cf51f34808f1ca4eb2b1a16fddb7548ba0acbde297d8c38

Observation 22bf5ec8-3028-4ca4-9873-d7ad5d388c63 · outbound

This paper cites How many members went to the zoo fewer than 2 times?.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning How many members went to the zoo fewer than 2 times?

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.372692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.683791Z digest=sha256:115956d9d496573a8052eec4485e0b7a2f63529a97b2b61d11309ce06278da6d

Observation 1f55effb-6617-47ee-be73-09c9c8f5733c · outbound

This paper cites fewer than 2 times.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning fewer than 2 times

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.351935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.690962Z digest=sha256:40960ad9ed79821b953487103bc8f4fc60b635d456714b8014c943a78ec91ea3

Observation a0510bc8-d4cd-4edd-a30d-f1790e2e30b8 · outbound

This paper cites fewer than 2 times.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning fewer than 2 times

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.334667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.700337Z digest=sha256:9ef3a75830a9f6c0b9a230bc2f163b554e60081d39e86c48f1cf26d845b1f185

Observation 651623cd-0644-4741-83a2-2d1a7ba97fc0 · outbound

This paper cites Fewer Than.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Fewer Than

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.316276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.706641Z digest=sha256:a533f3f26587fb3d6cbc3f76286102a05eafb59064df919c520474fb2bf0d03a

Observation 55a16e12-8896-4a63-9fd4-b821f7c694d2 · outbound

This paper cites </think> <answer>2</answer> R1-Reward Re�lection patterns! Figure 6:An example of the R1-Reward output.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning </think> <answer>2</answer> R1-Reward Re�lection patterns! Figure 6:An example of the R1-Reward output

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.295886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.712100Z digest=sha256:1be76f199a09b9f3b3666c8998e122a042850d2293d33f7e0d953aa66d1ffb0e

Observation 97d8f434-e052-4cc0-b888-6e8d90a513c4 · outbound

This paper cites Openrlhf: An easy-to-use, scalable and high-performance rlhf framework.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Openrlhf: An easy-to-use, scalable and high-performance rlhf framework.arXiv, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.281527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.716695Z digest=sha256:badf29952a1fcfadf0575c8475f64379478727fafd1a2149d7caf931889afb4d

Observation 75fe7630-2955-4703-8b7e-5f45e6a14300 · outbound

This paper cites Vision-r1: Incentivizing reasoning capability in multimodal large language models.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Vision-r1: Incentivizing reasoning capability in multimodal large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.262438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.721852Z digest=sha256:fedc9307fb96f533f25f469c24ae3f4e7cd09899c26a2796a36f8f41d16df01b

Observation 392f5164-3226-40b1-9fde-d910ea6e82f5 · outbound

This paper cites Minimax-01: Scaling foundation models with lightning attention.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Minimax-01: Scaling foundation models with lightning attention.arXiv, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.245856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.729815Z digest=sha256:f0c300a51ff4aa843885d8be0fb0fc00ffb1ff743cd631a05a16b25cbd28496c

Observation b8fa6320-d3dd-4f88-a8f9-f2189aa021ce · outbound

This paper cites Llava-onevision: Easy visual task transfer.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Llava-onevision: Easy visual task transfer.arXiv, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.225446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.734988Z digest=sha256:eb1fdd8be48066c2404caf337d276a6544bc61fa2288135196fe3940b4ecb4ca

Observation 285971ee-67f0-456f-8969-b20fce5f2c7e · outbound

This paper cites Vision-language intelligence: Tasks, representation learning, and large models.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Vision-language intelligence: Tasks, representation learning, and large models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.205364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.748877Z digest=sha256:ab9b849d2d8af05dab8db44f6b5c0311ffbfb605ffd6b057b258df01c9e4bdb5

Observation 06d1bb01-8744-4a2c-b76f-9d2877b485b1 · outbound

This paper cites Vlrewardbench: A challenging benchmark for vision-language generative reward models.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Vlrewardbench: A challenging benchmark for vision-language generative reward models.arXiv, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.182797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.755774Z digest=sha256:68d2ccc9abd47a78bb470a7f26227260b2f65b8063eb6191a109f210dc3130f5

Observation d80e271e-652e-4fdc-a78e-8e2a3234084b · outbound

This paper cites Silkie: Preference distillation for large visual language models.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Silkie: Preference distillation for large visual language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.168284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.760529Z digest=sha256:e9fc85e6a75cb78a83a52b76850d598befd27e718b5d135916294d179672c4b4

Observation 326f4dbd-c8ec-4924-ab5d-a171f74e2e4f · outbound

This paper cites Baichuan-omni-1.5 technical report.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Baichuan-omni-1.5 technical report.arXiv, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.154717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.768336Z digest=sha256:88405fb9c2a618d5eba12c648d8ca9a816f6ba7761dac0e5a50c7d4855d7b878

Observation 4856f6d7-c358-461e-be78-df354c8dabe9 · outbound

This paper cites Skywork-reward: Bag of tricks for reward modeling in llms.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Skywork-reward: Bag of tricks for reward modeling in llms.arXiv, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.138876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.776831Z digest=sha256:d85f065b735a1d98b759141907e4787b3d00f286e19dfb836dc7db32d1c5b867

Observation e6201145-dadc-4886-a7a3-143e5c127ca4 · outbound

This paper cites Understanding r1-zero-like training: A critical perspective.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Understanding r1-zero-like training: A critical perspective.arXiv, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.116111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.781722Z digest=sha256:0fea3037621c69830ba69ffd1972338a4b9091ffaf133164d261f86b1c3a9e20

Observation e3f3550a-d01d-4233-96c2-6d4f473e9a48 · outbound

This paper cites Inference-time scaling for generalist reward modeling.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Inference-time scaling for generalist reward modeling.arXiv, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:45:54.791273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:45:54.791273Z digest=sha256:87149716b134c31eabe3f0cdf1d8b2cf831e3c462f4d6156faa87cf64354409b

Observation 83c03cab-f52f-442a-824e-e8f9b52ff860 · outbound

This paper cites Visual-rft: Visual reinforcement fine-tuning.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Visual-rft: Visual reinforcement fine-tuning.arXiv, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.085846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.801405Z digest=sha256:5ac371a5d763e98ca4ec36957486cb06a35d784b939bbd4fb788f7d999b41a0c

Observation fa51adfd-45af-4502-b78c-2158638a176b · outbound

This paper cites Uncertainty-aware reward model: Teaching reward models to know what is unknown.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Uncertainty-aware reward model: Teaching reward models to know what is unknown.arXiv, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.071826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.807646Z digest=sha256:b5d7f1aeb4f3c67312ae887770adee4767d061ae46ef20c3879fe0dfb534f6be

Observation 2a1d03dc-ffae-4dc0-9404-fceadd76089b · outbound

This paper cites Dama: Data- and model-aware alignment of multi-modal llms.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Dama: Data- and model-aware alignment of multi-modal llms.arXiv, 2025

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.053529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.815407Z digest=sha256:e92e9b96b192996e96a806e6e1e42eed3b65a34eb29c555b5354c47e01348030

Observation 2dc2cd77-b5b7-4c09-8597-2f7d6b6ab41b · outbound

This paper cites Wildvision: Evaluating vision-language models in the wild with human preferences.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Wildvision: Evaluating vision-language models in the wild with human preferences.arXiv, 2024

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.032286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.819659Z digest=sha256:6d76fd911a9f55e99c812c22a057477b07948ef96b6864a96e64f1c33a013e4b

Observation 0e9e9852-610b-449f-b8ab-4623c874f83e · outbound

This paper cites Mm-eureka: Exploring visual aha moment with rule-based large-scale reinforcement learning.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Mm-eureka: Exploring visual aha moment with rule-based large-scale reinforcement learning.arXiv, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:56.016283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.824535Z digest=sha256:352af093a8e2739274fd8b2f0b9251b44a71482c0976fc4c3965fd9baa0c2a4d

Observation f4752bed-97a4-42c5-ac4a-fecb68d08260 · outbound

This paper cites Inf-orm-llama3.1-70b, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Inf-orm-llama3.1-70b, 2024

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:45:54.829038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:45:54.829038Z digest=sha256:ff68d1d958ea23500fe7577b0f95c33468d162032f2c3c6770a50ab7d6ec6eac

Observation 2e3de1cc-1edb-4a45-bd67-db7abbb59b52 · outbound

This paper cites Introducing openai o1-preview.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Introducing openai o1-preview

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.988798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.833612Z digest=sha256:afe20d7ba0a9ee08db3a6cd35a529baa0140db4b19db62807bfec4a97cf11b1f

Observation bb282e8e-29ee-426b-884f-ae9e8b78b4ad · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 2022.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 2022

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.975346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.838676Z digest=sha256:70ff4dc19e8c652067384d594cbe23f249e698f3fb311e25ef07419e2bd08b51

Observation 1b6e2fd9-13df-4532-89a5-3948b032c380 · outbound

This paper cites an unresolved cited work.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:45:55.958613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.844683Z digest=sha256:52aa1f2f90f4e5417952e7b5436ddcb5004035db029ba42b282a4b9ccc828c65

Observation 869b177d-1f3a-4b6d-8dd5-9c71d98604e6 · outbound

This paper cites Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl.arXiv, 2025

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.937245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.849152Z digest=sha256:3c62f872b64e898bf45d37e274f9fd19b0eb4fac04fc2ac109a51874cfb71b71

Observation 7337b76c-517e-43ca-912c-a73aef08800f · outbound

This paper cites Judge anything: Mllm as a judge across any modality.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Judge anything: Mllm as a judge across any modality.arXiv, 2025

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.909795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.855546Z digest=sha256:70e5c74fe7f02be968db85993274416f2197740501d5806cf07dbc0cd2cb4e68

Observation 59d9b895-07e9-4d3a-911b-907cf7295574 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T00:45:54.859765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:45:54.859765Z digest=sha256:bd19723cecb8d81f3618bdfc0446b5f771d4906e84ace531ad068646d1deaafa

Observation 0e9a62a2-969f-4ca0-9829-5f8e11288e03 · outbound

This paper cites Tapered off-policy reinforce: Stable and efficient reinforcement learning for llms.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Tapered off-policy reinforce: Stable and efficient reinforcement learning for llms.arXiv, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.882684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.868108Z digest=sha256:fe3db2f040e4891fc655add8f236a5ad2cb804a6193ea6da718ba319f70f638a

Observation 17647f99-f08e-4d4e-a706-ee3db05d3947 · outbound

This paper cites Proximal Policy Optimization Algorithms.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T00:45:54.873206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:45:54.873206Z digest=sha256:2f184331fa110d7c4155e231a0b623b77ddfc370db9f09b7c7e8659b00899cd0

Observation c898939a-3f5d-483c-a4c6-69dd25cbf92d · outbound

This paper cites A survey of deep reinforcement learning in video games.arXiv, 2019.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning A survey of deep reinforcement learning in video games.arXiv, 2019

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.867920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.877584Z digest=sha256:5d26d46a7643af8a8b86d088a334421aca51d15dd0d34a1b9550b49dd1e59195

Observation 457ef7e6-6e07-4bc2-9fa9-64107bb3c4d0 · outbound

This paper cites an unresolved cited work.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-16T00:45:55.854657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.882651Z digest=sha256:797c963732b19673efe092fc6f41d5e617c0011cee62341f931bc120e1a1cecd

Observation 76f93430-e0ee-4ea5-ae70-015f30d53ed1 · outbound

This paper cites Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.833166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.888879Z digest=sha256:dcd68d9f0187e022cc99271fbda0cc5f362c99ec6338e91989938a565dba5b19

Observation 8f074774-1988-4478-bde0-7de7fde38731 · outbound

This paper cites Vlm-r1: A stable and generalizable r1-style large vision-language model.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Vlm-r1: A stable and generalizable r1-style large vision-language model.arXiv, 2025

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T00:45:54.894302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:45:54.894302Z digest=sha256:12fc9b375bfcc883b795c799bd75c2e62d9cef7dd70f9251be0fbeb2fb201923

Observation dd6c7e1b-0bd9-40ff-a58d-7409a260058f · outbound

This paper cites Long-vita: Scaling large multi-modal models to 1 million tokens with leading short-context accuracy, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Long-vita: Scaling large multi-modal models to 1 million tokens with leading short-context accuracy, 2025

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.798113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.899951Z digest=sha256:9a637fb6b8078c9ee0692d756a98182bba9e5c5363e0908ae40972d7ab22a308

Observation 507636cd-09ca-41fa-92d9-0ec6824fff84 · outbound

This paper cites Reinforcement learning in robotic applications: a comprehensive survey.Artificial Intelligence Review, 2022.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Reinforcement learning in robotic applications: a comprehensive survey.Artificial Intelligence Review, 2022

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.780791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.907332Z digest=sha256:a0c2cb3c4c8cc129143d98ef8efc610279dc9b03970bfedfced024091e76d302

Observation 0a16e692-3fd3-45e0-b64f-413706271031 · outbound

This paper cites Aligning large multimodal models with factually augmented rlhf.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Aligning large multimodal models with factually augmented rlhf

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T00:45:54.912547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:45:54.912547Z digest=sha256:d7c44e3442b9235e87c2b370cfa77f7fdb2af58559216638963b6b5d06f1eb06

Observation 6dad2849-615c-4abd-8f29-f012d87d53ae · outbound

This paper cites Chameleon: Mixed-modal early-fusion foundation models.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Chameleon: Mixed-modal early-fusion foundation models.arXiv, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.755070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.918050Z digest=sha256:1b64b1db702065b081505b994521ac0d35f6e4548c8ef41019867b71793ed00c

Observation f205254f-772b-40fc-b5fa-c5387f33c3fb · outbound

This paper cites The llama 3 herd of models.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning The llama 3 herd of models.arXiv, 2024

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.737732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.923834Z digest=sha256:b2cbc44ab85b98af0ded1e3e60e67c116de07ecc2f2beee1c89ff1632d11115b

Observation bc20a886-4df0-495b-b118-778fff32e309 · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.arXiv, 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.711585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.932866Z digest=sha256:6917a9a26d1d96d529b412ed1ada162391aad6316ca8bf861c3a5a98bb01f9f6

Observation e9f27689-2eff-49bc-8524-5af18b0daec5 · outbound

This paper cites Visualprm: An effective process reward model for multimodal reasoning, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Visualprm: An effective process reward model for multimodal reasoning, 2025

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.686727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.939560Z digest=sha256:0dfe9738ef4ab06d259cea72788f4620ae4bbe5adcd8ae3a954277cf3e3251f4

Observation 8bfa70aa-7d6a-4bfa-88e1-3093e4ae0726 · outbound

This paper cites Zhang, Makesh Narsimhan Sreedhar, and Oleksii Kuchaiev.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Zhang, Makesh Narsimhan Sreedhar, and Oleksii Kuchaiev

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.661167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.946508Z digest=sha256:e5e86c6b61f2b204869b34644617c0286f6a49c5766403b2a75a7d5b30ac653b

Observation 33151fcb-0fc7-4b94-9f69-18902dc4b638 · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Show-o: One single transformer to unify multimodal understanding and generation.arXiv, 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.637820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.951268Z digest=sha256:b776f81d9dad47283bdcd6e08c795fdad4230968e977a449af7be3bddf871b7b

Observation 6bdcbe9c-c530-49d5-855a-d4da9a6385d0 · outbound

This paper cites Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning.arXiv, 2025

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.619227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.959321Z digest=sha256:f319139ce0791a40f679f676723adcb8a070ad337291cc87f7ff5baedc55f941

Observation 83dec038-2016-4f86-ab72-e0c84b32dceb · outbound

This paper cites Mme-unify: A comprehensive benchmark for unified multimodal understanding and generation models.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Mme-unify: A comprehensive benchmark for unified multimodal understanding and generation models.arXiv, 2025

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.601489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.966320Z digest=sha256:bdc477a971fffa4b3161708e81edcda4912a46c588c43c1d868f47572d609299

Observation 62d898d1-0f5f-400c-bd3e-362c787d7758 · outbound

This paper cites Llava-critic: Learning to evaluate multimodal models.CVPR, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Llava-critic: Learning to evaluate multimodal models.CVPR, 2024

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.584491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.972232Z digest=sha256:95bf1d39b272ff9ef8774f2470e56077d62d0398e8ded5ea74591ae0f5744c64

Observation 528c7a75-1878-4df9-81af-e5c079fea70f · outbound

This paper cites Multimodal rewardbench: Holistic evaluation of reward models for vision language models.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Multimodal rewardbench: Holistic evaluation of reward models for vision language models.arXiv, 2025

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.569880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:54.990104Z digest=sha256:2a4d59c050dedd7868ea9d0c947d1293a2e5d3124a5892b4f059fbd99e74b69c

Observation aac0d211-846d-40c9-bf9c-573071b2b175 · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Dapo: An open-source llm reinforcement learning system at scale.arXiv, 2025

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.555225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:55.002004Z digest=sha256:f6e7f68b7c6b7d6c7eab6b6721983ba497df0bc1b30ad74f522565ec28ac51f0

Observation 0b1c8a2e-6e2b-4d99-8b1e-3cfc264d91c1 · outbound

This paper cites Aligning multimodal llm with human preference: A survey.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Aligning multimodal llm with human preference: A survey.arXiv, 2025

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.538884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:55.009666Z digest=sha256:16b54e67f31e38d4fc783528b89a1aae1bdb4d055686b5c92d45d66a24f95071

Observation 03ce78c1-5db6-425f-8db2-685505623d6b · outbound

This paper cites Rlaif-v: Open-source ai feedback leads to super gpt-4v trustworthiness.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Rlaif-v: Open-source ai feedback leads to super gpt-4v trustworthiness

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.519764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:55.016127Z digest=sha256:0e846ff8d50a273dd9243bc2d263c367f577039f92d3f0e384cf7382723bd778

Observation a4c0cb6f-f282-4512-a554-0d768cfa02f0 · outbound

This paper cites Self-generated critiques boost reward modeling for language models.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Self-generated critiques boost reward modeling for language models.arXiv, 2024

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.503179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:55.022113Z digest=sha256:a16903be16ae2794577484bbda5437693856d19cab8cd40faa2d617dddcb504e

Observation 92cec4b3-6936-4796-84eb-a3167730e1bd · outbound

This paper cites Internlm-xcomposer2.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Internlm-xcomposer2

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.481839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:55.026722Z digest=sha256:4991af7b290985fbb90e1dfa05a69a898532da3263c74b0d91931a26b2f130d0

Observation a4f051c1-a2b4-4033-9295-a45670e74357 · outbound

This paper cites Benchmarking large multimodal models against common corruptions.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Benchmarking large multimodal models against common corruptions.arXiv, 2024

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.465118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:55.033201Z digest=sha256:3d769c4684276fed28e2ef9edfc4f9558904b12d1b4ead039d0b57d2ede6f765

Observation 81a79dd1-9db7-4108-b271-b5e2f1600257 · outbound

This paper cites Llava-mini: Efficient image and video large multimodal models with one vision token, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Llava-mini: Efficient image and video large multimodal models with one vision token, 2025

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.449361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:55.040131Z digest=sha256:9584cb1239bf727068bdba62db43387f51ae0dc6e6d2a1d2089b9ad87712ca14

Observation 4a458172-c9d5-42e8-a7ef-af9a43e89d8f · outbound

This paper cites Beyond llava-hd: Diving into high-resolution large multimodal models.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Beyond llava-hd: Diving into high-resolution large multimodal models.arXiv, 2024

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.433664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:55.045752Z digest=sha256:8a35b54a222b19f1eb958e171713ebdf2b1777f04d72c93e3e3fa89db9144009

Observation a2f11ca7-88fe-4900-b21a-5a96a4474275 · outbound

This paper cites Mm-rlhf: The next step forward in multimodal llm alignment.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Mm-rlhf: The next step forward in multimodal llm alignment.arXiv, 2025

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.418170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:55.050919Z digest=sha256:ab26c6af33a58e678759926a73ae4cdda1f1e10c81588a98ea8f55bac4641eaa

Observation 0f1a756e-2a93-42a6-94a1-2a6725489171 · outbound

This paper cites Debiasing multimodal large language models.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Debiasing multimodal large language models.arXiv, 2024

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.402478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:55.056159Z digest=sha256:1293e82f1592851c5a8f68b77cfc81e6df06311753f103604facc24da52c15a7

Observation db63b4c7-af3e-4e00-a030-5eb62b81ff64 · outbound

This paper cites Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans?ICLR, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans?ICLR, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.383192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:55.060999Z digest=sha256:f489b04a2cf7b587a0c15dac3fbad8fdad0dd33cc6fd9fb651e78f7ca0dd9c6d

Observation 256abbf7-03c3-4ec7-8b47-95948fded276 · outbound

This paper cites R1-omni: Explainable omni-multimodal emotion recognition with reinforcing learning.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning R1-omni: Explainable omni-multimodal emotion recognition with reinforcing learning.arXiv, 2025

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.349884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:55.066162Z digest=sha256:d1edb307bd46fb6aa68f6ed2270c229a29c7fce05d03f643c045248946aa3123

Observation 5eda74c5-2b04-4faa-bdf5-5a5b47387517 · outbound

This paper cites Aligning modalities in vision large language models via preference fine-tuning.arXiv, 2024.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Aligning modalities in vision large language models via preference fine-tuning.arXiv, 2024

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.327768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:55.072115Z digest=sha256:df1891140458344be7916b8445020e914644ae2e8aa27783c95f22fff7113cfa

Observation 254773cd-5e44-4800-a0f5-5b7d39e966bc · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models.arXiv, 2025.

R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models.arXiv, 2025

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:45:55.307320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-16T00:45:55.083191Z digest=sha256:06d781f9893320698622cd788d6a361a577817fd69cbd1223a9cf1976bf2fbcb

Pith citing papers

Observation e582f9d6-11e4-4597-8b15-4d8479bce771 · inbound

Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games cites this paper.

Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:04.359664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:04.359664Z digest=sha256:c55362e1f075c85f2f7e1546d4674b5331b87ffbe65900eca85bf1467e5a3249

Observation 6c8f6167-1acb-41c8-b638-8ff1fd1d1978 · inbound

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models cites this paper.

Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 146

Resolution
unresolved
no resolver link, observed 2026-08-07T14:31:22.005729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:31:22.005729Z digest=sha256:c0972035ebca9a9ed292c7596c18de014a3946864e1267dcc4d0677e7deaf280

Observation aecb0807-75d4-4cd1-9858-fbf31d88585b · inbound

Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning cites this paper.

Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-19T13:02:18.180265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T13:00:27.232079Z digest=sha256:3467862f29a5ec21b1b0e82d92a5e2a02d9d383e10da5058119bb40207372e0c

Observation 5c512187-0b0e-4e92-b81e-804a018e0106 · inbound

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark cites this paper.

Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:19.779262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:19.779262Z digest=sha256:cbd5df8291c4645a2170f381b8d99b16bcd07e3aa906b138ff07c8565fc96afd

Observation c14f5cbe-4c50-4d78-8a14-f4c8a8024c93 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.958836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.958836Z digest=sha256:822ddda943b9f31ed9965089493965759e9e63f444cfdd2d55e5944e18f6224f

Observation c7f2a4b8-a266-44fe-b9f7-48d2ba1cb883 · inbound

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning cites this paper.

SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T11:39:25.091160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T11:39:25.091160Z digest=sha256:4e8a82525320954e65139abc08df16f8bd1b5682eabd681ac8963f1f084bdd20

Observation fb148a72-3212-457f-8a67-f271635f405f · inbound

Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models cites this paper.

Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T03:37:50.232604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:37:50.232604Z digest=sha256:8b8a707b7ba79a71123e049f9bbdb2fe1b6f291af8bbad9ba8091bff9ddf4f5d

Observation 0e55930e-bd20-4101-a527-91a1f656848d · inbound

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation cites this paper.

Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T19:56:33.297932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:56:33.297932Z digest=sha256:74849a7b43a89340e256e05afc09b1db5d0cfcc09720c7f5e524acdfc558b3e6

Observation 20f7ae9c-bff9-4157-9b2c-6b5ac5d508e4 · inbound

Stabilizing Policy Optimization via Logits Convexity cites this paper.

Stabilizing Policy Optimization via Logits Convexity R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T19:53:08.856440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:53:08.856440Z digest=sha256:1b5381b8b3d39ae6bf77da02ecb169efb09cb6d40b496c89caa44e94dd5d980a

Observation 8e4116c8-75f2-4c32-9b37-a1a6eb7e9797 · inbound

StaRPO: Stability-Augmented Reinforcement Policy Optimization cites this paper.

StaRPO: Stability-Augmented Reinforcement Policy Optimization R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:20:59.984367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T18:11:49.056805Z digest=sha256:14a188744acac2e2c7c9a960bce588a215952b01ae0c12d4036c8b17a23d42f8

Observation 3dc8ab51-8e38-4a81-9a20-817baa9ba334 · inbound

Reward-Aware Trajectory Shaping for Few-step Visual Generation cites this paper.

Reward-Aware Trajectory Shaping for Few-step Visual Generation R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:45:21.293870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T11:43:55.499474Z digest=sha256:fa06cee2d721c66c22af42f4df13a54d56b970f734d9d178e2abf0dd2e58aed5

Observation 38709681-ae0f-4c6b-be99-0233d59b65eb · inbound

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling cites this paper.

DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:26:03.799793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T01:45:30.001398Z digest=sha256:3af65becad12acd1e9796835cd4d79ff9a45f5a46984c2498463f0ff395d021f

Observation 876803e7-4eab-4ff8-8f2a-be4e5e29c7ac · inbound

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection cites this paper.

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:36:17.471032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T04:46:16.497585Z digest=sha256:81cbf1973a0f799141d3941f2c103ed54b97c0a84564f8ffdc01003244d69c17

Observation 61e56765-442a-4be9-8d1b-b0dac92a37be · inbound

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling cites this paper.

Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:42:30.874973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T07:37:52.346280Z digest=sha256:ad1bf1872f72213c7442d9c2b070cc41db1d1af3986193d2e9ef1a1dc968ec71

Observation fe7449d0-f8bb-4331-ae5c-70692f8d852c · inbound

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR cites this paper.

Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:36:10.508835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-09T15:39:54.001327Z digest=sha256:7702983051610da2768658fd1cfb12988e16002ef70a4a3892e6fcc8f186e7f6

Observation fce63959-381a-47b7-97f9-ea0711fd5167 · inbound

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models cites this paper.

Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:58.385559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T02:18:20.880231Z digest=sha256:d4587a07047fbacdd88c2d2ff984913979c6e9d0c04967a0fc78ee632da7d89d

Observation a04a95b7-c373-4b04-8db4-ace8ef5aa026 · inbound

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification cites this paper.

DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:01:24.224742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:41:44.833354Z digest=sha256:e26a945efed862ba7aa29543ca8645336ea7befd00f68e367b96725c87f6db26

Observation c2321188-055e-40e6-8c19-fc000d9dd49c · inbound

Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs cites this paper.

Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:46:52.586888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T03:57:57.791351Z digest=sha256:0aff17875453a39c1180d80e2f98755408bce74241d914f8f5510bd3c2a2bd93

Observation bc2c86d0-8265-4c27-8395-a029444a507c · inbound

OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization cites this paper.

OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:21:12.728691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T07:19:57.063802Z digest=sha256:8d3e76d5039e7ea0dbcd846a2ce2bc350944967191a02cc4791ecd68889bf180

Observation 4eadc967-8a10-4d7a-aed1-4aec89d466db · inbound

Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning cites this paper.

Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:54:00.573718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-29T22:53:55.407461Z digest=sha256:07fefdc093b69da915e48dfff930ad7e42e4afe2f170e1758026abb6dc042c14

Observation fb54247d-0abd-426b-9a4a-ad51fb3942ac · inbound

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding cites this paper.

See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:57:29.063754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T17:35:07.667016Z digest=sha256:c59490687844b991b96b67b542786c2166e06c3e0a8e5402e6a9d4ca75f7b3e0

Observation 078cae58-a97a-4fe8-8ac4-a95bd8ec0d75 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 277

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:d7de8578b17b0e9bbdbe59cb0b7efc197a6cbaec621280e41d9e5a99c6bef238

Observation 59f68406-5f9f-496c-b3a2-8f6e11c5da4b · inbound

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification cites this paper.

FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T14:08:11.243120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T14:08:11.243120Z digest=sha256:ad51752a09874aefa3a35d1eff7555ac586e55cf8c2646cb022311a9bdead471