Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:45:55.083191Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 78 of 78 outbound references and 23 inbound Pith citation observations for arXiv:2505.02835.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:45:55.083191Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:04.359664Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T23:57:29.062496Z
78 of 78 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bb3da6f0-176d-4687-8ed3-ea513e5b55e9 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Pixtral 12b.arXiv, 2024
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5159fb98-5c63-4e79-b22a-264fd0f65345 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Qwen2.5-vl technical report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5b7fb8ec-7d7d-4531-92af-255041bfeb12 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Mllm-as-a-judge: Assessing multimodal llm-as- a-judge with vision-language benchmark
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d0847718-5782-4db8-81f8-7a6fe733c833 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.arXiv, 2024
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fb81d020-4d5d-42e8-90a3-e340f87c56a3 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.arXiv, 2023
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4b712200-ce9d-4d5c-ab84-7d5071faebce · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Sft memorizes, rl generalizes: A comparative study of foundation model post-training.arXiv, 2025
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 144d6781-ff14-453b-ae5c-76e5569b233d · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Process reinforcement through implicit rewards.arXiv, 2025
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5f7225c0-7e6d-4c92-b7b8-6ddfbaf51743 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Nvlm: Open frontier-class multimodal llms.arXiv, 2024
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b45b43a4-1f50-4f86-9368-af0f72100062 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning.arXiv, 2025
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e3b754c6-3549-447f-9814-4726b341c733 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models.arXiv, 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 06880698-8e6f-4e08-831d-1c0a50aa2d42 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Video-r1: Reinforcing video reasoning in mllms.arXiv, 2025
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 044552c2-92d9-409b-9fff-637ec2c0412c · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Vita: Towards open-source interactive omni multimodal llm.arXiv, 2024
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5e1422f8-75f4-4421-9ba1-bebec8f9aac1 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Vita-1.5: Towards gpt-4o level real-time vision and speech interaction.arXiv, 2025
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 94e1e321-9c7f-4ae4-bf68-276b742b1f0c · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Mme-survey: A comprehensive survey on evaluation of multimodal llms.arXiv, 2024
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f8504d21-e52b-4733-9ac5-ffdf864ceb6c · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Trips to the zoo last year
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bf6781c0-2064-4d0d-b120-a3bf4e541ef6 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning The question asks for the number of members who went to the zoo fewer than 2 times
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fe3fca72-1f83-4d48-8cd2-7d375d2ee741 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Trips to the zoo last year
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 22bf5ec8-3028-4ca4-9873-d7ad5d388c63 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning How many members went to the zoo fewer than 2 times?
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1f55effb-6617-47ee-be73-09c9c8f5733c · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning fewer than 2 times
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a0510bc8-d4cd-4edd-a30d-f1790e2e30b8 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning fewer than 2 times
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 651623cd-0644-4741-83a2-2d1a7ba97fc0 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Fewer Than
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 55a16e12-8896-4a63-9fd4-b821f7c694d2 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning </think> <answer>2</answer> R1-Reward Re�lection patterns! Figure 6:An example of the R1-Reward output
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 97d8f434-e052-4cc0-b888-6e8d90a513c4 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Openrlhf: An easy-to-use, scalable and high-performance rlhf framework.arXiv, 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 75fe7630-2955-4703-8b7e-5f45e6a14300 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Vision-r1: Incentivizing reasoning capability in multimodal large language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 392f5164-3226-40b1-9fde-d910ea6e82f5 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Minimax-01: Scaling foundation models with lightning attention.arXiv, 2025
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b8fa6320-d3dd-4f88-a8f9-f2189aa021ce · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Llava-onevision: Easy visual task transfer.arXiv, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 285971ee-67f0-456f-8969-b20fce5f2c7e · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Vision-language intelligence: Tasks, representation learning, and large models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 06d1bb01-8744-4a2c-b76f-9d2877b485b1 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Vlrewardbench: A challenging benchmark for vision-language generative reward models.arXiv, 2024
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d80e271e-652e-4fdc-a78e-8e2a3234084b · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Silkie: Preference distillation for large visual language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 326f4dbd-c8ec-4924-ab5d-a171f74e2e4f · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Baichuan-omni-1.5 technical report.arXiv, 2025
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4856f6d7-c358-461e-be78-df354c8dabe9 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Skywork-reward: Bag of tricks for reward modeling in llms.arXiv, 2024
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e6201145-dadc-4886-a7a3-143e5c127ca4 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Understanding r1-zero-like training: A critical perspective.arXiv, 2025
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e3f3550a-d01d-4233-96c2-6d4f473e9a48 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Inference-time scaling for generalist reward modeling.arXiv, 2025
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83c03cab-f52f-442a-824e-e8f9b52ff860 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Visual-rft: Visual reinforcement fine-tuning.arXiv, 2025
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fa51adfd-45af-4502-b78c-2158638a176b · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Uncertainty-aware reward model: Teaching reward models to know what is unknown.arXiv, 2024
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2a1d03dc-ffae-4dc0-9404-fceadd76089b · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Dama: Data- and model-aware alignment of multi-modal llms.arXiv, 2025
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2dc2cd77-b5b7-4c09-8597-2f7d6b6ab41b · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Wildvision: Evaluating vision-language models in the wild with human preferences.arXiv, 2024
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0e9e9852-610b-449f-b8ab-4623c874f83e · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Mm-eureka: Exploring visual aha moment with rule-based large-scale reinforcement learning.arXiv, 2025
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f4752bed-97a4-42c5-ac4a-fecb68d08260 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Inf-orm-llama3.1-70b, 2024
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e3de1cc-1edb-4a45-bd67-db7abbb59b52 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Introducing openai o1-preview
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bb282e8e-29ee-426b-884f-ae9e8b78b4ad · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 2022
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1b6e2fd9-13df-4532-89a5-3948b032c380 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 869b177d-1f3a-4b6d-8dd5-9c71d98604e6 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Lmm-r1: Empowering 3b lmms with strong reasoning abilities through two-stage rule-based rl.arXiv, 2025
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7337b76c-517e-43ca-912c-a73aef08800f · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Judge anything: Mllm as a judge across any modality.arXiv, 2025
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 59d9b895-07e9-4d3a-911b-907cf7295574 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e9a62a2-969f-4ca0-9829-5f8e11288e03 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Tapered off-policy reinforce: Stable and efficient reinforcement learning for llms.arXiv, 2025
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 17647f99-f08e-4d4e-a706-ee3db05d3947 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c898939a-3f5d-483c-a4c6-69dd25cbf92d · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning A survey of deep reinforcement learning in video games.arXiv, 2019
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 457ef7e6-6e07-4bc2-9fa9-64107bb3c4d0 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 76f93430-e0ee-4ea5-ae70-015f30d53ed1 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Deepseekmath: Pushing the limits of mathematical reasoning in open language models.arXiv, 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8f074774-1988-4478-bde0-7de7fde38731 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Vlm-r1: A stable and generalizable r1-style large vision-language model.arXiv, 2025
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd6c7e1b-0bd9-40ff-a58d-7409a260058f · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Long-vita: Scaling large multi-modal models to 1 million tokens with leading short-context accuracy, 2025
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 507636cd-09ca-41fa-92d9-0ec6824fff84 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Reinforcement learning in robotic applications: a comprehensive survey.Artificial Intelligence Review, 2022
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0a16e692-3fd3-45e0-b64f-413706271031 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Aligning large multimodal models with factually augmented rlhf
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6dad2849-615c-4abd-8f29-f012d87d53ae · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Chameleon: Mixed-modal early-fusion foundation models.arXiv, 2024
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f205254f-772b-40fc-b5fa-c5387f33c3fb · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning The llama 3 herd of models.arXiv, 2024
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bc20a886-4df0-495b-b118-778fff32e309 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution.arXiv, 2024
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e9f27689-2eff-49bc-8524-5af18b0daec5 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Visualprm: An effective process reward model for multimodal reasoning, 2025
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8bfa70aa-7d6a-4bfa-88e1-3093e4ae0726 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Zhang, Makesh Narsimhan Sreedhar, and Oleksii Kuchaiev
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 33151fcb-0fc7-4b94-9f69-18902dc4b638 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Show-o: One single transformer to unify multimodal understanding and generation.arXiv, 2024
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6bdcbe9c-c530-49d5-855a-d4da9a6385d0 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Logic-rl: Unleashing llm reasoning with rule-based reinforcement learning.arXiv, 2025
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 83dec038-2016-4f86-ab72-e0c84b32dceb · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Mme-unify: A comprehensive benchmark for unified multimodal understanding and generation models.arXiv, 2025
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 62d898d1-0f5f-400c-bd3e-362c787d7758 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Llava-critic: Learning to evaluate multimodal models.CVPR, 2024
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 528c7a75-1878-4df9-81af-e5c079fea70f · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Multimodal rewardbench: Holistic evaluation of reward models for vision language models.arXiv, 2025
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation aac0d211-846d-40c9-bf9c-573071b2b175 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Dapo: An open-source llm reinforcement learning system at scale.arXiv, 2025
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0b1c8a2e-6e2b-4d99-8b1e-3cfc264d91c1 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Aligning multimodal llm with human preference: A survey.arXiv, 2025
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 03ce78c1-5db6-425f-8db2-685505623d6b · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Rlaif-v: Open-source ai feedback leads to super gpt-4v trustworthiness
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a4c0cb6f-f282-4512-a554-0d768cfa02f0 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Self-generated critiques boost reward modeling for language models.arXiv, 2024
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 92cec4b3-6936-4796-84eb-a3167730e1bd · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Internlm-xcomposer2
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a4f051c1-a2b4-4033-9295-a45670e74357 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Benchmarking large multimodal models against common corruptions.arXiv, 2024
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 81a79dd1-9db7-4108-b271-b5e2f1600257 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Llava-mini: Efficient image and video large multimodal models with one vision token, 2025
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4a458172-c9d5-42e8-a7ef-af9a43e89d8f · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Beyond llava-hd: Diving into high-resolution large multimodal models.arXiv, 2024
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a2f11ca7-88fe-4900-b21a-5a96a4474275 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Mm-rlhf: The next step forward in multimodal llm alignment.arXiv, 2025
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0f1a756e-2a93-42a6-94a1-2a6725489171 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Debiasing multimodal large language models.arXiv, 2024
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation db63b4c7-af3e-4e00-a030-5eb62b81ff64 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Mme-realworld: Could your multimodal llm challenge high-resolution real-world scenarios that are difficult for humans?ICLR, 2024
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 256abbf7-03c3-4ec7-8b47-95948fded276 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning R1-omni: Explainable omni-multimodal emotion recognition with reinforcing learning.arXiv, 2025
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5eda74c5-2b04-4faa-bdf5-5a5b47387517 · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Aligning modalities in vision large language models via preference fine-tuning.arXiv, 2024
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 254773cd-5e44-4800-a0f5-5b7d39e966bc · outbound
R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models.arXiv, 2025
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e582f9d6-11e4-4597-8b15-4d8479bce771 · inbound
Divide-Fuse-Conquer: Eliciting "Aha Moments" in Multi-Scenario Games R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c8f6167-1acb-41c8-b638-8ff1fd1d1978 · inbound
Reinforcement Fine-Tuning Powers Reasoning Capability of Multimodal Large Language Models R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 146
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aecb0807-75d4-4cd1-9858-fbf31d88585b · inbound
Reason-SVG: Enhancing Structured Reasoning for Vector Graphics Generation with Reinforcement Learning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5c512187-0b0e-4e92-b81e-804a018e0106 · inbound
Evaluating MLLMs with Multimodal Multi-image Reasoning Benchmark R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c14f5cbe-4c50-4d78-8a14-f4c8a8024c93 · inbound
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7f2a4b8-a266-44fe-b9f7-48d2ba1cb883 · inbound
SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb148a72-3212-457f-8a67-f271635f405f · inbound
Joint Reward Modeling: Internalizing Chain-of-Thought for Efficient Visual Reward Models R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e55930e-bd20-4101-a527-91a1f656848d · inbound
Advancing Multimodal Judge Models through a Capability-Oriented Benchmark and MCTS-Driven Data Generation R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20f7ae9c-bff9-4157-9b2c-6b5ac5d508e4 · inbound
Stabilizing Policy Optimization via Logits Convexity R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e4116c8-75f2-4c32-9b37-a1a6eb7e9797 · inbound
StaRPO: Stability-Augmented Reinforcement Policy Optimization R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3dc8ab51-8e38-4a81-9a20-817baa9ba334 · inbound
Reward-Aware Trajectory Shaping for Few-step Visual Generation R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 38709681-ae0f-4c6b-be99-0233d59b65eb · inbound
DT2IT-MRM: Debiased Preference Construction and Iterative Training for Multimodal Reward Modeling R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 876803e7-4eab-4ff8-8f2a-be4e5e29c7ac · inbound
See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 61e56765-442a-4be9-8d1b-b0dac92a37be · inbound
Think, then Score: Decoupled Reasoning and Scoring for Video Reward Modeling R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fe7449d0-f8bb-4331-ae5c-70692f8d852c · inbound
Beyond Uniform Credit Assignment: Selective Eligibility Traces for RLVR R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fce63959-381a-47b7-97f9-ea0711fd5167 · inbound
Video Understanding Reward Modeling: A Robust Benchmark and Performant Reward Models R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a04a95b7-c373-4b04-8db4-ace8ef5aa026 · inbound
DeltaRubric: Generative Multimodal Reward Modeling via Joint Planning and Verification R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c2321188-055e-40e6-8c19-fc000d9dd49c · inbound
Perception Without Engagement: Dissecting the Causal Discovery Deficit in LMMs R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bc2c86d0-8265-4c27-8395-a029444a507c · inbound
OPERA: An Agent for Image Restoration with End-to-End Joint Planning-Execution Optimization R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4eadc967-8a10-4d7a-aed1-4aec89d466db · inbound
Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fb54247d-0abd-426b-9a4a-ad51fb3942ac · inbound
See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 078cae58-a97a-4fe8-8ac4-a95bd8ec0d75 · inbound
BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 277
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59f68406-5f9f-496c-b3a2-8f6e11c5da4b · inbound
FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification R1-Reward: Training Multimodal Reward Model Through Stable Reinforcement Learning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.