Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:29:18.144364Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.08802.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:29:18.144364Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation aaf6133e-29db-4dbb-8298-111327e5ab45 · outbound
Improving Generalization Robustness of Multimodal RLVR Qwen3-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86971bd2-66ba-472d-a305-fbe0af231641 · outbound
Improving Generalization Robustness of Multimodal RLVR Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b542ab2-b3a7-4ded-9438-d09b3f7b79ee · outbound
Improving Generalization Robustness of Multimodal RLVR e-SNLI:Naturallanguage inference with natural language explanations.Advances in Neural Information Processing Systems, 31, 2018
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4e81d58d-f92e-4832-9c99-2e0bc2a36852 · outbound
Improving Generalization Robustness of Multimodal RLVR Seibel, Yu Qiao, and Junjun He
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2778da20-1c78-458c-84fd-1805c863c686 · outbound
Improving Generalization Robustness of Multimodal RLVR When llm meets drl: Advancing jailbreaking efficiency via drl-guided search.Advances in Neural Information Processing Systems, 37:26814–26845, 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2a05e50-9ecb-4dfb-bf07-1055831f30e4 · outbound
Improving Generalization Robustness of Multimodal RLVR DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 604b35ad-5a7b-40a5-9229-abda0a2f9abb · outbound
Improving Generalization Robustness of Multimodal RLVR Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d723a6d4-3adb-441e-b604-cf36f23a40fa · outbound
Improving Generalization Robustness of Multimodal RLVR PathVQA: 30000+ Questions for Medical Visual Question Answering
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbbc29eb-604c-41c2-9ab1-1c06242fc7e5 · outbound
Improving Generalization Robustness of Multimodal RLVR REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7112acee-2383-4d63-9ac7-fbb30715048d · outbound
Improving Generalization Robustness of Multimodal RLVR Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c130c1b-ed88-4af7-816f-f97320837b96 · outbound
Improving Generalization Robustness of Multimodal RLVR Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43aeb306-63ab-4eed-b0e3-aad5661f587c · outbound
Improving Generalization Robustness of Multimodal RLVR Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 91e17568-08bc-4500-aff4-f4ca428ec461 · outbound
Improving Generalization Robustness of Multimodal RLVR Vision matters: Simple visual perturbations can boost multimodal math reasoning.arXiv preprint arXiv:2506.09736, 2025
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 444653c2-bcf0-45dc-b23f-6ed41ff704cb · outbound
Improving Generalization Robustness of Multimodal RLVR R1-fuzz: Specializing language models for textual fuzzing via reinforcement learning.arXiv preprint arXiv:2509.20384, 2025
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 63c451a9-577b-41e9-934d-5404afefecce · outbound
Improving Generalization Robustness of Multimodal RLVR Understanding R1-Zero-Like Training: A Critical Perspective
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 379df114-e619-4950-ba67-8c2d63059ddd · outbound
Improving Generalization Robustness of Multimodal RLVR Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 51f037b6-f459-494a-9bc9-1c14d17ef7da · outbound
Improving Generalization Robustness of Multimodal RLVR R-horizon: How far can your large reasoning model really go in breadth and depth?arXiv preprint arXiv:2510.08189, 2025
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df24018e-f579-457c-b9d6-ff05b4695165 · outbound
Improving Generalization Robustness of Multimodal RLVR Information-theoretic reward decomposition for generalizable rlhf.arXiv preprint arXiv:2504.06020, 2025
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 60f69cd1-d48e-4c15-b216-a4070237d181 · outbound
Improving Generalization Robustness of Multimodal RLVR Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation af77f218-4e1a-4cf7-a9ed-cb3da1b5eeaf · outbound
Improving Generalization Robustness of Multimodal RLVR MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 385a074a-b56d-44ad-adb7-a9ba6a80a30a · outbound
Improving Generalization Robustness of Multimodal RLVR Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74676aeb-72cb-4bb5-858e-5726ffc75fa6 · outbound
Improving Generalization Robustness of Multimodal RLVR Hashimoto, and Percy Liang
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ab29204e-da6b-418e-a008-cd80583451de · outbound
Improving Generalization Robustness of Multimodal RLVR Proximal Policy Optimization Algorithms
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecb6cbf9-65f5-41be-a109-5f6118658ed5 · outbound
Improving Generalization Robustness of Multimodal RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e29b9ae6-d993-412d-a2dc-9d42e90468b1 · outbound
Improving Generalization Robustness of Multimodal RLVR On the value of out-of-distribution testing: An example of Goodhart’s law
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ccb08b73-2409-4eee-b1ef-84683b71978c · outbound
Improving Generalization Robustness of Multimodal RLVR Blaschko, Sien Moens, and Tomasz Stanisławek
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6809d91b-f176-461c-b569-b01a3aa3d178 · outbound
Improving Generalization Robustness of Multimodal RLVR Rlhfpoison: Reward poisoningattackforreinforcementlearningwithhumanfeedbackinlargelanguagemodels
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0d8119a9-4d8d-4ce6-8f7f-4603a83ad0a3 · outbound
Improving Generalization Robustness of Multimodal RLVR Causally- enhanced reinforcement policy optimization.arXiv preprint arXiv:2509.23095, 2025
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15eeec20-fc75-485e-bf42-88045f548f37 · outbound
Improving Generalization Robustness of Multimodal RLVR Adversarial Preference Learning for Robust LLM Alignment
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 13d5d0d6-d30e-4e69-9807-018bf81c3189 · outbound
Improving Generalization Robustness of Multimodal RLVR Dynamic multimodal evaluation with flexible complexity by vision-language bootstrapping
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8eb48e98-1365-49e8-a380-1f6bbe8008eb · outbound
Improving Generalization Robustness of Multimodal RLVR Reward-guided prompt evolving in reinforcement learning for llms
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 69d47463-f76f-4dd3-9a75-8e5e8d431b75 · outbound
Improving Generalization Robustness of Multimodal RLVR Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?Advances in Neural Information Processing Systems, 38:57654–57689, 2025
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5a7101bf-fdf8-48a6-a434-2e89014b78d5 · outbound
Improving Generalization Robustness of Multimodal RLVR A Survey of Reinforcement Learning for Large Reasoning Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d555c56a-27a3-4ad9-8b58-b8a28e521b92 · outbound
Improving Generalization Robustness of Multimodal RLVR Improving reward model generalization from adversarial process enhanced preferences
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 10a383a7-8b2a-429f-8f40-b93264a1baa2 · outbound
Improving Generalization Robustness of Multimodal RLVR Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.