Pith. sign in

Paper Citation Record · LEDGER

Improving Generalization Robustness of Multimodal RLVR

As of 15 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.08802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08802 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:29:18.144364Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aaf6133e-29db-4dbb-8298-111327e5ab45 · outbound

This paper cites Qwen3-VL Technical Report.

Improving Generalization Robustness of Multimodal RLVR Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.011972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.011972Z digest=sha256:a810ea316a6cf8187c65903f59ed4d341d17a2d93e3f7124f0432ec495002dfd

Observation 86971bd2-66ba-472d-a305-fbe0af231641 · outbound

This paper cites Qwen2.5-VL Technical Report.

Improving Generalization Robustness of Multimodal RLVR Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.016100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.016100Z digest=sha256:e73998aca1a216815fbe38e2ee57f3646e40c8e710d915fa1d67044a546c5b0c

Observation 7b542ab2-b3a7-4ded-9438-d09b3f7b79ee · outbound

This paper cites e-SNLI:Naturallanguage inference with natural language explanations.Advances in Neural Information Processing Systems, 31, 2018.

Improving Generalization Robustness of Multimodal RLVR e-SNLI:Naturallanguage inference with natural language explanations.Advances in Neural Information Processing Systems, 31, 2018

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.800471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.020070Z digest=sha256:afd302afe11d8e02c884e38c968e6d1b9d7fc71bcacb5c0bfc9c690f8418dc80

Observation 4e81d58d-f92e-4832-9c99-2e0bc2a36852 · outbound

This paper cites Seibel, Yu Qiao, and Junjun He.

Improving Generalization Robustness of Multimodal RLVR Seibel, Yu Qiao, and Junjun He

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.789517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.023601Z digest=sha256:724f30471f3b1a6d71bc8e7d4d0db4a9a809a31c6d4b3eaf95f29675f6f230cb

Observation 2778da20-1c78-458c-84fd-1805c863c686 · outbound

This paper cites When llm meets drl: Advancing jailbreaking efficiency via drl-guided search.Advances in Neural Information Processing Systems, 37:26814–26845, 2024.

Improving Generalization Robustness of Multimodal RLVR When llm meets drl: Advancing jailbreaking efficiency via drl-guided search.Advances in Neural Information Processing Systems, 37:26814–26845, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.027061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.027061Z digest=sha256:940206818a06259dfa16e23d803ffc2e5d1e74fe6d7f807294e21444ce334d6d

Observation f2a05e50-9ecb-4dfb-bf07-1055831f30e4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Improving Generalization Robustness of Multimodal RLVR DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.031618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.031618Z digest=sha256:164e0806ed74ef9a09fc3264d0bc081385b047c0adb3b29663828733d7a6e8e8

Observation 604b35ad-5a7b-40a5-9229-abda0a2f9abb · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems.

Improving Generalization Robustness of Multimodal RLVR Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.773084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.037146Z digest=sha256:7211a4ee1098ca96443c1a0840e19cfeb24d36c7097f0ddc7af48cc12793ce4b

Observation d723a6d4-3adb-441e-b604-cf36f23a40fa · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

Improving Generalization Robustness of Multimodal RLVR PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.041870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.041870Z digest=sha256:3dbe0b771c9cfc8b4665ba0a98a93960d2a1c8f684f7e2cf6de4e28b8da85854

Observation bbbc29eb-604c-41c2-9ab1-1c06242fc7e5 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Improving Generalization Robustness of Multimodal RLVR REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.045713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.045713Z digest=sha256:4f418bb3ea71ae8ff925fdf667f45fe95ffa8429da98b84f670cda4ec76cd43d

Observation 7112acee-2383-4d63-9ac7-fbb30715048d · outbound

This paper cites Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models.

Improving Generalization Robustness of Multimodal RLVR Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.049711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.049711Z digest=sha256:9dde44f507b18a3b1753cce69231a043e62f63cfaa0197d2218d85f7866aa606

Observation 5c130c1b-ed88-4af7-816f-f97320837b96 · outbound

This paper cites Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning.

Improving Generalization Robustness of Multimodal RLVR Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.053542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.053542Z digest=sha256:7ad5c0163eb17be08379a858c4695e3c7f198ab9adad3f18f8740ed5de0fea0f

Observation 43aeb306-63ab-4eed-b0e3-aad5661f587c · outbound

This paper cites Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman.

Improving Generalization Robustness of Multimodal RLVR Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.763422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.057998Z digest=sha256:4ba5c4a1bddfbfae51183b7c711ef02989a3010da94a2a783d27a5dcee7fca54

Observation 91e17568-08bc-4500-aff4-f4ca428ec461 · outbound

This paper cites Vision matters: Simple visual perturbations can boost multimodal math reasoning.arXiv preprint arXiv:2506.09736, 2025.

Improving Generalization Robustness of Multimodal RLVR Vision matters: Simple visual perturbations can boost multimodal math reasoning.arXiv preprint arXiv:2506.09736, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.061645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.061645Z digest=sha256:ae502c428140b85be3f57a4bdf2edbef3e9b0758a6e05506774197b4ce74a7aa

Observation 444653c2-bcf0-45dc-b23f-6ed41ff704cb · outbound

This paper cites R1-fuzz: Specializing language models for textual fuzzing via reinforcement learning.arXiv preprint arXiv:2509.20384, 2025.

Improving Generalization Robustness of Multimodal RLVR R1-fuzz: Specializing language models for textual fuzzing via reinforcement learning.arXiv preprint arXiv:2509.20384, 2025

Reference 14

Resolution
verified exact
raw_fallback, observed 2026-08-14T04:29:18.543826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.064643Z digest=sha256:0566413d601cb6c30633ddecba7a9d41064d52bbb925165b05ae1eefca9b73fb

Observation 63c451a9-577b-41e9-934d-5404afefecce · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Improving Generalization Robustness of Multimodal RLVR Understanding R1-Zero-Like Training: A Critical Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.068390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.068390Z digest=sha256:1be0800758bfabfd12bed3859a8da8d77409a05b3946c3c5d87b82433778998e

Observation 379df114-e619-4950-ba67-8c2d63059ddd · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Improving Generalization Robustness of Multimodal RLVR Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.752938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.071481Z digest=sha256:788eba79db98e8dcb605207432151b1b32bcfa7eda3f4350474d287b30f926c0

Observation 51f037b6-f459-494a-9bc9-1c14d17ef7da · outbound

This paper cites R-horizon: How far can your large reasoning model really go in breadth and depth?arXiv preprint arXiv:2510.08189, 2025.

Improving Generalization Robustness of Multimodal RLVR R-horizon: How far can your large reasoning model really go in breadth and depth?arXiv preprint arXiv:2510.08189, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.074662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.074662Z digest=sha256:02b72a337ba146810dd020861f5bd6be6b9e45ab98bccddce3bc62d45d32c101

Observation df24018e-f579-457c-b9d6-ff05b4695165 · outbound

This paper cites Information-theoretic reward decomposition for generalizable rlhf.arXiv preprint arXiv:2504.06020, 2025.

Improving Generalization Robustness of Multimodal RLVR Information-theoretic reward decomposition for generalizable rlhf.arXiv preprint arXiv:2504.06020, 2025

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-14T04:29:18.427092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.078496Z digest=sha256:7b364030c218c5235b4dbc65cba6c854b002071918fdb0eb4812a53f8a359df4

Observation 60f69cd1-d48e-4c15-b216-a4070237d181 · outbound

This paper cites an unresolved cited work.

Improving Generalization Robustness of Multimodal RLVR Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:29:18.741000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.081528Z digest=sha256:4375908c9c86954fd23af8712a3a8ebec2ed4e7b4071cacf6c9d85d3eda0fc9e

Observation af77f218-4e1a-4cf7-a9ed-cb3da1b5eeaf · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Improving Generalization Robustness of Multimodal RLVR MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.084973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.084973Z digest=sha256:c28bee2635c20c6729d70c635c06a8c95dee2a0d76ee1328b37e9807f7d6dcf2

Observation 385a074a-b56d-44ad-adb7-a9ba6a80a30a · outbound

This paper cites Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming.

Improving Generalization Robustness of Multimodal RLVR Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.088848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.088848Z digest=sha256:7edf98fad5f74cf5bdaf75833f1548b97277d7ea9e26018496ce64722418a566

Observation 74676aeb-72cb-4bb5-858e-5726ffc75fa6 · outbound

This paper cites Hashimoto, and Percy Liang.

Improving Generalization Robustness of Multimodal RLVR Hashimoto, and Percy Liang

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.731708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.092652Z digest=sha256:7e7190284fe23c14e314b7e31d0b4920709d754b2e71192c563aa5df20812875

Observation ab29204e-da6b-418e-a008-cd80583451de · outbound

This paper cites Proximal Policy Optimization Algorithms.

Improving Generalization Robustness of Multimodal RLVR Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.095895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.095895Z digest=sha256:dcea53f45fcd3db19e33577f31a9f9de62495a3c5376f9318a0626a4ff8a4285

Observation ecb6cbf9-65f5-41be-a109-5f6118658ed5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Improving Generalization Robustness of Multimodal RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.100374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.100374Z digest=sha256:0b1054f62755c0db3cecf007df76091bd5d79c00184e7e62335d849f9ae3c29d

Observation e29b9ae6-d993-412d-a2dc-9d42e90468b1 · outbound

This paper cites On the value of out-of-distribution testing: An example of Goodhart’s law.

Improving Generalization Robustness of Multimodal RLVR On the value of out-of-distribution testing: An example of Goodhart’s law

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.722551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.103728Z digest=sha256:6f0bbea1900d7fad9f8d36cbfedc1f3c96db44b6afe3eecd2583105c95c29265

Observation ccb08b73-2409-4eee-b1ef-84683b71978c · outbound

This paper cites Blaschko, Sien Moens, and Tomasz Stanisławek.

Improving Generalization Robustness of Multimodal RLVR Blaschko, Sien Moens, and Tomasz Stanisławek

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.713377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.107091Z digest=sha256:da41d12061df2a632ae40dbe5716eff5548c6965e083594e6edbf7f95b45ed56

Observation 6809d91b-f176-461c-b569-b01a3aa3d178 · outbound

This paper cites Rlhfpoison: Reward poisoningattackforreinforcementlearningwithhumanfeedbackinlargelanguagemodels.

Improving Generalization Robustness of Multimodal RLVR Rlhfpoison: Reward poisoningattackforreinforcementlearningwithhumanfeedbackinlargelanguagemodels

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.702701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.110801Z digest=sha256:b78534c0a6d7e296f8c8ce60246c2c368073b4545edfbec8280ff564ccf56aeb

Observation 0d8119a9-4d8d-4ce6-8f7f-4603a83ad0a3 · outbound

This paper cites Causally- enhanced reinforcement policy optimization.arXiv preprint arXiv:2509.23095, 2025.

Improving Generalization Robustness of Multimodal RLVR Causally- enhanced reinforcement policy optimization.arXiv preprint arXiv:2509.23095, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.113870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.113870Z digest=sha256:ce2d359b2e28e73bca5336118ae477ef5f4f1c22215ddbebeaf008da7d182dbf

Observation 15eeec20-fc75-485e-bf42-88045f548f37 · outbound

This paper cites Adversarial Preference Learning for Robust LLM Alignment.

Improving Generalization Robustness of Multimodal RLVR Adversarial Preference Learning for Robust LLM Alignment

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:29:18.189926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.117327Z digest=sha256:d4bdd712e996e841a3010d20023c3a607588f74c9cecf845fbe4e2a91701081e

Observation 13d5d0d6-d30e-4e69-9807-018bf81c3189 · outbound

This paper cites Dynamic multimodal evaluation with flexible complexity by vision-language bootstrapping.

Improving Generalization Robustness of Multimodal RLVR Dynamic multimodal evaluation with flexible complexity by vision-language bootstrapping

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.694314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.124081Z digest=sha256:39998f99ee55eb3b591461b0f4f63d24cb199c4192221a2cf843d90a4bebb806

Observation 8eb48e98-1365-49e8-a380-1f6bbe8008eb · outbound

This paper cites Reward-guided prompt evolving in reinforcement learning for llms.

Improving Generalization Robustness of Multimodal RLVR Reward-guided prompt evolving in reinforcement learning for llms

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.684846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.128751Z digest=sha256:88b3566dace5928fe195856ea7847623b06b18d38403443b7b180dfe5a964c85

Observation 69d47463-f76f-4dd3-9a75-8e5e8d431b75 · outbound

This paper cites Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?Advances in Neural Information Processing Systems, 38:57654–57689, 2025.

Improving Generalization Robustness of Multimodal RLVR Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?Advances in Neural Information Processing Systems, 38:57654–57689, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.673253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.132174Z digest=sha256:868d491d0275e388ecaff70a3a14d66bbf15f5109e581f91eb2c410f538938f2

Observation 5a7101bf-fdf8-48a6-a434-2e89014b78d5 · outbound

This paper cites A Survey of Reinforcement Learning for Large Reasoning Models.

Improving Generalization Robustness of Multimodal RLVR A Survey of Reinforcement Learning for Large Reasoning Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.137778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.137778Z digest=sha256:4a9b749c21ceacffef1eb4897fd61f0a1346c735460092341977e39f89bbe408

Observation d555c56a-27a3-4ad9-8b58-b8a28e521b92 · outbound

This paper cites Improving reward model generalization from adversarial process enhanced preferences.

Improving Generalization Robustness of Multimodal RLVR Improving reward model generalization from adversarial process enhanced preferences

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.662872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T04:29:18.141262Z digest=sha256:b09f041d23443904b00407122e285245432e6516445ba5085df6919236399a17

Observation 10a383a7-8b2a-429f-8f40-b93264a1baa2 · outbound

This paper cites Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models.

Improving Generalization Robustness of Multimodal RLVR Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.144364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.144364Z digest=sha256:8e68a52923889f979a418e2de22a4f01296beee639b42dbb9035db2a45eed289

Pith citing papers

No inbound Pith citation observations are available.