Pith. sign in

Paper Citation Record · LEDGER

Improving Generalization Robustness of Multimodal RLVR

As of 14 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2608.08802.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08802 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:29:18.144364Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact3
  • verified fuzzy13
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aaf6133e-29db-4dbb-8298-111327e5ab45 · outbound

This paper cites Qwen3-VL Technical Report.

Improving Generalization Robustness of Multimodal RLVR Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.011972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.011972Z digest=sha256:936fdb611a77071b800a0fb927ce98e3cef5c68219a115c59d15880c4bdae0d9

Observation 86971bd2-66ba-472d-a305-fbe0af231641 · outbound

This paper cites Qwen2.5-VL Technical Report.

Improving Generalization Robustness of Multimodal RLVR Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.016100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.016100Z digest=sha256:f7bcd9c5c203cb73542879e421f9e2d2197fb4506a9e00ef7381458ec8c4f6ec

Observation 7b542ab2-b3a7-4ded-9438-d09b3f7b79ee · outbound

This paper cites e-SNLI:Naturallanguage inference with natural language explanations.Advances in Neural Information Processing Systems, 31, 2018.

Improving Generalization Robustness of Multimodal RLVR e-SNLI:Naturallanguage inference with natural language explanations.Advances in Neural Information Processing Systems, 31, 2018

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.800471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.020070Z digest=sha256:81c592ae1b5215f8f2727127ea0e4010c10d0981ca4fc573d5bf50b6e4f52e9f

Observation 4e81d58d-f92e-4832-9c99-2e0bc2a36852 · outbound

This paper cites Seibel, Yu Qiao, and Junjun He.

Improving Generalization Robustness of Multimodal RLVR Seibel, Yu Qiao, and Junjun He

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.789517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.023601Z digest=sha256:f233ed06eb534c1920e9013df8c86d9f53cf0e2480609cf254caf4f80921fff9

Observation 2778da20-1c78-458c-84fd-1805c863c686 · outbound

This paper cites When llm meets drl: Advancing jailbreaking efficiency via drl-guided search.Advances in Neural Information Processing Systems, 37:26814–26845, 2024.

Improving Generalization Robustness of Multimodal RLVR When llm meets drl: Advancing jailbreaking efficiency via drl-guided search.Advances in Neural Information Processing Systems, 37:26814–26845, 2024

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.027061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.027061Z digest=sha256:364a9d2aac712a3d4ed0155de272507fdc18e332bd8d326bfdcdce199b2df25e

Observation f2a05e50-9ecb-4dfb-bf07-1055831f30e4 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Improving Generalization Robustness of Multimodal RLVR DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.031618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.031618Z digest=sha256:464b49b1e7f549f725ab8e756e4f6d37d0d616cdf1702ee28965e87c55d01852

Observation 604b35ad-5a7b-40a5-9229-abda0a2f9abb · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems.

Improving Generalization Robustness of Multimodal RLVR Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.773084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.037146Z digest=sha256:df2c9acbd11097428895fcc11a0379971939c1fb0cb2e28b83d2608c07e013ed

Observation d723a6d4-3adb-441e-b604-cf36f23a40fa · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

Improving Generalization Robustness of Multimodal RLVR PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.041870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.041870Z digest=sha256:1f1eb5952712984b776f3a2bea364d7eeb68539b6c877a401c3561a6887fdfb9

Observation bbbc29eb-604c-41c2-9ab1-1c06242fc7e5 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Improving Generalization Robustness of Multimodal RLVR REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.045713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.045713Z digest=sha256:2e51b930070fd96eb661bf3cf23cd217ba6729cf538aac54710d4cfebcd744fc

Observation 7112acee-2383-4d63-9ac7-fbb30715048d · outbound

This paper cites Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models.

Improving Generalization Robustness of Multimodal RLVR Toward Evaluative Thinking: Meta Policy Optimization with Evolving Reward Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.049711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.049711Z digest=sha256:7ee76acbef78481d7e3f857fe201ef42aa768ad7e841a789f94435b46d002c13

Observation 5c130c1b-ed88-4af7-816f-f97320837b96 · outbound

This paper cites Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning.

Improving Generalization Robustness of Multimodal RLVR Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.053542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.053542Z digest=sha256:d435e155fe20001607d150ae911a68f6d505d9b0f631fa0ead02ceec9edd5d78

Observation 43aeb306-63ab-4eed-b0e3-aad5661f587c · outbound

This paper cites Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman.

Improving Generalization Robustness of Multimodal RLVR Lau, Soumya Gayen, Asma Ben Abacha, and Dina Demner-Fushman

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.763422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.057998Z digest=sha256:d44d655e15812b0621818316b8c09eb920c291666960da064477c0ac29886130

Observation 91e17568-08bc-4500-aff4-f4ca428ec461 · outbound

This paper cites Vision matters: Simple visual perturbations can boost multimodal math reasoning.arXiv preprint arXiv:2506.09736, 2025.

Improving Generalization Robustness of Multimodal RLVR Vision matters: Simple visual perturbations can boost multimodal math reasoning.arXiv preprint arXiv:2506.09736, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.061645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.061645Z digest=sha256:bdc2f5c9c6359fddc0b0e3078840dbcae9b61496e95cb70966febf5c5faf78cc

Observation 444653c2-bcf0-45dc-b23f-6ed41ff704cb · outbound

This paper cites R1-fuzz: Specializing language models for textual fuzzing via reinforcement learning.arXiv preprint arXiv:2509.20384, 2025.

Improving Generalization Robustness of Multimodal RLVR R1-fuzz: Specializing language models for textual fuzzing via reinforcement learning.arXiv preprint arXiv:2509.20384, 2025

Reference 14

Resolution
verified exact
raw_fallback, observed 2026-08-14T04:29:18.543826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.064643Z digest=sha256:e4fd4931118ff82c6cc40c18b90c7f62395ab5f89aa92cfc76ab1ab13ca09ffb

Observation 63c451a9-577b-41e9-934d-5404afefecce · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Improving Generalization Robustness of Multimodal RLVR Understanding R1-Zero-Like Training: A Critical Perspective

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.068390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.068390Z digest=sha256:0add0a35e847a523dda75d2d66cb2172684048e10b7a4ee811c6c9ecf0dcdb99

Observation 379df114-e619-4950-ba67-8c2d63059ddd · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Improving Generalization Robustness of Multimodal RLVR Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.752938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.071481Z digest=sha256:814620c39ec858ae27ce8215bcb4d2768c68849436991b94dc47602c0e4bd86c

Observation 51f037b6-f459-494a-9bc9-1c14d17ef7da · outbound

This paper cites R-horizon: How far can your large reasoning model really go in breadth and depth?arXiv preprint arXiv:2510.08189, 2025.

Improving Generalization Robustness of Multimodal RLVR R-horizon: How far can your large reasoning model really go in breadth and depth?arXiv preprint arXiv:2510.08189, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.074662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.074662Z digest=sha256:bbdc2cff7fa7cf3e26d396cf5d3a505267b0294df020c9f5d874c64540f21793

Observation df24018e-f579-457c-b9d6-ff05b4695165 · outbound

This paper cites Information-theoretic reward decomposition for generalizable rlhf.arXiv preprint arXiv:2504.06020, 2025.

Improving Generalization Robustness of Multimodal RLVR Information-theoretic reward decomposition for generalizable rlhf.arXiv preprint arXiv:2504.06020, 2025

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-14T04:29:18.427092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.078496Z digest=sha256:33cb0564a022462db44572d52f8a9d8c48781c798a2fcff85e9b9140e7ac148b

Observation 60f69cd1-d48e-4c15-b216-a4070237d181 · outbound

This paper cites an unresolved cited work.

Improving Generalization Robustness of Multimodal RLVR Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-14T04:29:18.741000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.081528Z digest=sha256:37c0a5187ffab9b407ae21e537b00a254149a6962cbf53502e93ccdd1c280430

Observation af77f218-4e1a-4cf7-a9ed-cb3da1b5eeaf · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Improving Generalization Robustness of Multimodal RLVR MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.084973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.084973Z digest=sha256:bf3625e266289886f630cd52437187eb4a3945d069a67c2837641f6243fa4919

Observation 385a074a-b56d-44ad-adb7-a9ba6a80a30a · outbound

This paper cites Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming.

Improving Generalization Robustness of Multimodal RLVR Addressing Benchmarking Gaps in Large Language Models for Health and Medicine with Dynamic Red-Teaming

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.088848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.088848Z digest=sha256:f3acaf8b84ee07c6116b627531d1048baa7372a2631d6baa01e2d47f90705ca1

Observation 74676aeb-72cb-4bb5-858e-5726ffc75fa6 · outbound

This paper cites Hashimoto, and Percy Liang.

Improving Generalization Robustness of Multimodal RLVR Hashimoto, and Percy Liang

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.731708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.092652Z digest=sha256:dcfbde5398578e5e35ef3bdded4a8d8405fe7bfe35d298eb901ff746c8fb0bdb

Observation ab29204e-da6b-418e-a008-cd80583451de · outbound

This paper cites Proximal Policy Optimization Algorithms.

Improving Generalization Robustness of Multimodal RLVR Proximal Policy Optimization Algorithms

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.095895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.095895Z digest=sha256:ee0d14302e6c8ca389fea87d54bf47c613a6ab3cd65c458a4bf817fb1e0627be

Observation ecb6cbf9-65f5-41be-a109-5f6118658ed5 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Improving Generalization Robustness of Multimodal RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.100374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.100374Z digest=sha256:053c28d3cf5cb424f0f6018bf832f6c4e2e60799f46fe100cce977e147ed0cc5

Observation e29b9ae6-d993-412d-a2dc-9d42e90468b1 · outbound

This paper cites On the value of out-of-distribution testing: An example of Goodhart’s law.

Improving Generalization Robustness of Multimodal RLVR On the value of out-of-distribution testing: An example of Goodhart’s law

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.722551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.103728Z digest=sha256:b9e54c20c86ece6bf187164d94e78356c44be0e7e3286708bc9c69acbf8c6655

Observation ccb08b73-2409-4eee-b1ef-84683b71978c · outbound

This paper cites Blaschko, Sien Moens, and Tomasz Stanisławek.

Improving Generalization Robustness of Multimodal RLVR Blaschko, Sien Moens, and Tomasz Stanisławek

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.713377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.107091Z digest=sha256:40891c1a8f043a715ce58533fdcdf1b0cd9895256d6641d9448ba6cdb2f4fd0d

Observation 6809d91b-f176-461c-b569-b01a3aa3d178 · outbound

This paper cites Rlhfpoison: Reward poisoningattackforreinforcementlearningwithhumanfeedbackinlargelanguagemodels.

Improving Generalization Robustness of Multimodal RLVR Rlhfpoison: Reward poisoningattackforreinforcementlearningwithhumanfeedbackinlargelanguagemodels

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.702701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.110801Z digest=sha256:0d54703a2237c64c14a440d094961a5ae88c25988725178ac6d67ea8464ee0cc

Observation 0d8119a9-4d8d-4ce6-8f7f-4603a83ad0a3 · outbound

This paper cites Causally- enhanced reinforcement policy optimization.arXiv preprint arXiv:2509.23095, 2025.

Improving Generalization Robustness of Multimodal RLVR Causally- enhanced reinforcement policy optimization.arXiv preprint arXiv:2509.23095, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.113870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.113870Z digest=sha256:b3d9f048886de2e2d9471366bd8622694dd1baa00812c8da8f2cf08bc07e9533

Observation 15eeec20-fc75-485e-bf42-88045f548f37 · outbound

This paper cites Adversarial Preference Learning for Robust LLM Alignment.

Improving Generalization Robustness of Multimodal RLVR Adversarial Preference Learning for Robust LLM Alignment

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-14T04:29:18.189926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.117327Z digest=sha256:0ffc2cad4ca68b7a50afe40c2e1b9b8ef80fe189ec73abd36731668d4bc716df

Observation 13d5d0d6-d30e-4e69-9807-018bf81c3189 · outbound

This paper cites Dynamic multimodal evaluation with flexible complexity by vision-language bootstrapping.

Improving Generalization Robustness of Multimodal RLVR Dynamic multimodal evaluation with flexible complexity by vision-language bootstrapping

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.694314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.124081Z digest=sha256:ef17a276f171d1f96f7b1211d902ebe151c756e2e6d103853dbd1392b79862a3

Observation 8eb48e98-1365-49e8-a380-1f6bbe8008eb · outbound

This paper cites Reward-guided prompt evolving in reinforcement learning for llms.

Improving Generalization Robustness of Multimodal RLVR Reward-guided prompt evolving in reinforcement learning for llms

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.684846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.128751Z digest=sha256:6d469eea3bb6ad4d0ad0c6359b678c5280327bb58a167bf9b71a2ec9f3b80f54

Observation 69d47463-f76f-4dd3-9a75-8e5e8d431b75 · outbound

This paper cites Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?Advances in Neural Information Processing Systems, 38:57654–57689, 2025.

Improving Generalization Robustness of Multimodal RLVR Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?Advances in Neural Information Processing Systems, 38:57654–57689, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.673253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.132174Z digest=sha256:42a8fa9d828b67257cbc97c1c8bdd3407b5ae8d7af864e68149e23d681766b83

Observation 5a7101bf-fdf8-48a6-a434-2e89014b78d5 · outbound

This paper cites A Survey of Reinforcement Learning for Large Reasoning Models.

Improving Generalization Robustness of Multimodal RLVR A Survey of Reinforcement Learning for Large Reasoning Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.137778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.137778Z digest=sha256:3f522bd6f4f589a91ce8fbbef540808d790c8e3c65149bada601cd13c12eafbc

Observation d555c56a-27a3-4ad9-8b58-b8a28e521b92 · outbound

This paper cites Improving reward model generalization from adversarial process enhanced preferences.

Improving Generalization Robustness of Multimodal RLVR Improving reward model generalization from adversarial process enhanced preferences

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T04:29:18.662872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T04:29:18.141262Z digest=sha256:c2a0b73fdb82021eed516eef0acdae7b176af57c251e4924f602f1c82d069349

Observation 10a383a7-8b2a-429f-8f40-b93264a1baa2 · outbound

This paper cites Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models.

Improving Generalization Robustness of Multimodal RLVR Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-14T04:29:18.144364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:29:18.144364Z digest=sha256:83a8e8d5aa1346e39cbfc3f704ee5f0618de0c87a968e5b7ac8dd57f3365438d

Pith citing papers

No inbound Pith citation observations are available.