Pith. sign in

Paper Citation Record · LEDGER

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents

As of 12 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 4 inbound Pith citation observations for arXiv:2512.03438.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.03438 v3

Coverage vector

measured 90 of 90 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:51:51.527625Z

measured 94 of 94 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T01:16:44.541597Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T01:37:30.528844Z

Reference resolution

90 of 90 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved89
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b1cf2288-fb57-4cc0-ade0-561e6fae198f · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736,.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:41.814971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:41.814971Z digest=sha256:555b5b5e686165a26f7a76f2c685776bbe7ad9291f3e1364ec90a09c6195933e

Observation b6691b1a-3c08-4130-ac42-9ea98a651408 · outbound

This paper cites Qwen2.5-VL Technical Report.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:41.863940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:41.863940Z digest=sha256:e458884fa0ea1577a229c840d57fa32a726e26d58524970d22a59d0414ceae8f

Observation a31de791-e45c-45f8-8294-2f91e5923d7e · outbound

This paper cites TransDreamer: Reinforcement Learning with Transformer World Models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents TransDreamer: Reinforcement Learning with Transformer World Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:41.949829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:41.949829Z digest=sha256:be336f3696557369b5982cd7068963ddf88e6fca9d2d95aa170abe3b08179208

Observation d9a87b94-e835-4154-aec5-210c4525180e · outbound

This paper cites Train- ing strategies for efficient embodied reasoning.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Train- ing strategies for efficient embodied reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:42.039837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:42.039837Z digest=sha256:50c7f3db7e35ea88ea05594ceb3a81483ad31d6affd3bd061a0c3cffe0f02e19

Observation fa07bc4f-f08a-407a-818f-923bbe67742e · outbound

This paper cites Vision-language models provide promptable representations for reinforcement learning.Transactions on Machine Learn- ing Research, 2025.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Vision-language models provide promptable representations for reinforcement learning.Transactions on Machine Learn- ing Research, 2025

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:42.137544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:42.137544Z digest=sha256:f2ad982c0d54f5fb616c042114e84a88cf9d0787bb71638395014ef9a7d62397

Observation bfc72d3c-103a-4f5a-a2fc-30c662f43eae · outbound

This paper cites Scaling instruction- finetuned language models.Journal of Machine Learning Research, 25(70):1–53, 2024.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Scaling instruction- finetuned language models.Journal of Machine Learning Research, 25(70):1–53, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:42.237400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:42.237400Z digest=sha256:0aece87af135d42dee7eb1573ce8d06642f93b528bb8b8ec7834db22dda42ac8

Observation c9418be7-cc7f-4ccb-a5c8-c188ce4f8d8d · outbound

This paper cites an unresolved cited work.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:42.346027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:42.346027Z digest=sha256:6a91b776c14f8fa86c6e12fb1e14f75de508fdc0bf7b0099462d3bb8d674f726

Observation b657edd2-03da-41c1-ab5a-6cd12e73d80d · outbound

This paper cites Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:42.409109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:42.409109Z digest=sha256:0a9d8019f4c89d28705c164f1705da6bfcc0412b804dd80314656b9206e820cc

Observation 7db0bd1d-abd4-475a-9833-be98c51d2dab · outbound

This paper cites Tool-lmm: A large multi-modal model for tool agent learning.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Tool-lmm: A large multi-modal model for tool agent learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:42.503617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:42.503617Z digest=sha256:73b4fae835576c1f1af442e8b29bda37d33c87e134eea66477b7754fd7b9d984

Observation 14490dd9-9557-4a00-a348-c48c74f2db82 · outbound

This paper cites Palm-e: An embodied multimodal language model.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Palm-e: An embodied multimodal language model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:42.579935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:42.579935Z digest=sha256:996385db7501ea10979aeb4cd3d4f2d296f044b0ec4cc4022d2fa8c83f73c307

Observation 966708de-edc7-413b-b3d8-bf06890198eb · outbound

This paper cites Agent ai: Surveying the horizons of multimodal interaction.CoRR, 2024.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Agent ai: Surveying the horizons of multimodal interaction.CoRR, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:42.629359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:42.629359Z digest=sha256:59597f4b2b127c067aaec10fb108180689df1a7d0dbb60b59592cb000347874b

Observation cad19910-7fe8-4123-a681-b87d362055f9 · outbound

This paper cites GRIT: Teaching MLLMs to Think with Images.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents GRIT: Teaching MLLMs to Think with Images

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:42.710473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:42.710473Z digest=sha256:2fdc4d4e7b00d7efc765123d52a183f45b2a7f41cfd4dca1eeb1b6686f41d748

Observation 59a66006-ca95-4f28-9e78-e8c6cb5aa430 · outbound

This paper cites Video-R1: Reinforcing Video Reasoning in MLLMs.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Video-R1: Reinforcing Video Reasoning in MLLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:42.838254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:42.838254Z digest=sha256:33732f987771a035850e42c95629184bcc13b3c380e2b6b7b928ebaac69c072f

Observation bafa6645-251d-49da-98a0-e45c78ed0176 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Blink: Multimodal large language models can see but not perceive

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:43.003439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:43.003439Z digest=sha256:6782c8e6e580ee38cc0f66fbe7c9a76a96fe924a66ba07fccf2c60a1f2932fb5

Observation 94b94af4-a048-4138-ae21-7b35f0abdf27 · outbound

This paper cites Gemini: A family of highly capable multimodal models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Gemini: A family of highly capable multimodal models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:43.114191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:43.114191Z digest=sha256:2e639f74421eddf42ec4f0e7e58cab8e0725693d6ccba19f2e685f0bc425e13f

Observation 8f509be3-c5fd-4a73-9e5b-ef5ee4edcf1c · outbound

This paper cites Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:43.246845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:43.246845Z digest=sha256:4b3ed92d79895bb2e2463245616594c2ff25e9482c63a529613c825c94b26b32

Observation 6927bdf5-b849-496f-9d59-03d25dde7db5 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:43.340881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:43.340881Z digest=sha256:666faf5a47ca8eec4921e03947dfd52a765b35a2b42d3212f1260614f9cbc8f9

Observation a408dd7b-1330-4754-8ccd-22f5fc8703b1 · outbound

This paper cites Regiongpt: Towards region understanding vision lan- guage model.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Regiongpt: Towards region understanding vision lan- guage model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:43.451963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:43.451963Z digest=sha256:f7e0cbda960e553da3713c15271d2aefc93f775f85d883259059bda55522efb4

Observation 0b8c1652-67e6-4256-9196-2368481fab8e · outbound

This paper cites Mastering Diverse Domains through World Models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Mastering Diverse Domains through World Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:43.577972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:43.577972Z digest=sha256:e953bb4bfd8553cb61ab66e42eb1c710b2fba0c24ee0145acd26546682977d1c

Observation 06f649a4-7afc-4633-8d0d-6e96de609fe4 · outbound

This paper cites Training Agents Inside of Scalable World Models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Training Agents Inside of Scalable World Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:43.701009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:43.701009Z digest=sha256:f89640691488ff1ba32fb4720cb6a5d9b00125bd590f98a054383f0fc8a2c402

Observation e88bb92c-2667-4108-9e58-cb22783dbc26 · outbound

This paper cites Ghil-glue: Hierarchical control with filtered sub- goal images.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Ghil-glue: Hierarchical control with filtered sub- goal images

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:43.791670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:43.791670Z digest=sha256:c65f73dbe1e9fa6d29db6d9336852cc0f08beadf2933eb930a5406134732ba3a

Observation ba529d4e-97b9-4e75-8b2a-8f6b9b26d030 · outbound

This paper cites Breaking the reasoning barrier a survey on llm complex reasoning through the lens of self-evolution.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Breaking the reasoning barrier a survey on llm complex reasoning through the lens of self-evolution

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:43.887743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:43.887743Z digest=sha256:5a55852d0ee40b249f580d1121987a710ac02fa62776a6954561fedf39063703

Observation 4c097d75-dfa5-421b-b6c5-c497715cd5b5 · outbound

This paper cites Glm-4.1 v-thinking: Towards versatile multi- modal reasoning with scalable reinforcement learning.arXiv e-prints, pages arXiv–2507, 2025.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Glm-4.1 v-thinking: Towards versatile multi- modal reasoning with scalable reinforcement learning.arXiv e-prints, pages arXiv–2507, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:43.979744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:43.979744Z digest=sha256:0e2277449ba5ad443e664767f8745f3490c6bcd421304cbd1be214e32f2bd01e

Observation ffa928b4-c1f2-4302-9eb4-8fb33326c27a · outbound

This paper cites Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36:31096–31116,.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36:31096–31116,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:44.131430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:44.131430Z digest=sha256:b6f541ab0935b8e0e1e03f9b98813c7aad3416b7f598a1f859f5f6157ab705e0

Observation 646e10c6-43fc-4eef-9e30-6f6d25d15c7e · outbound

This paper cites Visual language maps for robot navigation.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Visual language maps for robot navigation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:44.320900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:44.320900Z digest=sha256:11ec2af8509a58f3b8818194681588db9ae04cb27953d2423be0539419ee2a3c

Observation 05f55d8d-52ec-499b-9aa8-d814f3fc2811 · outbound

This paper cites Multimodal spatial language maps for robot navi- gation and manipulation.International Journal of Robotics Research (IJRR), 2025.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Multimodal spatial language maps for robot navi- gation and manipulation.International Journal of Robotics Research (IJRR), 2025

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:44.428998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:44.428998Z digest=sha256:2179d800b36fbabfabf586991db9d0310eb476b91436be66cdccf08dda987ee6

Observation 2b0f6080-163b-4528-902a-66b104e49156 · outbound

This paper cites VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:44.540761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:44.540761Z digest=sha256:b931c68d5ab0ef274625c98226a16391885fca2a50bfc93327c1ced068a6e3ce

Observation 5f818a9d-a60a-48ce-8154-b2c83884673b · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.arXiv preprint, 2021.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Scaling up visual and vision-language representation learning with noisy text supervision.arXiv preprint, 2021

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:44.654518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:44.654518Z digest=sha256:42e9db776b172903be90028e0e6215ef29ece026766f468965b2d73225a4c409

Observation c0ebbc32-2467-409c-acc9-d6f6275e9052 · outbound

This paper cites Beyond sight: Fine- tuning generalist robot policies with heterogeneous sensors via language grounding.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Beyond sight: Fine- tuning generalist robot policies with heterogeneous sensors via language grounding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:44.716678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:44.716678Z digest=sha256:65b3165fbca9ef73be1c80d12dc861cf08d7cdf66719e989ebd0041e588531c0

Observation fab59234-3c16-4559-98fa-5b2ebc2fdbd0 · outbound

This paper cites Openvla: An open-source vision-language-action model.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Openvla: An open-source vision-language-action model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:44.792965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:44.792965Z digest=sha256:d971c5d3865c83b169c867090bce0b4fcee535978bbb696f602bf4cfe401d936

Observation bf5ac4f1-f029-44f9-a35d-dd53891240f8 · outbound

This paper cites MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:44.900908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:44.900908Z digest=sha256:32684b4bc128c9190a2ea9f1e145c0bbeadbebbae0a41afcb5ca325196d7c95a

Observation 97ce085a-80b7-48bf-85aa-9cf22d158860 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:45.019738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:45.019738Z digest=sha256:b06d8b63b109c6098660b97c8d9f88b6c438224e94d5a6217a4baa79dabf8184

Observation 4703619f-16c7-4e4e-b1f0-51fd9cb3fff2 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:45.133168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:45.133168Z digest=sha256:51c965f146fc0e9bb4fbbe636ee5d802c736847a93a6db5d70fbd9245f45528f

Observation ec5d3588-2cf9-4939-b169-9d4cc35722ef · outbound

This paper cites Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:45.298957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:45.298957Z digest=sha256:2fbdf6573f75c72c693aa88d5dd4d9e92a87716b24d61a6859e537d0e5bc162f

Observation e9dd06be-0ac6-4558-8fe0-bfafba9c6e65 · outbound

This paper cites Libero: Benchmarking knowl- edge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Libero: Benchmarking knowl- edge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:45.353774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:45.353774Z digest=sha256:815b494d64a7bd81b068953ed14005043d56eb30dc323e7bfb9454ce0918097f

Observation 7a88e144-876e-4915-9400-e565ea15a902 · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:45.408449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:45.408449Z digest=sha256:e979906d0913605661f2ba9cc0bf07821aa39fb3c8e5f740576e9a3cd94afa3b

Observation ee7337f9-87a1-430d-8405-e2ab8aabfd36 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:45.518711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:45.518711Z digest=sha256:ccc15f15c624e43e4fd71341953afc2987d2f607c539af406136c6123b8601ea

Observation bbc43b82-a66d-4a22-94a2-eda8e21733c2 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:45.664347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:45.664347Z digest=sha256:f064e338b974f92618d109954d3a3afe4452ea9e08c2048b46f9f6c21139d9b4

Observation cc92eca6-c13c-45d7-ae70-a3c2b94a34a4 · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:45.837086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:45.837086Z digest=sha256:823aec99f164f7f5b0540dc01e5a6301dbe3a952b6fa237ff487f0289301598e

Observation 27d9199d-ba19-48e9-9704-b2c670d8c5e4 · outbound

This paper cites Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manipu- lation tasks.IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334, 2022.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manipu- lation tasks.IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334, 2022

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:46.012568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:46.012568Z digest=sha256:1aff948c9b551e392eefaf8d8d542052c2c8524a5955f5c2519752e176e878e9

Observation 6a0df662-3761-469a-b3d2-d42a28cd4a77 · outbound

This paper cites Grounding language with visual affordances over unstruc- tured data.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Grounding language with visual affordances over unstruc- tured data

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:46.125999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:46.125999Z digest=sha256:b893929b7407cad95d39b6f24fd85dbafe90e8434717fd9a66c46344cb9a1d92

Observation 93eb17f9-fa90-43ca-a291-44e6cb3b8ed0 · outbound

This paper cites Policy adaptation via language op- timization: Decomposing tasks for few-shot imitation.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Policy adaptation via language op- timization: Decomposing tasks for few-shot imitation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:46.262136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:46.262136Z digest=sha256:99e54a44762e0b782d1d09ad6da1d4e503104ad5ee081fa613054eae8b5d2636

Observation c15984cd-bd6b-4c46-83b8-223a1998eba0 · outbound

This paper cites Steering your generalists: Improving robotic foun- dation models via value guidance.Conference on Robot Learning (CoRL), 2024.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Steering your generalists: Improving robotic foun- dation models via value guidance.Conference on Robot Learning (CoRL), 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:46.346290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:46.346290Z digest=sha256:8b004232d5ddaaaf2e0336f5a42ebf9312e78ecbc1496c488e0a4c25f87ecdf3

Observation 34ac3fed-8674-4559-b95c-8d965880314d · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Representation Learning with Contrastive Predictive Coding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:46.394825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:46.394825Z digest=sha256:af1a08c7608d77737bbb3ee372cba1e811c348d40b4c5ee5f02f92c0b6b3a71a

Observation 0bfe305b-8c9a-4a59-815f-643496dd1596 · outbound

This paper cites Generative agents: Interactive simulacra of human behavior.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Generative agents: Interactive simulacra of human behavior

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:46.517463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:46.517463Z digest=sha256:3f406463c114d76ec7630d7a02b7b57336b95b88aee1dc768b922fefda826962

Observation 27454999-db97-47c2-b26c-6c5e042b97ed · outbound

This paper cites Fast: Efficient action tokenization for vision- language-action models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Fast: Efficient action tokenization for vision- language-action models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:46.638786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:46.638786Z digest=sha256:f0269dc71ad616e5e7a03598c7e4ec8989aab8505a5d2c06dfcd9ed3aca09dde

Observation f359e17f-95ec-409a-b452-504201331a10 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Learn- ing transferable visual models from natural language super- vision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:46.727873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:46.727873Z digest=sha256:8d7161759fee4fd1a00dbf274fd4a0dab4675c49896fa80cc1df84a4c6c04013

Observation 20e86bb1-b537-4db4-af9f-3b80378def10 · outbound

This paper cites Real-world humanoid locomotion with reinforcement learning.Science Robotics, 9(89):eadi9579, 2024.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Real-world humanoid locomotion with reinforcement learning.Science Robotics, 9(89):eadi9579, 2024

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:46.767211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:46.767211Z digest=sha256:b291ce193c0fd4fb0fbec15a3cbd17276d10ae6b90d2b23878f3c64e825af02d

Observation 27eee35f-1de0-4458-ba4b-437bffa5ecc5 · outbound

This paper cites Latent plans for task ag- nostic offline reinforcement learning.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Latent plans for task ag- nostic offline reinforcement learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:46.941419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:46.941419Z digest=sha256:c14ca2e77b27bee8dc07f8c53f7e55bd6c4842808519dd5a9419173d3ed81f2d

Observation 5795de7c-42bb-4675-851b-bab24139991f · outbound

This paper cites Toolformer: Lan- guage models can teach themselves to use tools.Advances in Neural Information Processing Systems, 36:68539–68551,.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Toolformer: Lan- guage models can teach themselves to use tools.Advances in Neural Information Processing Systems, 36:68539–68551,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:47.019376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:47.019376Z digest=sha256:37b4c9d74ac1a36f1fd79858af1c2467ed12f7ac3bba67c1a6615c41400ab530

Observation 2d3625bb-42ab-4e3d-ada3-808ed8820a40 · outbound

This paper cites Robovqa: Multimodal long-horizon reasoning for robotics.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Robovqa: Multimodal long-horizon reasoning for robotics

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:47.099422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:47.099422Z digest=sha256:d6f18bff391e928c9dcc51fa42083e14e441b33bd5f6e8af3934cf330fe2b700

Observation 9d94507e-de9d-42d2-8dee-7d3d51f2cc71 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:47.152960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:47.152960Z digest=sha256:2c2b9fbbbbfbbf22a1e830d643ca1a14a70aa91383f018c41f7e7b4be879911d

Observation ba02a353-00bc-4803-a3f7-8d3c3bab322a · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents HybridFlow: A Flexible and Efficient RLHF Framework

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:47.201232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:47.201232Z digest=sha256:d16a1db5984b2bbf6ae0cf6d2ee26f09f80d763b4a7cd151e7541320534643e0

Observation 3fc048df-a88e-4bce-91d2-e3307ef8f3c0 · outbound

This paper cites Koala: Key frame-conditioned long video-llm.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Koala: Key frame-conditioned long video-llm

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:47.302493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:47.302493Z digest=sha256:2625db734cfbd73b821d2d2231486992368aae3a8dcfd0b77beb9f61f2de8b0c

Observation 366d7f97-a531-4eea-8862-d05cbf46d5bc · outbound

This paper cites Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:47.364014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:47.364014Z digest=sha256:ec2037b9a69e6465b811aa67c7a974b3e028d054c81af46170acb3a9c0f15cad

Observation 88043d61-01f7-41b1-b149-a43c58a44a95 · outbound

This paper cites Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:47.427553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:47.427553Z digest=sha256:63e8caa7f5469d08d08dffaa0d66c3829a0d9a0cc7853441870a7106bccb078b

Observation d29d8af5-e775-48c2-bf6f-88a781b12b0a · outbound

This paper cites Self-Instruct: Aligning Language Models with Self-Generated Instructions.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Self-Instruct: Aligning Language Models with Self-Generated Instructions

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:47.534125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:47.534125Z digest=sha256:6780150288e541f3c69fbba74b3da98660acc08a5f02bae2cd6bdf2fd0dd34ef

Observation 3b5ae711-0391-4605-b1d1-dcfdfbf72183 · outbound

This paper cites Genartist: Multimodal llm as an agent for unified image gen- eration and editing.Advances in Neural Information Pro- cessing Systems, 37:128374–128395, 2024.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Genartist: Multimodal llm as an agent for unified image gen- eration and editing.Advances in Neural Information Pro- cessing Systems, 37:128374–128395, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:47.642128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:47.642128Z digest=sha256:d48be3762ccd73a176420b077175acc1db10ca58a285854c13ac5f4ebae0373c

Observation a68b1fc0-afd0-46c9-b561-3ad287ac16d7 · outbound

This paper cites Blip-3: A family of open large multimodal models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Blip-3: A family of open large multimodal models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:47.769021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:47.769021Z digest=sha256:ec6d0cc06a92a51aa8b733f012f75859afcc37058e8db462bc2ae1aa029e2368

Observation 4b885699-3a4d-4a5d-97e9-929a88efa11e · outbound

This paper cites Magma: A foundation model for multi- modal ai agents.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Magma: A foundation model for multi- modal ai agents

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:47.888856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:47.888856Z digest=sha256:21fd7b8b79a6c702dea5f39748f82f4dc5e9c224e39684965b6a1b35a5503bde

Observation f552b55e-3a86-4cb2-9843-7e8bfaf3efeb · outbound

This paper cites EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:48.033515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:48.033515Z digest=sha256:2b62a532f8c6875323cf0986297a3721fda261b9fc38d58c0981b15fedc6551e

Observation c6a3b328-4a7c-40f1-bbbd-828095bf693c · outbound

This paper cites Embodiedbench: Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Embodiedbench: Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:48.192562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:48.192562Z digest=sha256:b2587e4b93144416132395404bfd4b528db7c19de487263f81f33e2927163231

Observation 4c1e20f1-45fb-463f-851b-1ef2bff7fd45 · outbound

This paper cites React: Synergizing rea- soning and acting in language models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents React: Synergizing rea- soning and acting in language models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:48.308180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:48.308180Z digest=sha256:b75a6d6e216ff4ed7a44fcbae2b5c549f7108dd65dad53acb1fe74544763210a

Observation 92116bc7-0121-4aed-a4df-7aac4060b207 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:48.397328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:48.397328Z digest=sha256:a0757ab038996afefaccffeb34d3340fb8669881d41ebb18c29a5efe62d861e6

Observation ee0aa102-1432-428b-825a-6292d3d7afdb · outbound

This paper cites Spatial mental modeling from limited views.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Spatial mental modeling from limited views

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:48.519217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:48.519217Z digest=sha256:d9192f1892c7b476ef3088e389ffb855e6769df054a3bd59720e1f8920dd147f

Observation e9f9dc04-1acd-416f-ac6a-8f94a592e8ba · outbound

This paper cites MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:48.676604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:48.676604Z digest=sha256:6ef8bdd15717bf4f8eef532c5abe1f922967646951c5f181f2b2681ee466ac1e

Observation 4cad96f5-ebc5-461c-a586-3ed7ba30fb45 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:48.859923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:48.859923Z digest=sha256:ae9f1e47cc865a67c12bcd9c938fecfa1e16f62042e617e121b6b1e8f655c2ca

Observation df1568e0-223a-4b38-a150-4f7ffec200b8 · outbound

This paper cites Robotic control via em- bodied chain-of-thought reasoning.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Robotic control via em- bodied chain-of-thought reasoning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:48.997415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:48.997415Z digest=sha256:c6857eb8012828f769289c2c18650bcced32a4935ba1624849c94ccc3d76b07e

Observation 00327093-cbe3-4fe3-ade1-5fb576ca8d62 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:49.136696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:49.136696Z digest=sha256:e9d2495b77184a5903c984e1f5d870c5ad9c3b23d8f4771609318de28a8a89b8

Observation 246cfcd9-23e8-49f7-ae4b-bc94a50fb253 · outbound

This paper cites CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:49.249079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:49.249079Z digest=sha256:bba57e6a88b20f6b74b77dc70dfc68db56e3bf8c15a10b06fe4d2909a540d9b8

Observation ad87133e-003b-4ec7-921c-3fec66bddc16 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Multimodal Chain-of-Thought Reasoning in Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:49.421484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:49.421484Z digest=sha256:39dfce8137736eb2580791d97e7581d574856362f9d892692fd49002fc2b7053

Observation 59c822db-bcc1-4fc2-bd42-65e226d834ca · outbound

This paper cites Pareto optimal learning for estimating large language model errors.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Pareto optimal learning for estimating large language model errors

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:49.530353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:49.530353Z digest=sha256:32ee1ce584d65eaecc60f2e02b9cfb1f63a5b760eca5dc933156a9d3046a2807

Observation 2bd8d282-6f11-4455-b966-7dfe799193c2 · outbound

This paper cites TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:49.614162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:49.614162Z digest=sha256:384d096e927483544ca4ea128cc3829f7b0d94ba943e42d575d854267d9f3546

Observation 9fc2116e-7316-4876-b6f8-bacbbf998553 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:49.672073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:49.672073Z digest=sha256:9cf364ea55afec0f8a5bc8f55f9874368720c11c2fc88e1a2498784d61171238

Observation 14138785-c689-4ebb-b641-52723afdabed · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:49.723535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:49.723535Z digest=sha256:7dcad6493dd700a67e983ec676a26522a2dfa26d684a4cbb0e7393a60294f675

Observation 51270dc4-ba9b-4d15-a1eb-cd30ff187329 · outbound

This paper cites Around”, “Among.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Around”, “Among

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:49.796853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:49.796853Z digest=sha256:ac35d2aa6d30cdd43b685a4a81801c3ff33697365a2faeaa9acc9bbce64accd6

Observation 153f8c65-9b49-4f3f-80c5-eb866fc8a884 · outbound

This paper cites an unresolved cited work.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:49.958206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:49.958206Z digest=sha256:2b8df3515fed3dff532cc7f25f92b34fce321db5b1e180af53ebea9186f42dc5

Observation bd4f989b-8000-4259-8364-87a665419929 · outbound

This paper cites seedling with roots.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents seedling with roots

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:50.045829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:50.045829Z digest=sha256:b36b55f272e64f8eb1f978f565386d6b3487815fa6ee1e28b29bd6286f151c87

Observation 419d895a-0abf-43ef-a213-09d0f72fcfd9 · outbound

This paper cites an unresolved cited work.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Unresolved cited work

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:50.129764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:50.129764Z digest=sha256:ca93e2202ea8cbaf1b7c5b10a062bd92a18252f71dfb457ef56ab5deebe386ce

Observation f5abf25e-874a-42e7-babe-a9caea7b2c21 · outbound

This paper cites an unresolved cited work.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:50.215470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:50.215470Z digest=sha256:45b1e509b5bcc098d361f049dcd000c7c5df214492f785bbcd452dac7ce52bf9

Observation a0d5d264-79cf-4c2a-aba5-55defbadb592 · outbound

This paper cites roots”, “leaves.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents roots”, “leaves

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:50.274770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:50.274770Z digest=sha256:181e54b436c1d5289823700f072d8cedf8e2f4253dfadca704248a46cdfacc9e

Observation c6c5bc54-bf96-428a-ab7d-c4bfa536e35c · outbound

This paper cites an unresolved cited work.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Unresolved cited work

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:50.404241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:50.404241Z digest=sha256:976094eb5fc3b6d9463cbb1a8e53e89786f801f08804b3f51984919ef28b759f

Observation f190bb95-80d9-4add-84ae-6fccf9ee6395 · outbound

This paper cites an unresolved cited work.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Unresolved cited work

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:50.550354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:50.550354Z digest=sha256:b18d2e5735aabb5ce73cb514ca5dcac6c57426a8ac408c5603af51c16a1dc327

Observation 987f1d23-4f28-4176-a156-461d3db54ff4 · outbound

This paper cites observations.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents observations

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:50.664403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:50.664403Z digest=sha256:ffaa84339e86c04e4e45620ff33603ab5c70b8e25c0d95ae0aa2d1c52082cff5

Observation a9e79421-4b8c-445f-8bd7-929e5eacb754 · outbound

This paper cites anchor_text.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents anchor_text

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:50.802486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:50.802486Z digest=sha256:7e438e1622a49391b16b7b0efdd8cea7e39bfc5c3661e65839fc0919c772fd2c

Observation 7b18387c-e59c-4500-a00c-19effc0567bd · outbound

This paper cites anchor_text.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents anchor_text

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:50.961345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:50.961345Z digest=sha256:bd72caff7108edca316c65ce82bee4c85f0951a0a3767d64356a7fd83027fdfe

Observation e13dee2c-26e0-4a9f-a029-5361ba534362 · outbound

This paper cites anchor_text.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents anchor_text

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:51.124122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:51.124122Z digest=sha256:7f0bbb7be66343de1bc7e7be722f7070ac06274290c9e49c38d9fa2ce4c8af79

Observation 55a34ff7-19a3-4b1f-bb07-3a454b591014 · outbound

This paper cites frame": the nearest explicitly stated single frame number (case-insensitive.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents frame": the nearest explicitly stated single frame number (case-insensitive

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:51.264324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:51.264324Z digest=sha256:2b4d7471323201571659b6454eca7c27ac0e8ae0c94658b2407047b81c965f5c

Observation 60d51d6c-2b70-4969-a895-674698c3220d · outbound

This paper cites frame 6”) and/or a SINGLE time (e.g., “4.47 seconds.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents frame 6”) and/or a SINGLE time (e.g., “4.47 seconds

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-03T18:51:51.401535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:51.401535Z digest=sha256:b13dd1f9d5ddda059d2343f38dc2ebb2eb097637f7c4368324f74d3d9c2c83bf

Observation 4e90fe34-1b6d-4f2e-bfc2-2f00f278d650 · outbound

This paper cites frames 1–6.

Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents frames 1–6

Reference 91

Resolution
malformed identifier
no resolver link, observed 2026-08-03T18:51:51.527625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:51:51.527625Z digest=sha256:123a0b7f08121b8d43022eded68248addd60e41bf5821986c52094ea377ba32b

Pith citing papers

Observation 056266b2-5ce0-43dc-9a29-c00e181962ce · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents

Reference 191

Resolution
verified exact
local_arxiv, observed 2026-05-10T14:00:28.762336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:ed8056c9afb09e35d36b5a34c63baf3235f92337abc2b5dfe9f38b7225ce0ab5

Observation 05dd8d5a-264b-4ebf-a384-e86bce466c59 · inbound

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data cites this paper.

From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:55:00.226075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T18:53:07.734871Z digest=sha256:1b92c0aa5e1d41cf5778ecd59a40e920efd25f8d34affa88acfce0da71ae6919

Observation b86aef18-4663-47d9-952f-b024056ced4f · inbound

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization cites this paper.

Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents

Reference 283

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T01:37:30.530160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T16:26:34.918099Z digest=sha256:eee5308b5c972848bc23f5f656ffa27502286fbb26a48b68e6fecbe37c0cefbb

Observation 258d5d9b-27fe-4947-ae63-2174fdc12742 · inbound

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation cites this paper.

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-01T01:16:44.541597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:16:44.541597Z digest=sha256:645534bb87fb0bd6946c90ae3e9dbd9edd0baa775d405323fd242270ccd6a5ac