Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T18:51:51.527625Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 4 inbound Pith citation observations for arXiv:2512.03438.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T18:51:51.527625Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T01:16:44.541597Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T01:37:30.528844Z
90 of 90 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b1cf2288-fb57-4cc0-ade0-561e6fae198f · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Flamingo: a visual language model for few-shot learning.Advances in Neural Information Processing Systems, 35:23716–23736,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6691b1a-3c08-4130-ac42-9ea98a651408 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a31de791-e45c-45f8-8294-2f91e5923d7e · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents TransDreamer: Reinforcement Learning with Transformer World Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9a87b94-e835-4154-aec5-210c4525180e · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Train- ing strategies for efficient embodied reasoning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa07bc4f-f08a-407a-818f-923bbe67742e · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Vision-language models provide promptable representations for reinforcement learning.Transactions on Machine Learn- ing Research, 2025
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfc72d3c-103a-4f5a-a2fc-30c662f43eae · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Scaling instruction- finetuned language models.Journal of Machine Learning Research, 25(70):1–53, 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9418be7-cc7f-4ccb-a5c8-c188ce4f8d8d · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b657edd2-03da-41c1-ab5a-6cd12e73d80d · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Molmo and pixmo: Open weights and open data for state-of-the-art multimodal models.arXiv e-prints, pages arXiv–2409, 2024
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7db0bd1d-abd4-475a-9833-be98c51d2dab · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Tool-lmm: A large multi-modal model for tool agent learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14490dd9-9557-4a00-a348-c48c74f2db82 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Palm-e: An embodied multimodal language model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 966708de-edc7-413b-b3d8-bf06890198eb · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Agent ai: Surveying the horizons of multimodal interaction.CoRR, 2024
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cad19910-7fe8-4123-a681-b87d362055f9 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents GRIT: Teaching MLLMs to Think with Images
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59a66006-ca95-4f28-9e78-e8c6cb5aa430 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Video-R1: Reinforcing Video Reasoning in MLLMs
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bafa6645-251d-49da-98a0-e45c78ed0176 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Blink: Multimodal large language models can see but not perceive
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94b94af4-a048-4138-ae21-7b35f0abdf27 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Gemini: A family of highly capable multimodal models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f509be3-c5fd-4a73-9e5b-ef5ee4edcf1c · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Hallusionbench: an advanced diagnos- tic suite for entangled language hallucination and visual il- lusion in large vision-language models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6927bdf5-b849-496f-9d59-03d25dde7db5 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a408dd7b-1330-4754-8ccd-22f5fc8703b1 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Regiongpt: Towards region understanding vision lan- guage model
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b8c1652-67e6-4256-9196-2368481fab8e · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Mastering Diverse Domains through World Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06f649a4-7afc-4633-8d0d-6e96de609fe4 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Training Agents Inside of Scalable World Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e88bb92c-2667-4108-9e58-cb22783dbc26 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Ghil-glue: Hierarchical control with filtered sub- goal images
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba529d4e-97b9-4e75-8b2a-8f6b9b26d030 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Breaking the reasoning barrier a survey on llm complex reasoning through the lens of self-evolution
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c097d75-dfa5-421b-b6c5-c497715cd5b5 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Glm-4.1 v-thinking: Towards versatile multi- modal reasoning with scalable reinforcement learning.arXiv e-prints, pages arXiv–2507, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffa928b4-c1f2-4302-9eb4-8fb33326c27a · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Sugarcrepe: Fixing hackable benchmarks for vision-language compositionality.Advances in neural information processing systems, 36:31096–31116,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 646e10c6-43fc-4eef-9e30-6f6d25d15c7e · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Visual language maps for robot navigation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05f55d8d-52ec-499b-9aa8-d814f3fc2811 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Multimodal spatial language maps for robot navi- gation and manipulation.International Journal of Robotics Research (IJRR), 2025
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b0f6080-163b-4528-902a-66b104e49156 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f818a9d-a60a-48ce-8154-b2c83884673b · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Scaling up visual and vision-language representation learning with noisy text supervision.arXiv preprint, 2021
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0ebbc32-2467-409c-acc9-d6f6275e9052 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Beyond sight: Fine- tuning generalist robot policies with heterogeneous sensors via language grounding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fab59234-3c16-4559-98fa-5b2ebc2fdbd0 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Openvla: An open-source vision-language-action model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf5ac4f1-f029-44f9-a35d-dd53891240f8 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents MMCTAgent: Multi-modal Critical Thinking Agent Framework for Complex Visual Reasoning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97ce085a-80b7-48bf-85aa-9cf22d158860 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4703619f-16c7-4e4e-b1f0-51fd9cb3fff2 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec5d3588-2cf9-4939-b169-9d4cc35722ef · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9dd06be-0ac6-4558-8fe0-bfafba9c6e65 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Libero: Benchmarking knowl- edge transfer for lifelong robot learning.Advances in Neural Information Processing Systems, 36:44776–44791, 2023
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a88e144-876e-4915-9400-e565ea15a902 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee7337f9-87a1-430d-8405-e2ab8aabfd36 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbc43b82-a66d-4a22-94a2-eda8e21733c2 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc92eca6-c13c-45d7-ae70-a3c2b94a34a4 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27d9199d-ba19-48e9-9704-b2c670d8c5e4 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Calvin: A benchmark for language- conditioned policy learning for long-horizon robot manipu- lation tasks.IEEE Robotics and Automation Letters (RA-L), 7(3):7327–7334, 2022
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a0df662-3761-469a-b3d2-d42a28cd4a77 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Grounding language with visual affordances over unstruc- tured data
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93eb17f9-fa90-43ca-a291-44e6cb3b8ed0 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Policy adaptation via language op- timization: Decomposing tasks for few-shot imitation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c15984cd-bd6b-4c46-83b8-223a1998eba0 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Steering your generalists: Improving robotic foun- dation models via value guidance.Conference on Robot Learning (CoRL), 2024
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34ac3fed-8674-4559-b95c-8d965880314d · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Representation Learning with Contrastive Predictive Coding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bfe305b-8c9a-4a59-815f-643496dd1596 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Generative agents: Interactive simulacra of human behavior
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27454999-db97-47c2-b26c-6c5e042b97ed · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Fast: Efficient action tokenization for vision- language-action models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f359e17f-95ec-409a-b452-504201331a10 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Learn- ing transferable visual models from natural language super- vision
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20e86bb1-b537-4db4-af9f-3b80378def10 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Real-world humanoid locomotion with reinforcement learning.Science Robotics, 9(89):eadi9579, 2024
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27eee35f-1de0-4458-ba4b-437bffa5ecc5 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Latent plans for task ag- nostic offline reinforcement learning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5795de7c-42bb-4675-851b-bab24139991f · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Toolformer: Lan- guage models can teach themselves to use tools.Advances in Neural Information Processing Systems, 36:68539–68551,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d3625bb-42ab-4e3d-ada3-808ed8820a40 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Robovqa: Multimodal long-horizon reasoning for robotics
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d94507e-de9d-42d2-8dee-7d3d51f2cc71 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba02a353-00bc-4803-a3f7-8d3c3bab322a · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents HybridFlow: A Flexible and Efficient RLHF Framework
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fc048df-a88e-4bce-91d2-e3307ef8f3c0 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Koala: Key frame-conditioned long video-llm
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 366d7f97-a531-4eea-8862-d05cbf46d5bc · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Ego-R1: Chain-of-Tool-Thought for Ultra-Long Egocentric Video Reasoning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88043d61-01f7-41b1-b149-a43c58a44a95 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356, 2024
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d29d8af5-e775-48c2-bf6f-88a781b12b0a · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Self-Instruct: Aligning Language Models with Self-Generated Instructions
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b5ae711-0391-4605-b1d1-dcfdfbf72183 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Genartist: Multimodal llm as an agent for unified image gen- eration and editing.Advances in Neural Information Pro- cessing Systems, 37:128374–128395, 2024
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a68b1fc0-afd0-46c9-b561-3ad287ac16d7 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Blip-3: A family of open large multimodal models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b885699-3a4d-4a5d-97e9-929a88efa11e · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Magma: A foundation model for multi- modal ai agents
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f552b55e-3a86-4cb2-9843-7e8bfaf3efeb · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6a3b328-4a7c-40f1-bbbd-828095bf693c · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Embodiedbench: Comprehensive benchmarking multi-modal large language models for vision-driven embodied agents
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c1e20f1-45fb-463f-851b-1ef2bff7fd45 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents React: Synergizing rea- soning and acting in language models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92116bc7-0121-4aed-a4df-7aac4060b207 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee0aa102-1432-428b-825a-6292d3d7afdb · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Spatial mental modeling from limited views
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9f9dc04-1acd-416f-ac6a-8f94a592e8ba · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents MMT-Bench: A Comprehensive Multimodal Benchmark for Evaluating Large Vision-Language Models Towards Multitask AGI
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cad96f5-ebc5-461c-a586-3ed7ba30fb45 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df1568e0-223a-4b38-a150-4f7ffec200b8 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Robotic control via em- bodied chain-of-thought reasoning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00327093-cbe3-4fe3-ade1-5fb576ca8d62 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 246cfcd9-23e8-49f7-ae4b-bc94a50fb253 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents CounterCurate: Enhancing Physical and Semantic Visio-Linguistic Compositional Reasoning via Counterfactual Examples
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad87133e-003b-4ec7-921c-3fec66bddc16 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Multimodal Chain-of-Thought Reasoning in Language Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59c822db-bcc1-4fc2-bd42-65e226d834ca · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Pareto optimal learning for estimating large language model errors
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2bd8d282-6f11-4455-b966-7dfe799193c2 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents TraceVLA: Visual Trace Prompting Enhances Spatial-Temporal Awareness for Generalist Robotic Policies
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fc2116e-7316-4876-b6f8-bacbbf998553 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14138785-c689-4ebb-b641-52723afdabed · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51270dc4-ba9b-4d15-a1eb-cd30ff187329 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Around”, “Among
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 153f8c65-9b49-4f3f-80c5-eb866fc8a884 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Unresolved cited work
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd4f989b-8000-4259-8364-87a665419929 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents seedling with roots
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 419d895a-0abf-43ef-a213-09d0f72fcfd9 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Unresolved cited work
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5abf25e-874a-42e7-babe-a9caea7b2c21 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Unresolved cited work
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0d5d264-79cf-4c2a-aba5-55defbadb592 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents roots”, “leaves
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6c5bc54-bf96-428a-ab7d-c4bfa536e35c · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Unresolved cited work
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f190bb95-80d9-4add-84ae-6fccf9ee6395 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents Unresolved cited work
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 987f1d23-4f28-4176-a156-461d3db54ff4 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents observations
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9e79421-4b8c-445f-8bd7-929e5eacb754 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents anchor_text
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b18387c-e59c-4500-a00c-19effc0567bd · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents anchor_text
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e13dee2c-26e0-4a9f-a029-5361ba534362 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents anchor_text
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55a34ff7-19a3-4b1f-bb07-3a454b591014 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents frame": the nearest explicitly stated single frame number (case-insensitive
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60d51d6c-2b70-4969-a895-674698c3220d · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents frame 6”) and/or a SINGLE time (e.g., “4.47 seconds
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e90fe34-1b6d-4f2e-bfc2-2f00f278d650 · outbound
Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents frames 1–6
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 056266b2-5ce0-43dc-9a29-c00e181962ce · inbound
Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents
Reference 191
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 05dd8d5a-264b-4ebf-a384-e86bce466c59 · inbound
From Human Videos to Robot Manipulation: A Survey on Scalable Vision-Language-Action Learning with Human-Centric Data Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b86aef18-4663-47d9-952f-b024056ced4f · inbound
Proxy Reward Internalization and Mechanistic Exploitation: A Learned Precursor to Reward Hacking and Its Generalization Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents
Reference 283
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 258d5d9b-27fe-4947-ae63-2174fdc12742 · inbound
Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation Multimodal Reinforcement Learning with Adaptive Verifier for AI Agents
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.