Pith. sign in

Paper Citation Record · LEDGER

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization

As of 2 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 1 inbound Pith citation observation for arXiv:2604.06777.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.06777 v1

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T18:20:02.559108Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-01T06:32:01.292127+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-30T21:38:08.611674Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact45
  • verified fuzzy29
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0824279e-b1b2-48f0-90fc-9e29c3eb427d · outbound

This paper cites Qwen2.5-VL Technical Report.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.219986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:c2a3510cdc7fe5d977b93434f62f4347e638f9a6584650f0e59b8e6c6f5e3b9d

Observation 51bd245e-7b26-474a-af68-f51e37231b67 · outbound

This paper cites Seed1.5-VL Technical Report.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Seed1.5-VL Technical Report

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:c817c1d5651e96b9f26be30e8f60519c10dbaa57616e31494cd3e27025e1884c

Observation d4849641-f7b0-4652-aea1-879283ec4a98 · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Visual sketchpad: Sketching as a visual chain of thought for multimodal language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.448632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:2ea293a70d262bb570c9a38d9a09f3f5ecce2715b54fba2688b59a1ebbf01eed

Observation 57061bdd-2297-40f8-8564-6eed3fef0bba · outbound

This paper cites Thinking with images.https://openai.com/index/thinking-with-images/.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Thinking with images.https://openai.com/index/thinking-with-images/

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.451898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:c010a88339bba3bb7fd8f34bbf527b9e39d564016613ef4c53650127a2dff626

Observation c76b4e67-8fd3-4a5c-a447-1bf63fb47abe · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:42:57.181531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:17c4fa501d303838661c6daf04fc4dd1234ebff3acbb250c201627f113ed8729

Observation 2734c21c-f751-4c00-b0be-0b6995e451fe · outbound

This paper cites Thyme: Think Beyond Images.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Thyme: Think Beyond Images

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:33:29.495992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:304e073adb7fc7d573f334d7e5fd75bbbad5aac17cbde8288d0a6f1d757f872c

Observation 5faa6fd7-50f9-480d-be80-305e691693eb · outbound

This paper cites Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:17:55.768539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:91de054012a360a5e3f81be3007ffe5a2df6373238f44b8f196e5cb06bda5d2b

Observation 0a36ce3a-39eb-4ed4-a0ae-01d8473eebaf · outbound

This paper cites Deep but reliable: Advancing multi-turn reasoning for thinking with images.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Deep but reliable: Advancing multi-turn reasoning for thinking with images

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.046372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:fce94374712b4c5ecf436cec20b7a3b343742b948c7e1926b391ecaa7c20cf5b

Observation 91322220-7493-4638-8c19-141e56fe2a83 · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge.https://llava-vl.github.io/blog/ 2024-01-30-llava-next/.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Llavanext: Improved reasoning, ocr, and world knowledge.https://llava-vl.github.io/blog/ 2024-01-30-llava-next/

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.542218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:d80cdba7eed1f9ebf43360f538c0d6a330f6ead8d516af5ff9bd560723b825eb

Observation 99136b9d-790d-456a-8fd8-33e6f24f0326 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.049454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:229a23646d2d82610f8081144371c0f00a72142f988327f9f1a0e8209f530887

Observation 14d5b53b-7dab-4662-a12b-6e333a1b4f5e · outbound

This paper cites Kwai Keye-VL Technical Report.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Kwai Keye-VL Technical Report

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.095758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:d03d807c57cee73836cf04640083b392be63325fb11a407b63d8778567a697e7

Observation b602e8d3-fff2-446a-bea0-0b11e78b8181 · outbound

This paper cites Ovis2.5 Technical Report.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Ovis2.5 Technical Report

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:30:17.288454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:2d9cbf21c532524b1ba7db4ebb9934b5bc498965851bf6d19148ce62f826f8ff

Observation 9c676ea5-6e55-4fa4-b1ba-e95346b6a391 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Learning transferable visual models from natural language supervision

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.535839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:46d401f5178207c1f0820c7ff4db58571bd6530bd66cddbcb80e0fa8828895dc

Observation 74bfbb62-cd82-4e0c-a785-3622ec31e4ba · outbound

This paper cites BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization BLIP: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.539324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:f4eb4ed797162cfa4c0da29e4299f83d65bcd40bc035e1c4531b293e42537da8

Observation ed505a88-b6fd-424b-ae49-9b8d9cf4d93c · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.557477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:099ee3957eba5d2e775f6cbe636fc0e41ae4ef8e0d5fc047afc41f438e5fd3d3

Observation aef8e27e-8660-4c24-a180-2b865840a4e7 · outbound

This paper cites Sigmoid loss for language image pre- training.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Sigmoid loss for language image pre- training

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.554326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:34609c7cce02173edae428872cc8193aecd51fb2198376eaebdca2a477534a1c

Observation 7e46cd40-1d3d-4923-ae93-887f65da2b20 · outbound

This paper cites Show and tell: A neural image caption generator.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Show and tell: A neural image caption generator

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.517695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:01b7e879eb9b68c50d67669195249ec2e53df1a78c0c97a85499579ccd9f3884

Observation 24f1d174-58b8-41e1-8930-5bd8892a1e65 · outbound

This paper cites Show, attend and tell: Neural image caption generation with visual attention.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Show, attend and tell: Neural image caption generation with visual attention

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.486859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:198f196812849d2c7b77334b013d7768d2a34a64f56ece46c7a03eebff0134b3

Observation 1affeafd-4778-41de-8ea4-492c579c9f8c · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Flamingo: a visual language model for few-shot learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.499234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:e2ceab368a835e9fd4d82d3e5cd473add7d7d031b33ca9461e79e18adcc8fb68

Observation e592f24f-5b3f-44e7-bd48-d6409248ea92 · outbound

This paper cites Visual instruction tuning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Visual instruction tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.512973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:ad82b0d5afeb521869fb9282ad2861554fe3043300a42cfc7fbd8a4584281efb

Observation 3e6fb234-22c5-4528-ab7b-002c0213761a · outbound

This paper cites Language models are few-shot learners.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Language models are few-shot learners

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.551624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:6f4f0a69491b68c064aa950b9199506c56072fb41caa2321a561bed0dc9e59c6

Observation 1aad9eaf-692c-4b6c-9d23-69756211b4d7 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization LLaMA: Open and Efficient Foundation Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.208622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T11:08:05.851253+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:d4fafc9d48fbf9acebda2c9536d2895907976f552a37eb20db0b02dcffb48fb0

Observation e885d802-26b6-456e-bc1a-5c0a1868753f · outbound

This paper cites Training language models to follow instructions with human feedback.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Training language models to follow instructions with human feedback

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.458674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:0f7de90182a1846511a3194ce3207dc9cafc9ed4a39a27dba71a2b104715d83b

Observation 5a98bd98-33e3-4111-8a37-9cf928554a1f · outbound

This paper cites Improved baselines with visual instruction tuning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Improved baselines with visual instruction tuning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.528701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:2ce8cf202a13f190109827bdafe537de3bfce09809644832d74000cc94c46101

Observation a0cfa46d-cecd-4c59-b462-3ed711aa4c48 · outbound

This paper cites How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, page 220101.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, page 220101

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.548450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:313f0374e47780413045bd10c0c2366edf1daa12d239b601803d3c2a4249bc60

Observation f859988a-690c-41bb-97c1-c087a1441a09 · outbound

This paper cites Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.545536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:66cc3ac1e9337618900fecd7d7cf4f7027c06eb93eab7b5f29516483b08d3abc

Observation 9c2bf4f9-43de-4769-aff3-12657367f6ae · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.212455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:8b108dac1b78792c7c341c89e8f4e3023331613f0985a397e33f4cc8d1a7bbcd

Observation e937d26b-8bb0-4fb0-b115-b42e3001aba0 · outbound

This paper cites Qwen3-VL Technical Report.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Qwen3-VL Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.200508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:2ea4b11517da1ce2d3d1357a1cdb3a719af711a197cf28624348c2a1ad8b8527

Observation dc01909e-d858-4ee6-975c-c22717e2873b · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.462568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:8467272c14a1c943dd5be62c0283f93d2f3897658ba85ae18153f4cb036dbcb8

Observation dbb5e3f9-8286-4f06-a8c7-c91cfe99e2a7 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.204490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:356f1b0f49e5072aa22b585ad2d093e9ee22cd1f20325722e38ddf69818984e4

Observation ef8d8350-e08d-4672-ad60-e45d15abe2c8 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.062623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:ab814a05c2c919f2e812f35e85b30e0ef85c61e3b9dc67eb91099dd78cfc76d1

Observation c5a2c74d-25d1-474d-85cf-87a7b22718e4 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.076778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:ed146bba285fb8bc20daf12a0d1e991b223380edef0bcdbe9759eb3ada7c3fa3

Observation 749cd877-3525-4518-b60b-54876c5b3f08 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization LLaVA-OneVision: Easy Visual Task Transfer

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.192737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:7bbb75feee679a7e066fe803fd23fe7d9053809eaed2c8bb3377776d10b120f7

Observation d3cff9b5-13fd-4e3f-a775-4f383422ced1 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Multimodal Chain-of-Thought Reasoning in Language Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:12:27.664923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:0010f88856b9e5f6af53055a5b1b877ef091f89c99b9a6adf0a0985584960e7b

Observation eff8e9f7-7997-4900-91e8-1a4f2be97088 · outbound

This paper cites an unresolved cited work.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-05-17T05:14:05.531990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:243dad0e91ffc2cddf076da14dbff80e97ec7272510c9666e9e9fe0ccabc81e9

Observation b6ecd65e-7569-467c-8290-ccb324a1b228 · outbound

This paper cites Satori-r1: Incentivizing multimodal reasoning through explicit visual anchoring.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Satori-r1: Incentivizing multimodal reasoning through explicit visual anchoring

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.180651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:f15411f960096761bcbe009284c096b8aea8f6e57dacf93cc64be6585c7451f1

Observation fd9edc39-c727-4ea7-9be8-a6c37f628a02 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:30:15.750849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:df623ea36bd583db47469ffbc0b969527ea4bc301f58c12b2521d6b0c907f89e

Observation 745a9d40-ea66-4c4d-bbf3-3a16a50535e7 · outbound

This paper cites Toolformer: Language models can teach themselves to use tools.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Toolformer: Language models can teach themselves to use tools

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.476636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:ba2a171097c1813439c109bcff920e672824985c41c0fa81affc3e909444f47a

Observation 67f11fdf-69db-492a-b679-a7bfa042a3a2 · outbound

This paper cites React: Synergizing reasoning and acting in language models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization React: Synergizing reasoning and acting in language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.473160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:c5e9dd3adab0ed5d2a9a3f33cd9b353f4ce7c225a2d2639f9ae783fcc90cc9e2

Observation fd2f3094-b3b1-469d-aeba-98b6bcbed5c5 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-14T01:17:58.977376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:4dee2b346e8484a768666a021ead74ecb1317c31c3e1f2158c302b216d530c2c

Observation dec36dd2-bf1e-4ce1-9f0c-2b9e015ebdbf · outbound

This paper cites Visual programming: Compositional visual reasoning without train- ing.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Visual programming: Compositional visual reasoning without train- ing

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.483554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:6b53455f7007a5efc0b5aca1439958bc487f738e744ef7ec839c3fe911ec86a5

Observation 3a1d39f3-e894-4a73-ba7c-f4e92d44d3e4 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:21:45.673967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:b83fa1d16193e156223fe67a67fbc9572d8eff0bc4617fdb4105493392a09320

Observation 0732976c-d11f-4b3f-933b-3480377d370d · outbound

This paper cites Proximal Policy Optimization Algorithms.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Proximal Policy Optimization Algorithms

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.162903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:ff0c1925ac12cf582764f1b3c17a72ee266968e8406de7e30035cdd0d04b5a93

Observation 212f6652-4e15-4cbe-b609-52ebbbbe7b1a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.116532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:b9fa61cbe3feed5514c951a2c847e9ec35c6a5ae57e75e0110f551a49aee0cd8

Observation a90c4980-cc53-440d-a8c2-8112c29fea9a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.059446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:9a826a5e0031f6ffd8893c439c15ed83645e4fd0b775f2fd2140762712bf664c

Observation f020a5f1-104d-4d5c-b581-f4def7661036 · outbound

This paper cites Group Sequence Policy Optimization.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Group Sequence Policy Optimization

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.184268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:e1d848f9920a32debdc3fc911ed61bd6aab841d10f4ce6c0a6092b90cf9a99a1

Observation 4ee9354d-9266-4470-b354-2f84ff34aca4 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Kimi K2: Open Agentic Intelligence

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.069676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:0b4cc55f54aafb309c5674cd2c490a4af2c2b6f10dbf4c4699df0b10fcac9b23

Observation 7b074a6a-42d9-4f4d-b6d9-ebadc9e61e0c · outbound

This paper cites WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization WebWatcher: Breaking New Frontier of Vision-Language Deep Research Agent

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:56:24.054213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:6ba3029c6454939ac452e73ab0e6e65254082a20b9b395a1d316a7a281cdddc4

Observation 5cd4db6e-d670-4757-9327-b65a1ed2e3eb · outbound

This paper cites WebSailor: Navigating Super-human Reasoning for Web Agent.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization WebSailor: Navigating Super-human Reasoning for Web Agent

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:37:09.773663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:256676fd866a28fbf2b08a45a97d3f9166efa7b6c5b312900de6a47e0ba40ee1

Observation 13040d64-8ccd-45cf-b01d-be7d383320dc · outbound

This paper cites WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization WebShaper: Agentically Data Synthesizing via Information-Seeking Formalization

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.158468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:693166b1b14917c3af9b07549ce507b6273b9c999ab2e865d4ed6397be56422e

Observation fb5d1019-2f23-4e62-9b69-c82919ee8576 · outbound

This paper cites Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Agent RL Scaling Law: Agent RL with Spontaneous Code Execution for Mathematical Problem Solving

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:45:50.167268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:d655ac59726929c5b4fdca009fcbd4ef5c711930b3e0e75a62613dd2f3b5fc02

Observation 91ff23a9-3d9f-4891-9a74-03c8c64502b9 · outbound

This paper cites SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization SimpleTIR: End-to-End Reinforcement Learning for Multi-Turn Tool-Integrated Reasoning

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.171881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:010f37e2372169adf7a08de00c798e852d6d55ab794cb42d5d94e38a6bf0b8ed

Observation 83bf4041-421f-4799-826c-c324198a0115 · outbound

This paper cites Explorer: Scaling exploration-driven web trajectory synthesis for multimodal web agents.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Explorer: Scaling exploration-driven web trajectory synthesis for multimodal web agents

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.502562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:fa37d4d1dfed6eb8c9532e5f063400a3c15fc00213d57a5853e4094b1421b31a

Observation 943c6789-6844-4526-b83a-c2c61459f484 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:13:34.687091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:c85c2b09e3b96da85fd0f38bfce9f37963448db921240014b98dec6bdf6e7211

Observation 65d14b92-343e-4f1f-b38a-dfc12d969005 · outbound

This paper cites Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-13T08:34:23.540456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:5f40e67a9421a82a117dddb190d64e57059d6d1a7909fbd7be9b6606437deee8

Observation 3a010412-b2be-4ea9-be83-e9400d76d655 · outbound

This paper cites Latent sketchpad: Sketching visual thoughts to elicit multimodal reasoning in mllms.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Latent sketchpad: Sketching visual thoughts to elicit multimodal reasoning in mllms

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.145785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:e8f96906d9631a85a16edcc78b9a86596b0106dc565bfeec1866fd5c7fdeccea

Observation 17de2ce0-8681-41ed-b83c-c66a1fa2614b · outbound

This paper cites Monet: Reasoning in latent visual space beyond images and language.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Monet: Reasoning in latent visual space beyond images and language

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.149921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:946ea28fbc5bb921134eea7d2ed8a38ca600895c2dd298e10429fe0e71d2da5f

Observation e1aa2929-42ee-4b61-91c8-3aa65be2d3db · outbound

This paper cites Latent Visual Reasoning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Latent Visual Reasoning

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:41:30.500607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:3654cfb1a667f0fc0b289774a4a66214866fd4d91ca9aba785e2caf858b86636

Observation 598ea679-32ab-4f53-878f-369b9ebb2ee2 · outbound

This paper cites Interleaved latent visual reasoning with selective perceptual modeling.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Interleaved latent visual reasoning with selective perceptual modeling

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.087660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:860c6243441ad1558e7aa7677890d3204db9123b6fd3c0792fe8169670f7833f

Observation 31f3e513-db5e-4382-af03-15f82102ef10 · outbound

This paper cites DeepEyesV2: Toward Agentic Multimodal Model.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization DeepEyesV2: Toward Agentic Multimodal Model

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.694372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:2ceb9dc4f6890f051606132c7726d5196c092e56b5f312ed16c9c588dc04ff46

Observation 052aa34e-c3d5-4b46-8620-57bb85519c02 · outbound

This paper cites Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:35:13.359401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:9a90e79941c8a1377f3362b92271e678fadf4886deed263ff2a08a0729fe923b

Observation 4101407c-c167-430d-89bf-2f14a241def1 · outbound

This paper cites Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:22:27.090901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:ba09671e6becc3c6971894c3931fbe601bd2f9a223624dc76fd35c57526eee31

Observation f5beb4e0-9243-4940-a28a-9be871b815b3 · outbound

This paper cites High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization High-Resolution Visual Reasoning via Multi-Turn Grounding-Based Reinforcement Learning

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.112672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:50f38f6582f4916b9ed45c3fa89cfebfe652a64b312d2881d6d4d4046a9c01a7

Observation 98f7b2a9-196c-4e0a-bc8e-28aa5b145015 · outbound

This paper cites MMSearch-R1: Incentivizing LMMs to Search.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization MMSearch-R1: Incentivizing LMMs to Search

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:27:04.573030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:2969f54319f3c84e13fd1fa180f98348548193a8ed2c7e176ebd9a2c97995345

Observation f55e675c-9e83-4278-b74b-1bb613346095 · outbound

This paper cites VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization VisionThink: Smart and Efficient Vision Language Model via Reinforcement Learning

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.083817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:40879ffe5e94dbee78c548f0741ba7270f5fc7b6802bc6e5053c86b98f1075ef

Observation a3849417-c099-4ddc-a179-1b1a96fbc4f9 · outbound

This paper cites ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization ACTIVE-o3: Empowering MLLMs with Active Perception via Pure Reinforcement Learning

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-06-09T03:07:59.867366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:e67427a53d2a349e78b58a4d096fa3685ff8555c8427eacf5b05cb9ed933e19c

Observation 68a4a4b5-454f-45a1-b45a-c0602ff4824d · outbound

This paper cites Thinking with programming vision: Towards a unified view for thinking with images.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Thinking with programming vision: Towards a unified view for thinking with images

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.133336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:481f67c9350ca85df244e3b089c9d55920d0f0e58a9242852560d3e68303397b

Observation 0d0649de-7dc6-4641-9d38-0bd76b0a6402 · outbound

This paper cites Vacot: Rethinking visual data augmentation with vlms.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Vacot: Rethinking visual data augmentation with vlms

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:45:50.052410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:831fe288319cffa5880e1d3e1d36ad77115174defb9ec6a5f1311fa1b604cc3d

Observation 260d141c-5131-4670-8b54-78b2e1aa3899 · outbound

This paper cites arXiv preprint arXiv:2602.12916 , year=.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization arXiv preprint arXiv:2602.12916 , year=

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:45:50.141554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:b7c1d8b027f36891ce6ee9f7e5457870709c67bf4906030246f3156edeaf1f14

Observation 04c301e5-6b5e-4bec-b822-39af19c9d3ae · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Hybridflow: A flexible and efficient rlhf framework

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.466157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:5ce0e5c4a9d4e252546c9a1520b39b68e1d0922af8cbd2d32edb803330130869

Observation 605af6cb-b296-4c94-a8a3-a7e2f7df4831 · outbound

This paper cites Momentum-based variance reduction in non-convex sgd.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Momentum-based variance reduction in non-convex sgd

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.509515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:b8292370e253eed144dd0d893a9a96b4992c71ae3fc970d2483a40833dbfa43e

Observation 6ec6489e-66a8-4588-873a-de8b58b255bf · outbound

This paper cites Gpt-5 system card.https://cdn.openai.com/gpt-5-system-card.pdf.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Gpt-5 system card.https://cdn.openai.com/gpt-5-system-card.pdf

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.455333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:62ddc27e79c95dd75c1782d8d25df28413456ba577917c6c58f4ac4a5b987126

Observation 289660dd-1155-4bb1-9317-45204dd24262 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-11T00:45:50.137506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:63735ea278555e55a17f08737fc0760f050dd850f50c0368d13f318dcfcbb679

Observation 898d3559-3658-4270-b87a-7cb07bf1ace5 · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization V?: Guided visual search as a core mechanism in multimodal llms

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.490077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:c712e06b28859c06cc67e9e3d81917d8db6591a5155f3484b7d291d0d4bc3fdf

Observation a3d5a24d-02da-4bce-99e1-d61d13540fdb · outbound

This paper cites Divide, conquer and combine: a training-free framework for high-resolution image perception in multimodal large lan- guage models.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Divide, conquer and combine: a training-free framework for high-resolution image perception in multimodal large lan- guage models

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.493060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:342e9f39b6127df89ff9ea90e6ac37a04ba4d1115b46483cc2ae130c3ccad7cf

Observation c18c734a-b9b2-4098-9036-f621d47f9e38 · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:59:32.958879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:db56c6aa34895fd9351b9bc7f144d5ea5f5f3488a0cfdda2baba47173b85c29f

Observation cc9c9c16-8d6b-44c9-b715-aab5dc443c4b · outbound

This paper cites an unresolved cited work.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-17T05:14:05.495829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:d58d318d0526642396b15adc1f9eda519c5e419d418643feea1ca591f143f685

Observation 548b0822-a4b1-491f-8913-215fcdb13900 · outbound

This paper cites Varr(ˆgout|τ) =∥∇ θ logπ(τ)∥ 2 ·Var(r out) =C τ ·σ 2 out (21) Varr(ˆgsem|τ) =∥∇ θ logπ(τ)∥ 2 ·Var(r sem) =C τ ·σ 2 sem (22) whereC τ =∥∇ θ logπ(τ)∥ 2.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Varr(ˆgout|τ) =∥∇ θ logπ(τ)∥ 2 ·Var(r out) =C τ ·σ 2 out (21) Varr(ˆgsem|τ) =∥∇ θ logπ(τ)∥ 2 ·Var(r sem) =C τ ·σ 2 sem (22) whereC τ =∥∇ θ logπ(τ)∥ 2

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.524735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:f693eb8ad85af1b2ff47ce32bcf28f594438ce55d6994bbde29c1feee96d296c

Observation 1acc7250-554a-45a4-94eb-41be46b8fce1 · outbound

This paper cites an unresolved cited work.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-05-17T05:14:05.469357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:440b020ae72138bc796390581f0b2d0eb43a374a7d7b7af4ff21339153eaf90f

Observation abf79b2a-3732-49bc-abe1-de71f31464d2 · outbound

This paper cites an unresolved cited work.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-05-17T05:14:05.521174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:77ebfbe2d081bc123999a26e5db671e189bd0414fcf2ca09fa8b4404e70d1b69

Observation 57af2e20-13f1-428d-8ffc-563dc82ad2a8 · outbound

This paper cites type":"function.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization type":"function

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.480070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:2c1683b647e958ba6c8407a426e3d74307171658a36d4b4748cb35619994bd24

Observation 3cea1cd7-16df-420a-9025-49d5f5ac1de9 · outbound

This paper cites The input image resolution is dynamically handled, with a pixel constraint range of[10 6,2×10 6]pixels to support high-resolution visual reasoning.

Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization The input image resolution is dynamically handled, with a pixel constraint range of[10 6,2×10 6]pixels to support high-resolution visual reasoning

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T05:14:05.506196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-01T06:32:01.292127+00:00.

source=pdf_text observed=2026-05-10T18:20:02.559108Z digest=sha256:e701c29dc44fe8640d192ee2c2551061771ddb53a2455429685808edb9d8f147

Pith citing papers

Observation 9b906d49-9f1c-4e96-80b6-e2f12fa6e915 · inbound

See2Think: Do Multimodal Models Really Use Intermediate Visual States? cites this paper.

See2Think: Do Multimodal Models Really Use Intermediate Visual States? Walk the Talk: Bridging the Reasoning-Action Gap for Thinking with Images via Multimodal Agentic Policy Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-30T21:38:08.611674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:38:08.611674Z digest=sha256:99c03def1a3bcdcac74e21c7207cf9a847d40cb94bc438289e15819fe7f640e4