Pith. sign in

Paper Citation Record · LEDGER

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 100 of 125 outbound references and 8 inbound Pith citation observations for arXiv:2506.07905.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07905 v1

Coverage vector

measured 100 of 125 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:27:05.282142Z

measured 108 of 108 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:57:07.475851Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:39:46.104997Z

Reference resolution

100 of 125 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved99
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a2f85b3b-0818-41ce-9958-2cae1fc7f270 · outbound

This paper cites Qwen2.5-VL Technical Report.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.910573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.910573Z digest=sha256:a9a5c9f6d1a538f3d74d7c4f1be8617b14e1c0c94e8100194f2d6cdd22a13155

Observation 978c1cab-a1dd-4444-89c9-6ad7b5172600 · outbound

This paper cites Openai o3 and o4-mini system card, 2025.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Openai o3 and o4-mini system card, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.914834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.914834Z digest=sha256:618dc7c7d21217295acd7d77a9e7b659e18ab636d55e2934f0351d081b54f5c3

Observation 4843fe97-3092-45f2-bf8c-925fce292f34 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.918664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.918664Z digest=sha256:4d83955812363806c68a78e308517523482be6149eaddea1083a2c16cb55981d

Observation f9fbe446-e9d6-4e54-8d3b-2e4157656162 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.922660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.922660Z digest=sha256:b9774ae9d59eac5327f5ddc149b3842b9065ce2cb0a831ad71d7208077af0922

Observation 03394a63-8b2d-453b-831a-db6468728ed7 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.926476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.926476Z digest=sha256:951ff59f4313204ce2330126090728333ed8ed06c7ed34a7f5ce450ab1ba8152

Observation 3fb95a9b-1dc8-4365-8085-9f50c9b4469e · outbound

This paper cites s1: Simple test-time scaling.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning s1: Simple test-time scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.930282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.930282Z digest=sha256:17aedfae6eb3426ede78251d8ed148618e1d5fba6d2a7fce6e38ca75a67fdf18

Observation 822dc9e1-c613-4d34-ac15-2d56b77e2daa · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.934233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.934233Z digest=sha256:e2e7e5c0ff4f52c608c49b7a59bad85f9828d14190b17e6aedb98b768279ac26

Observation 495b0a75-d03b-463d-83cb-f010a3d40fbf · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.938064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.938064Z digest=sha256:0b4574c6bf4e7318293d40a78e2bfa26bb3f9af5dcf91f152930c0b6855b5b15

Observation 5894646a-2c5d-4728-b863-800dbed0db4a · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.941900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.941900Z digest=sha256:4d9a8ddec4bd639c36e57df200dbab3acccd9af3a3e02c16a83e1339bd58d922

Observation b32720cc-9242-4df5-9434-3ef7c2fcbb91 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.945637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.945637Z digest=sha256:3b5add27f20ea594775668b06c8d0ebd5ee07f6d8bebae3aa67bb69900285429

Observation b65f2783-cb81-4d0b-81c6-d4f0b5b9c89e · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.949338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.949338Z digest=sha256:c399c1a235c9deb583afd4e17c32f2a0be6dc2c3d0da147314ae8b1a20a1356e

Observation 52f93ccf-1a8f-4b7b-a910-5f065c564f82 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.953540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.953540Z digest=sha256:1264af9afc5f2cf78f8837dab764d01e978cd2bd9fa1510c1c81f8bbd76d0768

Observation 853b1ea0-8045-4c13-85fe-1bc958a89ac2 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.957024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.957024Z digest=sha256:e21ee053a2d72803941fbe6ca7379e54363b52e874eb76326a6853487ffec187

Observation c9527092-abfe-40c2-90fc-066cb510a9a0 · outbound

This paper cites Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Skywork R1V2: Multimodal Hybrid Reinforcement Learning for Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.960499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.960499Z digest=sha256:7c704f4da831151dde2a9055793f240538e94949e7a07407a9722379728cb668

Observation 2be5a1d3-a000-4c5f-9957-538d64e23047 · outbound

This paper cites Noisyrollout: Reinforcing visual reasoning with data augmentation.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Noisyrollout: Reinforcing visual reasoning with data augmentation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.964036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.964036Z digest=sha256:530fac80b523a10440534249591de2f1e775c1027396ee570d92c7ea72d0d605

Observation 931c915c-c4eb-46fc-a16f-3c8a5d0e56df · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.967572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.967572Z digest=sha256:d1fc41315314c96e41718ecf75d2b470403cb5199ff27a07807cd70b0fbd251e

Observation 4360f939-c564-4e2d-b0db-dd4d8b278137 · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.971025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.971025Z digest=sha256:64ff1051490d140dda5db9c4664ffd0fa495134027419f370f958ce54a355a75

Observation 734dfa59-790b-4812-9760-d728fd635084 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Wavlm: Large-scale self-supervised pre- training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.974789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.974789Z digest=sha256:d65833af6fa0b713e85565fb5dc2127289d8daf43dad545976f18afb6bf5b768

Observation 43752052-9380-4f73-bef0-41ec3dc11c36 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning VideoChat: Chat-Centric Video Understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.979169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.979169Z digest=sha256:11349bfc921aa987e82059f597a64ba756b738a7766187728f9f4d497e7796b5

Observation ad8d57d7-e276-4a9d-acab-25b25134dd35 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.982714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.982714Z digest=sha256:f70b0872e2b74d37c70d996936da8ee9f495423041f94e4763038084f0e8a9a4

Observation 53b78e31-cdda-46c9-8643-9501c5f25b5c · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.986725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.986725Z digest=sha256:c8cc7899a5741a3bd83e2cfe2a78fd9844305a581a5179660e741b95ffc8d464

Observation 0bb33cbc-892d-4c18-906c-d4aadfe706ef · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.990004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.990004Z digest=sha256:0c5c5977900f0bcc40e0aedd668b242520b5e605aad08a8123dec1d7d2cd3823

Observation 3691229f-741c-4818-8bb2-a29dd3b1b139 · outbound

This paper cites Cogvlm: Visual expert for pretrained language models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Cogvlm: Visual expert for pretrained language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.993533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.993533Z digest=sha256:e40a61644011860a658374408119198353858cf7a4038f96a6e4e79bd778301d

Observation cca83f3e-479e-4246-af88-9e0ac8d0177e · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Sharegpt4v: Improving large multi-modal models with better captions

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:04.997106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:04.997106Z digest=sha256:7a7203fa94d02ee2467777f86c860984f45b1d22de21eed5d31c1389e2dd91ee

Observation 773ba099-657e-40ee-b8de-4024deea19f1 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.000341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.000341Z digest=sha256:74606953cdf472ee60d46a3d2194138611931610dff90ca44bca1f15b14659a3

Observation f2523b7a-b2f7-45dc-9ef8-6112f043af83 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.004085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.004085Z digest=sha256:5868f95fc3b9ae11c74d7fd7f5c1e0c2a79f9e66407e7b9ef47653f68cf426e4

Observation ea755578-ed6e-48c1-a8fc-4b270399c075 · outbound

This paper cites Improved baselines with visual instruction tuning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Improved baselines with visual instruction tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.007607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.007607Z digest=sha256:257e6eecbaa26c1d98674f03262b943c2b07d63b8b9b0be2abac00800ac945b8

Observation c6e4bad6-63ba-4c71-bc04-1cd724b58701 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning LLaVA-OneVision: Easy Visual Task Transfer

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.011264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.011264Z digest=sha256:8bc7f1763a1b699352d1c3337ed8b82da1b0beb90dc044e6a624ed98bb0450be

Observation 85856486-bfae-44a9-a157-c0638ef44a41 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.014866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.014866Z digest=sha256:e2e051a41601760fa45844e866b6ac2f18a608707ab64f6e7572d38307ee9dbf

Observation 866e57e5-febe-4c62-a48f-569875f5fdc9 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.018179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.018179Z digest=sha256:c3aee237e97fe32443097d894d2ce37a0419d9d6ccf0fa63d2428e2a1123b9bd

Observation 4d41cacd-b65d-4fad-ad0e-98d679aad327 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.021821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.021821Z digest=sha256:8130537997b6431e132855b2516ccb32d10af6e42117a2db85fb832512bb646d

Observation 743ef82d-e42c-42e6-9bdd-0f9672ef23e6 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.025413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.025413Z digest=sha256:951d69cf88e310d5a639d4e45d7db7c073a012521fe1f7f7e5c2e30e7540e623

Observation c9f02f92-f888-404a-bb3a-02e0f1357460 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.028712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.028712Z digest=sha256:afc7828648ddc7fd241d65b65c39e2e04a90b9700fe51cc82ac3868569856bd3

Observation a55093e7-af8c-4648-80c8-294a51b3d510 · outbound

This paper cites GPT-4o System Card.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning GPT-4o System Card

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.032184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.032184Z digest=sha256:515eb336ba2030a24c9d7ae8ede4ae0f54f7a70e9ee665fe644a73786ddfb527

Observation c1f62a4c-2475-4afe-8820-7f6546cb2512 · outbound

This paper cites Claude.https://www.anthropic.com/.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Claude.https://www.anthropic.com/

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.035668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.035668Z digest=sha256:0bfac1a1d535438db83160bfec659c1f9ed3feb6f8f9850c694ce2c16e39cc79

Observation a1f81c7e-41dc-4c9f-a147-ca620750ba25 · outbound

This paper cites Grok.https://x.ai/.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Grok.https://x.ai/

Reference 36

Resolution
parse uncertain
no resolver link, observed 2026-08-07T05:27:05.039295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.039295Z digest=sha256:f1c2f11e22815757f361de7f6724a8e9bfb238ca6e009273d447f8af1ab6d459

Observation 580d4b2b-7a81-4cf8-a2d8-cf1b67f926ee · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Chain-of-thought prompting elicits reasoning in large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.042756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.042756Z digest=sha256:254b72a1c6a22424189b83ec83d4f39d8b745dbb5bb674b92cddb5d3b26b8163

Observation 217880b1-42c6-488a-b0f8-96e668dc1cee · outbound

This paper cites Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Large language models are zero-shot reasoners.Advances in neural information processing systems, 35:22199–22213, 2022

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.046056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.046056Z digest=sha256:e917cc6287dd1e6fa45d7a68be19c3ff29d961ab454d39f1517ea2867bf8e091

Observation 373b2e0a-d748-48a8-967e-69b399de2ec1 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.050109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.050109Z digest=sha256:1eb3247a8558f70dc699b05c5f01e1af94276d53a9b55ad89dc8e41ee56f12a2

Observation f00f1f6c-04fc-4e3f-8008-a8d0088b3fbf · outbound

This paper cites Chameleon: Plug-and-play compositional reasoning with large language models.Advances in Neural Information Processing Systems, 36:43447–43478, 2023.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Chameleon: Plug-and-play compositional reasoning with large language models.Advances in Neural Information Processing Systems, 36:43447–43478, 2023

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.053861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.053861Z digest=sha256:a190eb4dffdb70ca22349a54b764cc94341ed62274f99c7d903d23c32c0d03f1

Observation eed55363-e79c-4b63-834e-d9e1eb30dab3 · outbound

This paper cites Layoutllm: Layout instruction tuning with large language models for document understanding.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Layoutllm: Layout instruction tuning with large language models for document understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.057374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.057374Z digest=sha256:f6926e85da5015d43147afcdea6b593500c1ab2206c5bd7be81ef92fd843d006

Observation 2b587f56-f0a4-4f1f-a570-c273cdc5878a · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.060976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.060976Z digest=sha256:4c62d7cad5e2926098cc37904f3a7d1a1c1484401e801621ca30a48343f72841

Observation 5793574b-1f18-435d-8eb9-a6f95939c91f · outbound

This paper cites CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning CoCoT: Contrastive Chain-of-Thought Prompting for Large Multimodal Models with Multiple Image Inputs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.064735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.064735Z digest=sha256:807cb98b1aede5f566fedea7f7a8c2d6d406750aca07299b5a41396f180ba4b2

Observation 25e96661-8911-472f-b823-13befa0f4b47 · outbound

This paper cites Compositional chain-of- thought prompting for large multimodal models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Compositional chain-of- thought prompting for large multimodal models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.068394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.068394Z digest=sha256:c1db72fde347410794132c6d455b43165c2ea70d7bc581a357f4b1c618fc102a

Observation 4093cbf7-d757-4e22-86d6-75ffac09711c · outbound

This paper cites Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neural Information Processing Systems, 36:5168–5191, 2023.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neural Information Processing Systems, 36:5168–5191, 2023

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.072062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.072062Z digest=sha256:9f50ced6dc174236f88899834e088c605bc35d05790729983569ffb20ca19b8d

Observation c88c960c-7624-418f-870a-1bde9f3e909d · outbound

This paper cites Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual-O1: Understanding Ambiguous Instructions via Multi-modal Multi-turn Chain-of-thoughts Reasoning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.075376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.075376Z digest=sha256:e00f529a2cc54b48651a13662a39fa2963b17841eaa8b71d26ec75edec288167

Observation 722c6d23-e56b-4f1e-8965-a73b0ea02771 · outbound

This paper cites Visual chain-of-thought prompting for knowledge-based visual reasoning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual chain-of-thought prompting for knowledge-based visual reasoning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.078937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.078937Z digest=sha256:c5e9ad62bc803acdaf46bb3609f197ef58ebe5a8f3cf0b739a37388f5f10c084

Observation 8b2136dc-12f3-45aa-8bda-ac0ed1d08a1c · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.083376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.083376Z digest=sha256:b1fc6651d4689771074bdc01e3c835e02c0296aec5b16dab8e7e37af2959bef8

Observation b21cc11a-7cb7-4090-b1d7-b43f9e8a02c6 · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.088125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.088125Z digest=sha256:68dbb468ad8a2e066892536c3c2c20cc8f42a4b34853df29a6ad73923b2de2ec

Observation 38fca845-86fc-4356-abf4-22a29e52629f · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.091700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.091700Z digest=sha256:c2c45d704e8ba4c22e0d77f6524f903a5cac37e9f90697c86897acde275bdc37

Observation c4e8ca4e-e792-4925-8635-57e9b487c9aa · outbound

This paper cites Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Reinforced MLLM: A Survey on RL-Based Reasoning in Multimodal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.095888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.095888Z digest=sha256:cd7f782d93e67e39bc528d8c5988fe9f28c8268b00521002bddfa28ae06124f7

Observation d226f46e-6381-44ed-a5dd-84a08c35a2d9 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Training language models to follow instructions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.099935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.099935Z digest=sha256:aa3784bae98b954d5367a1ab52e60217f6d61e479bdb8ef46a8c9b458f745f40

Observation f129975a-ae54-4a41-a608-82aca32d2b88 · outbound

This paper cites Relation-r1: Cognitive chain-of-thought guided reinforcement learning for unified relational comprehension.arXiv preprint arXiv:2504.14642, 2025.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Relation-r1: Cognitive chain-of-thought guided reinforcement learning for unified relational comprehension.arXiv preprint arXiv:2504.14642, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.103227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.103227Z digest=sha256:2fe1147228fa644a5acd6ba5db4596f8fa34f7bda66c021a34ec9643d47596a9

Observation aa3c6233-1240-410d-b735-770dffa707f5 · outbound

This paper cites Compile Scene Graphs with Reinforcement Learning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Compile Scene Graphs with Reinforcement Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.106433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.106433Z digest=sha256:6d4b9183d77f4aff0d72a0535cd5474620bf022374f2f89f98849a24b57e5b4a

Observation fb0606e7-931c-4f6b-a038-f2783832d086 · outbound

This paper cites Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Embodied-R: Collaborative Framework for Activating Embodied Spatial Reasoning in Foundation Models via Reinforcement Learning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.110383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.110383Z digest=sha256:493cfe1be11cd1e0d128906477807f2b00681f67da1ce0a9563d2773baec47fc

Observation 900592f3-cb75-485e-8f6d-57cc9329113b · outbound

This paper cites Improved Visual-Spatial Reasoning via R1-Zero-Like Training.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Improved Visual-Spatial Reasoning via R1-Zero-Like Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.114508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.114508Z digest=sha256:c2484c22c37e6edd6caf91cbac4fd237731cc6fd1f903bf9b81ab6972ee55300

Observation d4216eaa-f78b-4267-8cf5-86dd6946c4ec · outbound

This paper cites Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.118106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.118106Z digest=sha256:19915a3cd060e69a83bb4307b0665a79d31b994e02b28e06b507498b15f1b60f

Observation 103c64de-bd83-4be8-9b3f-cd361c7ee95f · outbound

This paper cites R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning R1-Zero's "Aha Moment" in Visual Reasoning on a 2B Non-SFT Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.122122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.122122Z digest=sha256:6949a53c00ee34dcee6016b2f1d061483df0668b7d5ebeb37f2a5b2d9e251c85

Observation d120e7ae-903b-489a-aa89-14f38d5d8a2f · outbound

This paper cites Perception-R1: Pioneering Perception Policy with Reinforcement Learning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Perception-R1: Pioneering Perception Policy with Reinforcement Learning

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.125750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.125750Z digest=sha256:fd8351cfdc396eb4da16b1fb2d62105d382228b8cae1f8eed7604d6258c014c3

Observation 37fd741a-0b79-46df-9e00-a1c67bf22259 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.129486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.129486Z digest=sha256:d327a0a2d03a411da583c19f4c4fbac14c85e3c2bc94d7f63750dd162b10c296

Observation 7679a474-d1ab-4409-abd9-9c4aaa7000b6 · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.133072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.133072Z digest=sha256:4d9d43b7e662416c9f5a9841762845e91fdb1161f4263c55864ba6049b34798c

Observation 6348801c-8c22-4cb9-8943-fd8dbb1ab974 · outbound

This paper cites Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Boosting the Generalization and Reasoning of Vision Language Models with Curriculum Reinforcement Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.136689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.136689Z digest=sha256:78a673f6b38e65469c2858362665d1499e684a49c654a042415986ad0a726d3b

Observation 22bc375c-4e64-45da-a171-f57595902bb2 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.140729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.140729Z digest=sha256:803eda5d8562cb8c46422a7fb0be62366674c64e93c4de3604c3e22df682884e

Observation d157d2d5-3671-4249-8e5b-5e86a6654f91 · outbound

This paper cites Crowdvlm-r1: Expanding r1 ability to vision language model for crowd counting using fuzzy group relative policy reward.arXiv preprint arXiv:2504.03724, 2025.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Crowdvlm-r1: Expanding r1 ability to vision language model for crowd counting using fuzzy group relative policy reward.arXiv preprint arXiv:2504.03724, 2025

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.145217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.145217Z digest=sha256:ce211e868565129c86cd2572174254dd3e2438cfc7e50781041b3e9f3ec5b349

Observation 4217f07e-78ef-41aa-84e8-8aa534724142 · outbound

This paper cites Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Reason-rft: Reinforcement fine-tuning for visual reasoning.arXiv preprint arXiv:2503.20752, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.148735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.148735Z digest=sha256:066448e5bf77eee72d3bc0f8f04cf2b19ff18286551c7e871994e10aa046f910

Observation 5fc15cf8-7aa9-4a9e-8c30-3b2432984ada · outbound

This paper cites OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning OThink-MR1: Stimulating multimodal generalized reasoning capabilities via dynamic reinforcement learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.152302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.152302Z digest=sha256:a219718d10096f93a99e5df445db5b7834483fd7d86711097332e39aa34af372

Observation c581d90e-8ee5-493f-9b53-e0e95a7b9f11 · outbound

This paper cites Microsoft coco: Common objects in context.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Microsoft coco: Common objects in context

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.156233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.156233Z digest=sha256:070e1c17eda622b3fcfe3f1070babf7d9e752eec50e62e40db54ec47869010b4

Observation 2d04fc18-ed40-4c4a-a6e7-95235979c46f · outbound

This paper cites Segment anything.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Segment anything

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.160052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.160052Z digest=sha256:6d65b753f3e05eb8a68fa1c2d242078c69a79efa8a75c9da039a70633dab61c2

Observation 89861dc0-b9bf-4763-a63b-48c10926ab6c · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.164402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.164402Z digest=sha256:998797b6ff0cd03c09bdd4f31f02b6d69774a267f4a9e34f14ccc4fea679f6f3

Observation e381a6a3-9853-40d7-8bf7-9d101078f2ce · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.168039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.168039Z digest=sha256:4b6ede09b2c4b85d39ec002b305147ee97f22347461a39144699877a0789e4ab

Observation 8f3867a4-b03c-4462-b451-93b616482945 · outbound

This paper cites Dual-glance model for deciphering social relationships.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Dual-glance model for deciphering social relationships

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.171480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.171480Z digest=sha256:5ce9ecd588bcf4c14e99365ec4dc062084d445dc6c74b86dc3547497a92542f7

Observation ca7b2491-d32b-4813-98bb-6b7c5dc20abb · outbound

This paper cites Towards vqa models that can read.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Towards vqa models that can read

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.175095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.175095Z digest=sha256:5f6e8c70ff3fb7f07bce65d67ebe2ed1ba4b22fd75b171b2d091c19ae63800ba

Observation 291d4e8a-b359-415c-b459-b6c1c0a85b02 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Docvqa: A dataset for vqa on document images

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.178737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.178737Z digest=sha256:fe4150ef605a6d9ba18a9dff8cf6dc7fd3fda0e48b45c92d1396f42bc28d6444

Observation a10dcfd0-96a0-4c4e-8c9c-f3242c3acade · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Ocr-vqa: Visual question answering by reading text in images

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.181987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.181987Z digest=sha256:f9bc360e355fa3d7ba54772b5856a8b2cddbcdec6d0dfdd042995b1cdc08e0fc

Observation 068b58b0-13bb-4fbd-9d9e-84d6034677ae · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.185608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.185608Z digest=sha256:7aa5b30b09fd507791ef236a5d48984b1399f8d3c97dce4bbea68fbd2c580b89

Observation b9aa5595-ee68-44df-bb08-01cfb3932d8d · outbound

This paper cites An augmented benchmark dataset for geometric question answering through dual parallel text encoding.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning An augmented benchmark dataset for geometric question answering through dual parallel text encoding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.189776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.189776Z digest=sha256:e324cc6a5c6811f127a37acae55a1f90512894da20e9ff7813e598407098cf9f

Observation 7d62ce74-5137-4911-8044-048da7265ebd · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.193367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.193367Z digest=sha256:91ddb9dd96bec50458ee1a11bfb22bfeb037a732fae5d54e9c60b8e626a1f95e

Observation b06efcd8-19f4-4144-843f-9821c78a36cc · outbound

This paper cites A diagram is worth a dozen images.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning A diagram is worth a dozen images

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.196950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.196950Z digest=sha256:b19b013455b1e32ef6d8813c8caf36e8634a33a2f95d5708caf15b2729c6ba8f

Observation eabeb259-7914-4223-8ee5-9a69dae88dbc · outbound

This paper cites CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning CLEVR-Math: A Dataset for Compositional Language, Visual and Mathematical Reasoning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.200325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.200325Z digest=sha256:4ac7fa0be1b5c9499a9d0fb09ee256b204930dc688243872fec075a6573da296

Observation 6cca9462-d10f-4096-afbf-ed6e80f2ff0e · outbound

This paper cites Large-scale Classification of Fine-Art Paintings: Learning The Right Metric on The Right Feature.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Large-scale Classification of Fine-Art Paintings: Learning The Right Metric on The Right Feature

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.204087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.204087Z digest=sha256:6810f9fd8bba00b8cd6c65a014864d9f7c78223ab320b1fd2dcd92ef2592bd79

Observation de39bb08-53f5-4639-9513-e13008944536 · outbound

This paper cites Quality at a glance: An audit of web-crawled multilingual datasets.Transactions of the Association for Computational Linguistics, 10:50–72, 2022.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Quality at a glance: An audit of web-crawled multilingual datasets.Transactions of the Association for Computational Linguistics, 10:50–72, 2022

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.207819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.207819Z digest=sha256:d338efddddfaf99e254a973b942e32304197b01b6cbd240f7c0a00351ee4da8b

Observation 453ec9e4-fd2a-452c-b484-b24161f7711f · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.211152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.211152Z digest=sha256:3a735050f3cfc798418b79823ba27de9cbeb0f0807c0ef2bbbc94e039cc5ecc5

Observation 2217614f-6fb3-46a2-bb53-017653874857 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.214816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.214816Z digest=sha256:e40c5279cccf49304f13d302080442053d5da288ffe33732bda7c5b69eb245d4

Observation 31eea80b-7b86-4b09-a24c-f21b2b5d36fa · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Measuring multimodal mathematical reasoning with math-vision dataset

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.218695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.218695Z digest=sha256:1e0eef4c3acfa6f51a9f84f52c39c2e78d15f1452a454b58a50e2e31df18b6c0

Observation dc998b26-193b-44a7-8a6f-d61669a8b7a9 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InEuropean Conference on Computer Vision, pages 169–186.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InEuropean Conference on Computer Vision, pages 169–186

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.221981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.221981Z digest=sha256:47a7bbefb63c43a30978c1ca6b7aa85bcbb094bf658176b0f74a5c42ac55a7c4

Observation a5bc9403-a0d6-4efd-b92f-bc0e7a252ec1 · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.225253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.225253Z digest=sha256:f9b4a7f34265265d7c3a1b3353edaec6cb5e16c747b093ebdc006fba94aabe34

Observation d11d3584-06d5-4374-867b-ff60a6a0b96f · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.229610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.229610Z digest=sha256:b76cafcdfc09f8e36433b74aff24616ae5d078154bcd5d145212a0b77f7703c0

Observation 2583034b-a82f-497d-8e02-8cfb62091901 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.233360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.233360Z digest=sha256:81866be89ab5f592ee7862781c0d17e6a522c1fc4cf14f6aaf28ff8f054805dc

Observation 7e6b901e-2631-4aa3-8984-2716b6426aa0 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.237152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.237152Z digest=sha256:cf39751f486ff40faa24621cc3ea7bf68c25f98bdb4303626e0bb44f29bf17df

Observation 6017241e-0ff7-461c-8c6f-b5ba0517dc04 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.240629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.240629Z digest=sha256:36c13dcf02b01a02ba223734502d283e3d55f21958d465afd5c58d1ebd5459e8

Observation c3dbcfa2-99e2-44f3-b95c-5f8a0ae48840 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.244024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.244024Z digest=sha256:9c8eeb732ab60583cb479f1dd11dda03f49ad0dcdc123335f0cb9143a9e8d5f8

Observation 98464604-acdf-499c-acf3-4e38f8e7e286 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.248472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.248472Z digest=sha256:db14342afc068c0481620613c83061df6629352749a7466f9031c89e0aa02682

Observation f24ffa82-7d1f-4a5a-818f-553dd3864afa · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.253190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.253190Z digest=sha256:0cdc616ee33bbed8afe7fe823409ea4011e22029bd5a49f1361cc3f33bd300ef

Observation 00d2368b-9ee9-4596-bafa-2c10545d8f6f · outbound

This paper cites Grok-1.5 vision preview: Connecting the digital and physical worlds with our first multimodal model.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Grok-1.5 vision preview: Connecting the digital and physical worlds with our first multimodal model

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.256916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.256916Z digest=sha256:9b1fd08928254f755eafd7d75f1017b9de4d625b69a60cd8a7db6fa58094cda1

Observation fed61cc2-61ea-49e4-bd12-eddbad3687f9 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.261384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.261384Z digest=sha256:3cc014aa4fe37e64929492e91e3261d4bb5fc18b041b142655b435500e3a132a

Observation 2c09f3aa-a7bc-4afb-93a0-874cd4d96168 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.265419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.265419Z digest=sha256:aeaa06c77acebc40cd4c4c8cdc0650b63ba9fe2430d56e52d34441714837a7cb

Observation 0738cffd-d5a9-412b-a106-c939684827a1 · outbound

This paper cites DeepSeek-V3 Technical Report.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning DeepSeek-V3 Technical Report

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.269421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.269421Z digest=sha256:ce6ddc4beef8de955a821a44899d5da056c3900630749437ea0134153ceaf513

Observation dcc8fc6f-9d29-414e-90d5-719c3458ef99 · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Easyr1: An efficient, scalable, multi-modality rl training framework

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.273613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.273613Z digest=sha256:f5b59b02cf80c0c5f6b98e685ec6de2267fda774cfa7f3e2d257582d214ed235

Observation 40909008-8a02-4e26-98bd-9a1ed8bbe2eb · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.277827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.277827Z digest=sha256:79835be903585b41dbf721f3a041e7de8d67839035a7aba7436451207b0684af

Observation 522bca22-8ced-4f8f-b805-8fd38471bdb4 · outbound

This paper cites an unresolved cited work.

WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning Unresolved cited work

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:05.282142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:05.282142Z digest=sha256:02a736a156374e55e6386fb5be472c5894db63ff808fbe49337cac10b35822eb

Pith citing papers

Observation db245629-0bcf-49c4-9b7c-5aa648952409 · inbound

Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention cites this paper.

Interpretable Open-Vocabulary Referring Object Detection with Reverse Contrast Attention WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:57:07.475851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:57:07.475851Z digest=sha256:261515cdf36aa4b5751623f2c4565aa33cac62c8ac52b3ddb64be31ebe7d36a6

Observation 81f9e82d-4567-4976-9de0-b0ba9a4a3bb7 · inbound

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models cites this paper.

MathReal: We Keep It Real! A Real Scene Benchmark for Evaluating Math Reasoning in Multimodal Large Language Models WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T23:03:10.158919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:03:10.158919Z digest=sha256:0a927fc68c50c863b747d3706ad9c269015ed31ef7a52493826cc51ac34101b9

Observation 48e25eed-7b27-47cf-834e-9b44876c6c93 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 238

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.853224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:7f6066777184a0f12b7552987ca35bab856c93c88281bc7fa6db1c6c6d0cc288

Observation 0664b863-da57-4521-b0a4-9e6706ad160d · inbound

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models cites this paper.

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T14:57:29.176406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:57:29.176406Z digest=sha256:18b78d968402bb3316602645aed828b8c3f5acdd2108d1e856a19726a4734bdd

Observation c6b81be2-5b5c-4e7c-a979-e9098ce6c857 · inbound

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents cites this paper.

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:07.983014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T03:21:30.732925Z digest=sha256:4adec6b373b4a0715d7e7565319aca6abe5d72059a85b6eb409b57f91b9fbfcc

Observation 6983eb48-fcc7-4ed4-9e77-1de6f3922bfd · inbound

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models cites this paper.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.563530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:c612ac37cb7f583fdc0941d3699c5bddb199e0a97dd648d20c9dde6a95587d7a

Observation 09b8ffd8-cfb9-4af7-b9d0-3ed26b71c759 · inbound

VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct cites this paper.

VeriEvol: Scaling Multimodal Mathematical Reasoning via Verifiable Evol-Instruct WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:39:46.106878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T08:34:27.719022Z digest=sha256:6160dec4be462d8df7ecb05ee044ad83e9cf87c554797388da7ea1f662e8eed3

Observation a38f450c-7479-419e-a93a-2d14d92ed709 · inbound

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning cites this paper.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:28.164827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:28.164827Z digest=sha256:de820c71feb89363dce94b631f8f0efd197d1003ebdd0b55268075b9001d768a