Pith. sign in

Paper Citation Record · LEDGER

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

As of 7 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 6 inbound Pith citation observations for arXiv:2506.01078.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01078 v1

Coverage vector

measured 84 of 84 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:55:15.498636Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:37:19.575834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.331710Z

Reference resolution

84 of 84 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved62
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa0bce29-ef9c-4bd5-901f-3e684ab67a49 · outbound

This paper cites GPT-4 Technical Report.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.774631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.774631Z digest=sha256:deed101464a238003322c22122a4265463710d5c303d794255b1e506fccc70d3

Observation 12445f70-6053-460c-9e82-f2362571e48e · outbound

This paper cites Qwen2.5-VL Technical Report.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.821566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.821566Z digest=sha256:a18b273c054dc3fcc005ecb5bfb1729df9b4587d51783e024987cdbacc9cb91b

Observation 38f64302-aa0c-470f-99a1-8791e16daa41 · outbound

This paper cites Perception Tokens Enhance Visual Reasoning in Multimodal Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Perception Tokens Enhance Visual Reasoning in Multimodal Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.856611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.856611Z digest=sha256:4a52deebca4b78565ab9ad951ba389cad886fe7ce1b42dc296686264fd353197

Observation 9ecec8f3-af2c-44e1-964e-f995accd4dbc · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.906301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.906301Z digest=sha256:5196687d55709b01441f70bca7a5bd360ddca79c03b571146e660a4c7d06e2a3

Observation 023d15c5-a23f-41eb-9564-d013bc58a441 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? InThe Thirty-eighth Annual Conference on Neural Information Processing Systems.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Are we on the right way for evaluating large vision-language models? InThe Thirty-eighth Annual Conference on Neural Information Processing Systems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.930951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.930951Z digest=sha256:50338dc0707e012b084651197eb5ec9f6083db41fb67b831564149931a045643

Observation ebac17ce-1efc-4864-8848-c04bf13502bd · outbound

This paper cites R1-v: Reinforcing super generalization ability in vision-language models with less than $3.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking R1-v: Reinforcing super generalization ability in vision-language models with less than $3

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:11.969152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:11.969152Z digest=sha256:64939d82ca8c3c41e8629c61b08ff8e23d40d012d1f9e9f770e3cf7acfaf4f3a

Observation c9ff8af9-fef5-462a-8545-4a856cbd90e8 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.014182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.014182Z digest=sha256:1b009b789813c0a2a002386ee25d66216d3449c7de662507626e2f59efafdb61

Observation b0a05cc2-2c03-4da4-bf31-5f88bf5219d4 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.061180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.061180Z digest=sha256:22090cdc92e7af201fd314faa102763385d77a56be38f81c35554e7f097ed11e

Observation 4d17a14d-d0f8-4327-9080-26d7492ad706 · outbound

This paper cites Gemini 2.5 pro preview model card.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Gemini 2.5 pro preview model card

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.108675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.108675Z digest=sha256:9b9226dc926c639605e95765f473027bcf453f82cb0116dc2b063917cb6b66a3

Observation f1e0314c-df3b-4a52-8752-6ab3de98c9fb · outbound

This paper cites Virgo: A Preliminary Exploration on Reproducing o1-like MLLM.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Virgo: A Preliminary Exploration on Reproducing o1-like MLLM

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.149124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.149124Z digest=sha256:af7b289064e26bfae551db867a61244c20deef073cb852eb884429322d44083d

Observation d879ea55-65b0-4411-9b04-86a11e4af277 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.213124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.213124Z digest=sha256:fc75481fde18ec61dcc7c152fcaf542edd1f35621fab3939c6830a0802ef25ea

Observation c25cc8e7-7529-47dc-9568-0ef4fd50a058 · outbound

This paper cites Cantor: Inspiring multimodal chain-of-thought of mllm.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Cantor: Inspiring multimodal chain-of-thought of mllm

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.465045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:12.262102Z digest=sha256:b298bc0f79b579db6fcc76e0cc7b388b717570bbde5b6544e5ad352dc268cec3

Observation f24ee158-fd32-435c-8c58-fa50cdb828ab · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.301566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.301566Z digest=sha256:60b77b855226813c9c4b54dd8225f50e88156b248ea23d3bc20d9301da4fb149

Observation 5021add6-3abe-4e19-9a91-c1f21a66fc32 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Measuring mathematical problem solving with the math dataset

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.290620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:12.345211Z digest=sha256:7afc44e6bd606efe37bb07277715bcbd7bc9002ef9d45e32320358f55676ec44

Observation 7efaa70e-e358-41cc-8c91-60c873221cd6 · outbound

This paper cites The abduction of sherlock holmes: A dataset for visual abductive reasoning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking The abduction of sherlock holmes: A dataset for visual abductive reasoning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.168420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:12.398632Z digest=sha256:7de26571ec6653304b57cd9cb574a0a893b1bba054d47cc59c218f5b06518dad

Observation f35b0048-75c2-496e-b003-074ef2392571 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.450962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.450962Z digest=sha256:d5112663bb3325c4a724c3f66c45b84962e3386d0d3c139cb81b503dfd43ac4e

Observation e41420bb-eee7-4219-a427-1d2970955726 · outbound

This paper cites GPT-4o System Card.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking GPT-4o System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.506068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.506068Z digest=sha256:39dac0128c9d887b6c5bc9db4ece61b8bb4e1e32300195c8502d18a73998a5ad

Observation 11301c7e-0e1d-476e-9355-b371fd77dd1b · outbound

This paper cites Math-Verify: Math Verification Library, 2023.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Math-Verify: Math Verification Library, 2023

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:19.006825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:12.549003Z digest=sha256:829dcf239e42da7528a020042da62bdd313f7ced2e7f6515667c46489b0519d3

Observation 7e1d0daf-db50-469e-be9e-1c0f31cb1850 · outbound

This paper cites OpenAI o1 System Card.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking OpenAI o1 System Card

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.601005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.601005Z digest=sha256:2a13c0f6ff0891cea6794fad2593346ec86f0faabfd60f4873220e51f48701bd

Observation 46c18466-ed5b-4404-b7fc-2097d1c65101 · outbound

This paper cites Abstract visual reasoning with tangram shapes.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Abstract visual reasoning with tangram shapes

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.814552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:12.636382Z digest=sha256:d68d7c061bb2fe067c7fc8e0ab8e857552b8d0dfe0a7037732a53be5d9327cdc

Observation e0ac755b-9f00-462d-ab09-3b38b42ce4ec · outbound

This paper cites Dcot: Dual chain-of-thought prompting for large multimodal models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Dcot: Dual chain-of-thought prompting for large multimodal models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.642575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:12.689346Z digest=sha256:5534406cfd13da3b2945388d8b4e4845c071dfbcb760f43e822711e9e091fb7f

Observation 3b2c1749-b529-46e6-b946-27af2bc38d7e · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking LLaVA-OneVision: Easy Visual Task Transfer

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.734446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.734446Z digest=sha256:719520efa147a916786ba89cc81ebe2a53ce04c5a0117a2541b007820e85e0b7

Observation 0eb33f72-d27a-4a76-930a-3606de24e0fc · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Silkie: Preference Distillation for Large Visual Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.792022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.792022Z digest=sha256:46951ffcc849115989511667b9a2640b3b31688d588bdfa99ef0cd0f210126cc

Observation 62f27d10-92d2-4000-9c11-92909d519264 · outbound

This paper cites VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking VoCoT: Unleashing Visually Grounded Multi-Step Reasoning in Large Multi-Modal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.887272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.887272Z digest=sha256:8af408a8690a310649f4bfa16d6e22aa5b385ac28265ba86a281524bc9677908

Observation 1208e7d4-dfad-4aa9-ba0d-592b73ae6df9 · outbound

This paper cites Diving into Self-Evolving Training for Multimodal Reasoning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Diving into Self-Evolving Training for Multimodal Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.938364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.938364Z digest=sha256:27ccbea0e5b62f78f7979f791f5d3a34a8939d06c17ec23c6148816d7f392017

Observation 3194040f-a74e-4226-818d-3234125cc836 · outbound

This paper cites Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Chain-of-Spot: Interactive Reasoning Improves Large Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.978031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:12.978031Z digest=sha256:3505511a136a0259af27bb66235b1a49f25b17a9abeb5f48dcdcf34b60d164a7

Observation d0bb0f6b-b5c6-47ba-b795-aefcf70039dd · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.510804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:13.031238Z digest=sha256:0aa4141dd820827ad5f46248a7270998aaaa2813c2636b3d15f29dd1face9d25

Observation e43b2e78-e870-4fba-83c5-a9fc52b45e50 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Learn to explain: Multimodal reasoning via thought chains for science question answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.090908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.090908Z digest=sha256:d74bd06deb48fdf01468116a2e44720ddad211aff23f73c604a8b9ce7cef45ac

Observation baf9e354-97f7-452c-aeb2-9d34a4aef0c0 · outbound

This paper cites Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Dynamic prompt learning via policy gradient for semi-structured mathematical reasoning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.438773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:13.159210Z digest=sha256:05056a57e4fac9b4dd3d4446250763b6db8adfc586ae7892d45368d5a1aeeb50

Observation f3ea50c4-9b66-4532-954b-befd96fe0118 · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.202412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.202412Z digest=sha256:1322738b9ea9f9da6a7cf194e66afb1e2c563fd8b07d3de89fc828cdec534206

Observation 768f4ee9-819d-409a-82a3-6511bcca86bf · outbound

This paper cites TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.230798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.230798Z digest=sha256:3ca6e847968f906101914ae42cf52ddc0815bf94064eb0554e94c5dd4eeefe46

Observation 85d412e4-cd40-4c36-8b56-4669993426c0 · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.261487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.261487Z digest=sha256:d94c5d3d31add02314da0ccc749f2a56d7e610a63413a8255f638cf1cbf67e60

Observation 86254fe8-b730-415d-851c-388c2eae1875 · outbound

This paper cites Compositional chain-of-thought prompting for large multimodal models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Compositional chain-of-thought prompting for large multimodal models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.282182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:13.311730Z digest=sha256:cefc49a3f9f61255fe5fa35f89221b5d2ece8b63d36b8d0f483014f6bdb49669

Observation d1e06d9c-a647-4703-aa0f-1712997adb57 · outbound

This paper cites O3 and o4-mini system card.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking O3 and o4-mini system card

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.169036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:13.342588Z digest=sha256:83a956b626452e68e5ad7544fc941b23b25e799513bf6918f2020cb2ac1f6bba

Observation d2cd426f-ae60-4c5b-a659-a910706b907c · outbound

This paper cites Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Skywork R1V: Pioneering Multimodal Reasoning with Chain-of-Thought

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.380294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.380294Z digest=sha256:076f3fef5a23a887cd7a946b2eb9338cacda601eed1705c4a400c61f02f1b52f

Observation 7f4e8530-e3d7-4994-a8a2-82b77e6730e9 · outbound

This paper cites O1 Replication Journey: A Strategic Progress Report -- Part 1.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking O1 Replication Journey: A Strategic Progress Report -- Part 1

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.434001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.434001Z digest=sha256:abbcae38e05409b2d85556faeaa9e6ad5eaf68a9c2982ea443a0b440572f60ae

Observation 7b8c2e92-d821-40c1-b36f-71bedf13125f · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36:53728–53741, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.463903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.463903Z digest=sha256:8845fa1e0b0f7f6a7217fccd17caec73fc734a698edb241493f95782857d5916

Observation c6d3d20f-5991-4e78-8cee-5c35904d0653 · outbound

This paper cites Rethinking Reflection in Pre-Training.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Rethinking Reflection in Pre-Training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.515733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.515733Z digest=sha256:ac3a3ca43bcb044f4ac05bb76464599907f0eded5c370abc13dca8cc32f54653

Observation 9ab4e716-36fc-488d-9445-0b2094067eea · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.537018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.537018Z digest=sha256:8619e2a2bb34123f9f117c2921223cdf119bc944f539980e2fedde0ab6aa7b85

Observation c9f8922b-8b8f-40fb-ad7f-a5e4a598770b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.586726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.586726Z digest=sha256:d51fb3f4eae07692582020c46004886ea14346847a997aacc46aad88e14aef23

Observation 983461eb-e95b-4f81-bc9b-e7161ec4e2e9 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.667498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.667498Z digest=sha256:11ed137ac10821376c62227b96c67494f0f1a99b557588ed6301e07d00f4746b

Observation 9623ce81-8064-42c0-b353-2e803c167ba6 · outbound

This paper cites Kimi-VL Technical Report.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Kimi-VL Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.706010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.706010Z digest=sha256:a47e33d6c965867bd2540ad6ed40e5046d6362b5064687c3f0b077964d3321bb

Observation 443fff95-c174-4f1b-b701-1aa212e2e4d1 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, 2025.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Qwq-32b: Embracing the power of reinforcement learning, 2025

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.751900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.751900Z digest=sha256:d1cc65ec28625a67ffdfc3642f94f3df23a2e3e5b34b8a279cda8d5378460eca

Observation cf62bded-4187-46b7-9d7e-cff9c6edfacc · outbound

This paper cites LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.783825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.783825Z digest=sha256:136a9cef3806e7e123f13ec0f7304888b84ac5f5844386596ae4c2d40890ef5b

Observation 9a19dd4a-05ab-423d-8799-5b2d39bfd03e · outbound

This paper cites Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Automatic Data Curation for Self-Supervised Learning: A Clustering-Based Approach

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.815176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.815176Z digest=sha256:ef68480f0476cb2a6aed505948ada31ad6bcd4f168d35183b291265a23a8e4da

Observation 7e2bbec1-3765-486d-be83-792e41dc3a9c · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.Advances in Neural Information Processing Systems, 37:95095–95169, 2024.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Measuring multimodal mathematical reasoning with math-vision dataset.Advances in Neural Information Processing Systems, 37:95095–95169, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.847422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.847422Z digest=sha256:328e656ffa5535949d530cd388752daafb9ed09b8e733e5bde4e04b26e942d16

Observation db4734aa-bd8f-48fb-8b6d-ad871887ec71 · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.922969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.922969Z digest=sha256:01f5652f331c3ea78eeaad88bf689a80ffe514cf8a340cbe9a975ccd7c11a0b9

Observation a17f981c-2931-4cd4-a188-00780aeb2972 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:13.976791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:13.976791Z digest=sha256:5214fac029c9c43408ed1dcbf2583c64130db1416cdb28a1ac9fd0006371626f

Observation af52cd2a-3648-4a19-8272-91c349c85350 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.016315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.016315Z digest=sha256:612831adb153d3a1a6e8284f9e5bd98ff15ebe5247f93c2ab87fa70dd767814d

Observation be5d73bb-39cd-4d6c-9ab7-4903ac947a3b · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking V?: Guided visual search as a core mechanism in multimodal llms

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.047663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.047663Z digest=sha256:f491d7b12bfd5fe8d42ebde7508b6567f392ee2ecd3a2ea73f139ff787921a9d

Observation 6f645a90-1413-4158-bc97-bd2620bea1e5 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.082910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.082910Z digest=sha256:77e6ebbcabbb8a7598f1a35ca0dd18fd9918f8c5e3bb807fbfd36cf9091b68cd

Observation f7ffcbb5-699a-45e6-a7b4-69e78ef5f9e6 · outbound

This paper cites Valley2: Exploring Multimodal Models with Scalable Vision-Language Design.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Valley2: Exploring Multimodal Models with Scalable Vision-Language Design

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.123082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.123082Z digest=sha256:c5b2815f27fd336ec2162c768416265ebfd3f7684cd56563a2271887f0b5a9f0

Observation 325687c6-cd93-4575-9e89-36946bb2e03e · outbound

This paper cites Grok-1.5 vision preview.https://x.ai/blog/grok-1.5v, 2024.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Grok-1.5 vision preview.https://x.ai/blog/grok-1.5v, 2024

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:18.044283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:14.154557Z digest=sha256:8fe32f57c3be3dc0738b7224b080ec6581e7ac25b68c6aef8e8dbdeff193fe4b

Observation 8216b34c-7b24-4c7d-b699-180713ba760f · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.196592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.196592Z digest=sha256:9633f60b91bfc6c6efcaa2dbf05fd066c8f3d1db80ff94b4ef81f873b8a3354e

Observation ef1a0837-1386-407b-829c-63f946f193cb · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.243341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.243341Z digest=sha256:0bc518906a7497d0ebbc919b29c136d2fec0468dc8dcedb4135a352906ee046c

Observation 126c7a05-0e26-4e47-af71-db622ff377b7 · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.280040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.280040Z digest=sha256:f4038f5f062c1b6eb68c40574d8b36238b5584de3144ed69536e24609a79d515

Observation 5cb9a2fa-18d6-447e-a62d-cf036c3f1d21 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.320981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.320981Z digest=sha256:5bc3399517edbd61f8b6f757cdd43a531e694c91bf3dabdea9a39561315d7e9d

Observation f2a6f081-263a-452c-8fcf-e02f9268efb1 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.363044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.363044Z digest=sha256:bfae14817a49e239e6b36db29b28bcdbc6a26ef8dfaddb316102d64e3981e564

Observation be287d4e-fbca-4dcb-bccf-494b8b5b684d · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.arXiv preprint arXiv:2405.17220, 2024

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.405560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.405560Z digest=sha256:247d75fe7a0d876030820671b6805624e663d10fde990655b35d2f6217ffdfaa

Observation eba62b12-8943-4eb0-90cc-1e83da13778c · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.467663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.467663Z digest=sha256:63d79a533c893a198d309636be1c3561dd0b927a4ef084943079fbf6372392a5

Observation 5fc1b220-d50f-4c6e-8177-b57b88ed8c5e · outbound

This paper cites Griffon: Spelling out all object locations at any granularity with large language models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Griffon: Spelling out all object locations at any granularity with large language models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.944518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:14.511694Z digest=sha256:b813b0e8a86f8d683ca0e4cca147ecee71d3190ddcc78deab80b517054c7aa84

Observation 2d6a7a05-c0da-436e-bf2d-44ed9fd03936 · outbound

This paper cites Ferret-v2: An improved baseline for referring and grounding with large language models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Ferret-v2: An improved baseline for referring and grounding with large language models

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.766547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:14.559871Z digest=sha256:6018311858458dc1b21923539d4e3ac443354b9b8b9efa4d3f9c29aeca074cbc

Observation 9d312232-0a92-41c3-a7d8-a8b6e7113e02 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Improve Vision Language Model Chain-of-thought Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.605390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.605390Z digest=sha256:d7bf52adfea2547e27fcc4584e0f353526296b1ca2788ec89096f1aef8384804

Observation 4a5719e8-17ef-4d2f-994f-0806687c0102 · outbound

This paper cites SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.660602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.660602Z digest=sha256:0cb435adb1819034e555852e9588c6e91917f2a951a1b81693e41c26d0789c5a

Observation af67a684-595c-4119-b025-89b1851ed955 · outbound

This paper cites MM-RLHF: The Next Step Forward in Multimodal LLM Alignment.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking MM-RLHF: The Next Step Forward in Multimodal LLM Alignment

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.702270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.702270Z digest=sha256:bd694dd2bd3000e24345dd4788e53df4f30206e95b35245e5c3c92958ff8e83a

Observation 5e94eb29-f3b2-481f-9344-cd4227812f63 · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Multimodal Chain-of-Thought Reasoning in Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.744601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.744601Z digest=sha256:7e1316ac377bb1f611e64893035ad08648fcbb9ecaa9bac1e2a8e89638a122ff

Observation b61ee54b-df04-4705-99f2-a932e7f313c4 · outbound

This paper cites Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neural Information Processing Systems, 36:5168–5191, 2023.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Ddcot: Duty-distinct chain-of-thought prompting for multimodal reasoning in language models.Advances in Neural Information Processing Systems, 36:5168–5191, 2023

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.808518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.808518Z digest=sha256:ac5e2785636affa754b15491ea695075b251676bed1ee1b8dd9b6cc4c5e25ae5

Observation ab1a015a-c659-4ae2-979b-d101bc05bf37 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:14.847482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:14.847482Z digest=sha256:3b396819e3cd5334b2300f38cc85bc58326bca0ed6176693a143baaf54c21145

Observation 154118fb-2756-4d03-a647-3194e61ddeed · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:17.641879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:14.880283Z digest=sha256:86ce408b395331839972bf7fd80cf08e7be29f422cd520cec389444138321eb0

Observation 49b668c6-1839-40b1-b341-9a14a4da1826 · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:17.545042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:14.915753Z digest=sha256:cdb729355eea58a34eeac45dcff2c4db12b1046a10dcf072d5d3ea41428235fa

Observation 8b139858-c0d0-47db-92a4-2475a2512a97 · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:17.460148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:14.970907Z digest=sha256:4bbfde9540d600477d02c5e3a522652a9458caf2eede3cd7ec057b611c7a75b1

Observation 3bb65ca0-0f86-4c04-9581-9da1db0c5bde · outbound

This paper cites model’s chain-of-thought (CoT).

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking model’s chain-of-thought (CoT)

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.375472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:15.015345Z digest=sha256:8f4b075d61cbdb14f8e041d458779afe04899c8542397b5f92b27e34878b26b2

Observation 0337b9b7-4972-4951-80b1-eb404c4cef5c · outbound

This paper cites - Then, wrap the model’s entire thought process in <think></think>.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking - Then, wrap the model’s entire thought process in <think></think>

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.288363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:15.075607Z digest=sha256:f7880b9f5a6784a50d8b9fdbd952c95a4a4a6a4efe3341d2a74196c56b1d6f7a

Observation 03b5c2e2-ab2a-49f7-85a9-c3c61f023780 · outbound

This paper cites <vcues_1>, <vcues_2>,.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking <vcues_1>, <vcues_2>,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.211611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:15.125236Z digest=sha256:d4c92781bc86b3494c54f33afbf1eaceb565d6a7e641ea7f0517403c9f94f8b1

Observation 92a74542-590d-4e50-a054-f3ad6876cce2 · outbound

This paper cites All the data to be processed now concern reasoning errors based on visual cues rather than errors in visual cue perception.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking All the data to be processed now concern reasoning errors based on visual cues rather than errors in visual cue perception

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.139294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:15.156197Z digest=sha256:08ad2c5d5a703c032e26d723f44e32d12f16ae01d05f4e14a05f933afdf86670

Observation b2715fd1-294c-463c-987c-a6d13ec84414 · outbound

This paper cites based on the rationale.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking based on the rationale

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:17.067581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:15.182656Z digest=sha256:3d40ae52c5268f2f705a39605e4bda1a6b6e934574d5274fc8d91d9e171df479

Observation 9abc4366-1090-4d11-b7c8-1a1192610bac · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 79

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:16.981376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:15.223161Z digest=sha256:1cb2e41388903f36bdbd6ddf11adda91f47c7e962cec8c89aae76e875bd09e16

Observation b3501fb9-f946-451a-ae49-8aa9c067f709 · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:16.906529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:15.258999Z digest=sha256:07937bac6b3a4512f8c25742ef75598f35c9f8d3e5bf08f48a14e13ee0810fdf

Observation 0ec0a9da-6c2b-42c2-82e7-bd83727448e0 · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 81

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:16.830330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:15.298530Z digest=sha256:3a4eccdf696f59ac12c84d88b29427f516959cef9facb52c4544d8a8369d8412

Observation 7efba94b-a213-4ed4-a25a-0e79e112b29f · outbound

This paper cites Let's verify each visual cue and its reasoning before finalizing the answer.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Let's verify each visual cue and its reasoning before finalizing the answer

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.740232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:15.335118Z digest=sha256:fdf756bab01186605327b0d7d17774d1aec8ebafa96b8be10c4d197596666695

Observation 37267687-4afd-40a6-98bb-e09353490526 · outbound

This paper cites an unresolved cited work.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:55:16.651087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:15.400221Z digest=sha256:87d3fe61427f474bf96eee69ea16f09ef6603dd5e994c5faa95d7d3a02d82255

Observation 8c3afcc1-d706-4bba-8d3f-c0c4259cfd39 · outbound

This paper cites - The angle 78° is an interior angle of the triangle, and angle 1 is 42°.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking - The angle 78° is an interior angle of the triangle, and angle 1 is 42°

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.564150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:15.434526Z digest=sha256:69dd4327fe976d23716ee621a69fd5b02b35a2761621ca4876990adc5dc22864

Observation 1993d855-5ca4-4b5a-a7dc-9b2a6750ef4a · outbound

This paper cites However, upon reevaluating the problem, it appears there might be a misunderstanding in the interpretation of the angles.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking However, upon reevaluating the problem, it appears there might be a misunderstanding in the interpretation of the angles

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.470842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:15.463161Z digest=sha256:b0e25f3c792a7d3f15aab743da0b717fc34ba1fb5ff148e7cb0701cf17178991

Observation 27496727-abc3-4c17-824f-ace683444b01 · outbound

This paper cites - <vcues_2>Angle 1 is 42°</vcues_2>.

GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking - <vcues_2>Angle 1 is 42°</vcues_2>

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:55:16.279875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:55:15.498636Z digest=sha256:98190b06f10bc392ffba65d6d09064f26b1ed76a4026a679c952d68de1c12fd1

Pith citing papers

Observation 9bf62d37-bacf-46c9-9528-1e292c4f2fe2 · inbound

SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law cites this paper.

SafeWork-R1: Coevolving Safety and Intelligence under the AI-45$^{\circ}$ Law GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T14:37:19.575834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:37:19.575834Z digest=sha256:01f9f5619b8f9f4aebc08ce2687d6af9e682948c1b0abd0c2246f24ef1f3f429

Observation 51434a38-27d4-4cdd-9a2e-fd34b8e95e25 · inbound

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey cites this paper.

The Landscape of Agentic Reinforcement Learning for LLMs: A Survey GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 240

Resolution
verified exact
arxiv_id, observed 2026-05-18T19:21:48.600332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T19:19:36.427337Z digest=sha256:e39cfe4cd29e52c6c0053ce5bcd0cc1626d8d993ffd659a9cb9da91b1db32c0a

Observation 74e75a2a-7660-4070-9f2a-35d038cf8ac1 · inbound

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle cites this paper.

Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 234

Resolution
unresolved
no resolver link, observed 2026-08-04T16:07:46.730039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:07:46.730039Z digest=sha256:c3329a9f488de278e0e9d450ad8cf32bf6027621a0811844a7bea5d82900cc56

Observation 3b5c11b0-909c-4144-9dbe-7e7e7e6e00a9 · inbound

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection cites this paper.

See Further, Think Deeper: Advancing VLM's Reasoning Ability with Low-level Visual Cues and Reflection GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:36:17.576402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T04:46:16.497585Z digest=sha256:8209c1cda7b88f7a0a6023a6e4e62340694c72845659a5bc51e64ac490a429ef

Observation 8160789e-f389-48e2-ab36-3b6de0efeffc · inbound

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models cites this paper.

From Seeing to Thinking: Decoupling Perception and Reasoning Improves Post-Training of Vision-Language Models GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:13:21.576609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:13:03.237427Z digest=sha256:c38b3457d251195b70c7d4c9f1c39993cf96693910dd1ac8dee0e0690ac0db24

Observation 4ea4b481-5c01-42f1-8ad5-12453ff09425 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models GThinker: Towards General Multimodal Reasoning via Cue-Guided Rethinking

Reference 217

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.333992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:49b111ec1ea2901d86b5fc2355171f336ffae28a53a199d76218ab800b04f29d