Pith. sign in

Paper Citation Record · LEDGER

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning

As of 19 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2607.10004.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.10004 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T01:07:06.861298Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69a58c99-618a-4049-9a3d-f2fc857e730c · outbound

This paper cites Qwen2.5-vl technical report, 2025.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Qwen2.5-vl technical report, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:fbc9e0373960500a34151ab75b6ccdd9341f08a7cc9f744445812d90951582c4

Observation 0e4f160e-9c36-4d44-8251-13a76a3702fd · outbound

This paper cites Reasoning in the dark: Interleaved vision-text reasoning in latent space.arXiv preprint arXiv:2510.12603, 2025.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Reasoning in the dark: Interleaved vision-text reasoning in latent space.arXiv preprint arXiv:2510.12603, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:99af51be850da60712b10df3c80bbaa139e0ed0f2ed90cfb40a4cdd39866c49f

Observation 38b977a1-16bb-4a65-9a9d-7361fec5b5cc · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:7382437f568b7a23016427eea106025c5cca4b39ed1a387332c3a9b4d17ddc8e

Observation 223faf9f-2329-42f4-a446-0399647c6251 · outbound

This paper cites MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:34a6c7a45a3f9e3c2a5eb821a0d1fa99a50a3f767efd0ae3527f17a2f829bdb1

Observation c86377e5-063c-4b8d-ac55-921c6d72079c · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:7c9eec06022499fecb74974a1f1dde38a77e9f63652ef15ac49bc3ef64f057c3

Observation 41e34dee-a2fe-407c-9dd8-a5313d7f5253 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Emerging Properties in Unified Multimodal Pretraining

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:0d576a5e80ed6a87f1de5d2542b627dc3de25b21a7f5a81d06a378efd9b0c827

Observation e636d312-d701-4694-a1dc-29b0bdefab24 · outbound

This paper cites Codeplot-cot: Mathematical visual reasoning by thinking with code-driven images.arXiv preprint arXiv:2510.11718, 2025.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Codeplot-cot: Mathematical visual reasoning by thinking with code-driven images.arXiv preprint arXiv:2510.11718, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:ae6fc369ef3bb83cb842029b3f287a0fdf4286d3d09b79d1b48503569124d5a8

Observation cec7d7d6-c797-4468-a227-9ac89c904e7a · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:38ed06481224b728c545616aaa4e3b1207c5d7db9d23a34c3b88e40b5d190bbf

Observation 5725318c-0a66-4463-a36d-c9198221cbfc · outbound

This paper cites Interleaved-modal chain-of-thought.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Interleaved-modal chain-of-thought

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:b6a47a973d1e8a6a0d5cb5561fd8bedf19476af09d09a59e3b844c0707d97725

Observation 1730a604-eb8c-49ea-9e3f-50043610d767 · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:bf3970ddca344e39b6a2146ad32e078c99c8e90aa62ec3e153c801a96ea29115

Observation f322e518-5c2a-4435-b8b8-fabde3996f6b · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Visual programming: Compositional visual reasoning without training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:fc177063467678eeb97a0058c78092eed30760697afd3cf0f81196c34087724c

Observation ce384a81-1936-4c91-81f3-010ca7dd6e48 · outbound

This paper cites Unicorn: Towards self-improving unified multimodal models through self-generated supervision.arXiv preprint arXiv:2601.03193, 2026.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Unicorn: Towards self-improving unified multimodal models through self-generated supervision.arXiv preprint arXiv:2601.03193, 2026

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:94fff81147d52adc88d38c5f3b8459144acbf49cf343a11e5e69a787d21d8238

Observation 953ff31a-902b-4624-b9fa-abc84c68bc07 · outbound

This paper cites Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Olympiadbench: A challenging benchmark for promoting agi with olympiad-level bilingual multimodal scientific problems

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:b81f851661e0dcaa2d22143f7f61ee67f171551af7118a0366632ac8511f34c9

Observation f06ed034-8ad5-4d0d-881c-a52deae6483a · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.Advances in Neural Information Processing Systems, 37:139348–139379, 2024.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.Advances in Neural Information Processing Systems, 37:139348–139379, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:cbac6039e99b518b24e6118afc479fa1a1dba8215e4d44b45a439c899bccf86b

Observation f459b292-e795-4203-9dfa-a44a4d617e20 · outbound

This paper cites GPT-4o System Card.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning GPT-4o System Card

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:0881c504928b5c42e5442d18dd8fd11bd9e4b14a93b46dc626a6e8cdcea1fa57

Observation 0ae3453d-01e2-4625-b431-27095c5e4fa8 · outbound

This paper cites Zebra-cot: A dataset for interleaved vision language reasoning.arXiv preprint arXiv:2507.16746, 2025.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Zebra-cot: A dataset for interleaved vision language reasoning.arXiv preprint arXiv:2507.16746, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:9cfcfd2ae3d24ee6cc93a48341458b3c871ae9d73338b58f4e83e60941c9c915

Observation 291dd5fd-2f9b-4bcd-8d82-bbbf42fe7f0c · outbound

This paper cites Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:f5e68804fa103c007b670660bb55ee51a85ca92467155af68b8618ec6c8cc39f

Observation ccc1023c-1020-4d32-8fe7-0795c0589f22 · outbound

This paper cites AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning AIM-CoT: Active Information-driven Multimodal Chain-of-Thought for Vision-Language Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:853e8bbb33fa6800711e3a00b6ae76b6c0ac634a509177ffeac34e30daa82246

Observation 8572acd0-e2bb-4609-b66b-c4c5d7976d52 · outbound

This paper cites Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:0cacb1e0ae5246293c6475952c27c768b3a48509af94cd475e488120e7a1130d

Observation a1a73ab9-c06d-46bd-880e-52615ef940d4 · outbound

This paper cites Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Let’s think with images efficiently! an interleaved-modal chain-of-thought reasoning framework with dynamic and precise visual thoughts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:cc4ec9eb1230ed932973ca676f27fd384a39d4fce1b97848e08c150f9edb674e

Observation cb79be7a-55e3-4535-99d5-e6aa9a91bfa4 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:3d496b9b3db96b637548bca65ecdb97995a598d5a51b75e112b036fda35ad514

Observation 75fc4684-086d-4358-8cb7-ba8102369c62 · outbound

This paper cites UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:85a70d17466b04cbb00a7b781b5b81ce2bda6ffa3e5cd40289e05aafb75e66d5

Observation 1b1df96b-c9bb-4e24-8364-eae0a0a2fd22 · outbound

This paper cites We-math: Does your large multimodal model achieve human-like mathematical reasoning?, 2024.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning We-math: Does your large multimodal model achieve human-like mathematical reasoning?, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:b3994f6b024845ce4a945e4ebe324ff4602d4f19f780b7f0fcc96b2ffbe804a3

Observation ac34e4d7-a7cb-4fbc-a915-d3a94461d78a · outbound

This paper cites Zoomeye: Enhancing multimodal llms with human-like zooming capabilities through tree-based image exploration.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Zoomeye: Enhancing multimodal llms with human-like zooming capabilities through tree-based image exploration

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:4d9c75bbdd6a64ee2e74030ba139bac2f215fd9478c447af3fa25818bfcf9eaf

Observation 0eeacc86-db98-4c73-9996-10309e73b28c · outbound

This paper cites Mathcanvas: Intrinsic visual chain-of-thought for multimodal mathematical reasoning.arXiv preprint arXiv:2510.14958, 2025.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Mathcanvas: Intrinsic visual chain-of-thought for multimodal mathematical reasoning.arXiv preprint arXiv:2510.14958, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:bca56315611eaeb8e459ddb17947fb007eb3e1e61561c12450f7d4cdf368d7e2

Observation c08a7f4e-7a9d-4e37-8503-fa604a3edb66 · outbound

This paper cites Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:f78604246cfdec3ea261ff571a15b297f62b6fc42f7a67aa5c672eae0425bb83

Observation f21d88a5-67c8-4458-b5b9-09c3371864dd · outbound

This paper cites Llamav-o1: Rethinking step-by-step visual reasoning in llms.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Llamav-o1: Rethinking step-by-step visual reasoning in llms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:655d97127a4fc8e7a3bc9177b006f26bcdfcb8472f565b31b89fc4aed830c5cb

Observation eb88b1f6-4543-469a-abd5-950832919542 · outbound

This paper cites Unirl-zero: Reinforcement learning on unified models with joint language model and diffusion model experts.arXiv preprint arXiv:2510.17937, 2025.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Unirl-zero: Reinforcement learning on unified models with joint language model and diffusion model experts.arXiv preprint arXiv:2510.17937, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:e0401f2a98f51af94fb8a4e3da0236e8a00d1a54bd9544b67bd5741878c246f8

Observation 428361c6-ce4e-4d4a-a58a-93819de614f3 · outbound

This paper cites Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:3ade3a9f74c461e330df2973ee0a190dccf18cda45c95c73208e7e678e5dce94

Observation 58d45922-b98c-42d7-8cf3-5bca708a1934 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Measuring multimodal mathematical reasoning with math-vision dataset

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:318f5ddea10a656ecfb70e9ad3e0af123b2b935689dd50b96c81c90d4e8a58ed

Observation 630449c4-f800-4b15-b9f5-767023bc813c · outbound

This paper cites Mathcoder-vl: Bridging vision and code for en- hanced multimodal mathematical reasoning.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Mathcoder-vl: Bridging vision and code for en- hanced multimodal mathematical reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:065825334b57bbf641bfa2d2f5ae29b6be03f38bc903c499500c2748ae4b1248

Observation cce4ce6a-476a-430f-9557-7ee0aabf3965 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Emu3: Next-Token Prediction is All You Need

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:9e1e91390953566a801c9c8f394c76ad14829de6269444f29a088e7c160dae22

Observation 936b2489-440e-4bfb-b4b2-a333137daa9a · outbound

This paper cites Visuothink: Empowering lvlm reasoning with multimodal tree search.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Visuothink: Empowering lvlm reasoning with multimodal tree search

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:af1d9f853bed4118bc9128c677b3d90a2aff1ba70827cfbfe916fe97317670cc

Observation a22b2653-c294-4940-beea-19cade8475c4 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Chain-of-thought prompting elicits reasoning in large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:2fc8bf9a9b705b03d1683319c780225d910f6684bd9d1e3beee530146db3a890

Observation 2f2ef16c-5b13-4d3b-a712-f50fefdfe3e1 · outbound

This paper cites Muse-vl: Modeling unified vlm through semantic discrete encoding.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Muse-vl: Modeling unified vlm through semantic discrete encoding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:be82f05cb0bc4a310d210da5e5fdb92084e554420091aeeda9bdafb62fd55bc5

Observation b13ac0bf-9f70-42a3-b165-cf056b68049a · outbound

This paper cites Llava-cot: Let vision language models reason step-by-step.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Llava-cot: Let vision language models reason step-by-step

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:18e757b343cfde5536d4d82eb3df810384ef5887d100bc937dbd78feabf0ec14

Observation 5e933b81-ad73-442f-8fe3-4765e11e0c6c · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:adbd7120e65c6a5ce6a8bad94b133865846eaa008fde6cc823622a5c7df45af7

Observation 7054361a-4db7-44df-b67f-bc57e2443442 · outbound

This paper cites ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:8898178c38b0cb8008ec689d88a0e20cce84e576c68fdcfe5a0108e80ad56971

Observation d21586ff-b7b5-45eb-932c-c74f9cd391e6 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InEuropean Conference on Computer Vision, pages 169–186.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InEuropean Conference on Computer Vision, pages 169–186

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:c220f93cb1eeb459f387cae1f6a31cce30a15e0efbda4b3300a4014419336856

Observation 76c03e6f-6774-41c4-bc14-a322c95b0a9d · outbound

This paper cites Multimodal Chain-of-Thought Reasoning in Language Models.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Multimodal Chain-of-Thought Reasoning in Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:e22560c68ce5f99b6d30c186b54ca0d20efc75a7fb3cbffe040184a739452375

Observation 899d1151-7cee-4937-bffc-4768efccdcc6 · outbound

This paper cites Unified multimodal understanding and generation models: Advances, challenges, and opportunities.arXiv preprint arXiv:2505.02567, 2025.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning Unified multimodal understanding and generation models: Advances, challenges, and opportunities.arXiv preprint arXiv:2505.02567, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:238ffb1615a778229b55197cc3e035c48b46565c63184b70f3d71bcd02a861a0

Observation 12ae6135-258f-48cd-bab4-be5afc0ad744 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:5bfd4ecee43119ff8919ef438c41eb1972d222700378c7c2d5f6bdf343ac6a6a

Observation ee54c62c-eef4-4ccd-ac6c-f89ced35fad5 · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:c9019307ad69384294bcc491bd8a70a16469c738d3c7f88d0df693345830aa4b

Pith citing papers

No inbound Pith citation observations are available.