Pith. sign in

Paper Citation Record · LEDGER

V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2312.14135.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.14135 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 39 of 39 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:29:19.034618Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c4752a49-937d-4f37-8801-70f4bc348d93 · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 208

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.239118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:26a86a3f2392ffa20437b28f804a6c7e80d53d2df77d679bc8e657acb42bf995

Observation 5d2ab2cf-19a8-45c1-8e0d-6a1e5c44ce3f · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.236393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:33ce2732b282ace8b8ee597231e7f6936c20611ab161a0a3fccc677da62badea

Observation 1d2ef04f-9865-412b-9b35-1307004e607b · inbound

Seed1.5-VL Technical Report cites this paper.

Seed1.5-VL Technical Report V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:05.692676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T05:26:04.960844Z digest=sha256:c704ab3de1d0149a68e7ced6646e42f33c6da51abc9e5b89111755924589f02a

Observation 5684eb9b-6dd6-4071-a5a2-23d709b86867 · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-22T01:05:52.086561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:dcd1599405c244a8ba0ca027712848f6c9b3e58e81f251b5cb3b17d7386962bf

Observation 71f3e515-41de-4d7a-b0bf-ad75ebc445e9 · inbound

Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research cites this paper.

Unifying Language Agent Algorithms with Graph-based Orchestration Engine for Reproducible Agent Research V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:19.034618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:19.034618Z digest=sha256:57ee9b49de3832bd518862830a509f8b9db685a90fafae9c505c8bbfec79c497

Observation 7599ee8d-24d9-4414-ba22-ad3aa2d74bf5 · inbound

MiMo-VL Technical Report cites this paper.

MiMo-VL Technical Report V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.241662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.241662Z digest=sha256:ea6f378cdd93f2efc5abc13071b5273af23949e70d4408d22ef9a899befac030

Observation 858d4b93-a6c5-4dbe-9a08-d68c203ae2d6 · inbound

Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations cites this paper.

Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:23.231324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:23.231324Z digest=sha256:2084b1c27578655e341f8800e7aac8894290e0b55ac1552013e7609f631b5ffd

Observation e3c446d9-6703-4203-825c-4baea7897aba · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.885634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.885634Z digest=sha256:ce895e00bc6dd03c35d080794ec72a613fcce9d0e77b448dc4f5eef7141ecafd

Observation e64e0d55-fe91-403e-8dc5-24610c1afd0b · inbound

Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning cites this paper.

Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:36:25.246023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T13:33:12.508639Z digest=sha256:ae983b5f7e64524be1a4443de8cb39a343619b42d8750a79070b9f6e23cead50

Observation a052efc3-2ef7-4196-8175-27680ac2b9d1 · inbound

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling cites this paper.

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:14:22.807207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T21:11:47.804851Z digest=sha256:03f5a90357d70837aa538479b9014c2b44685109d337e918dea731c7b50a42fa

Observation e7edd0ef-5bc5-42cd-8a8d-baaa308ee84a · inbound

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling cites this paper.

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:58.310244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:42:58.310244Z digest=sha256:190a0cca5c01478c478b8a2ea552a0c0b17d97737c3bf44ba982d7c7ab805f79

Observation 0bead9c2-52c2-44ea-9125-0979bab87762 · inbound

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models cites this paper.

Beyond Reasoning Gains: Mitigating General-Capability Forgetting in Large Reasoning Models V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-04T08:15:57.799831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:15:57.799831Z digest=sha256:e5445e4fd49f003f8eaf257c103c2abe3d27e0fbca8e5998f40be9751667c74d

Observation 978dc477-32ba-4afb-8e4c-f600cbfa318b · inbound

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding cites this paper.

Blink: Dynamic Visual Token Resolution for Enhanced Multimodal Understanding V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T17:09:38.165356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:09:38.165356Z digest=sha256:315b19e9ad55abd7fc83bc67769ac8eedee5071477b389fd287c7e74aee9f970

Observation 92a875fa-dc60-4fdb-91e4-ed1ceac05f94 · inbound

Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension cites this paper.

Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T03:30:33.046213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T03:27:53.694506Z digest=sha256:eab59ed6e29b9455b458760dc1f6f1da4c47bea0b4fd44f93e9800fc3500e9f6

Observation 1c20257b-e448-47cf-bad4-28b558dd828a · inbound

DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model cites this paper.

DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:05:04.972295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T12:04:08.487678Z digest=sha256:63e3cde125e567b78386a84b46a9d66deb6837a9c8f76fc4405f0ce2d26094c0

Observation 1e1dfc9c-74da-4b71-a0dc-5347b54124fa · inbound

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning cites this paper.

SpecEyes: Accelerating Agentic Multimodal LLMs via Speculative Perception and Planning V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-13T19:34:58.789459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T19:34:58.789459Z digest=sha256:42cb263609f323126dd38b792f7ad28c203eda322cb128c23adf65fe292512d2

Observation cc98f7b9-99a6-49ed-a866-90b9366f1cf1 · inbound

LanteRn: Latent Visual Structured Reasoning cites this paper.

LanteRn: Latent Visual Structured Reasoning V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T17:26:40.239147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T17:26:40.239147Z digest=sha256:6dc2cae1c8b196a388352654d6fdf8aab612e0b0cb258ea1d6f858fd76f14655

Observation 40ed16f7-83f2-455c-a1fc-4d451d3d4b24 · inbound

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models cites this paper.

Q-Zoom: Query-Aware Adaptive Perception for Efficient Multimodal Large Language Models V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:00:51.617866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:46:26.869644Z digest=sha256:b45cc38ce21c732e11d2e83439ef03c2948c26a639001ce65af8cfb6a01a944b

Observation 7eb14478-c1d8-4d8c-9be8-ebd448c8b100 · inbound

Multimodal Latent Reasoning via Predictive Embeddings cites this paper.

Multimodal Latent Reasoning via Predictive Embeddings V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:46:52.537656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:26:27.683269Z digest=sha256:a2daa17427ce1fcb1be740aed7337c560f155cf2b7ac00f2b83438cd20c2b780

Observation d99eb290-4d03-4448-8e02-44676820aed7 · inbound

Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models cites this paper.

Entropy-Gradient Grounding: Training-Free Evidence Retrieval in Vision-Language Models V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:16:10.605697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:14:11.941977Z digest=sha256:af92a2b03136668bc110b62412c54ae20e30654cb9beed7b542e1e62ca95bca6

Observation 83b8a0dd-9add-4223-8527-8a5659dfeeff · inbound

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding cites this paper.

MAG-3D: Multi-Agent Grounded Reasoning for 3D Understanding V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:51:18.794113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:25:31.097385Z digest=sha256:a2d7de010c0714e552f20e257c4617b23630c2c43987192e0728f3ab8e5ee814

Observation d3c74dd9-a762-4006-bd93-04f769f35297 · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:25:58.822046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:03:15.222571Z digest=sha256:0c152bc5589e77dfb79e02232c53ce1b88b5ba8fb2bdf8d9adc9cf7c437bc66a

Observation 983d9ea9-53f7-41d6-abae-b8d9662b0c16 · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 93

Resolution
unresolved
no resolver link, observed 2026-07-12T22:48:45.647588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:48:45.647588Z digest=sha256:862c484417548c16ca6fb34cae687cddfd8571bb6a1f7e401876927deedabc6d

Observation 814f6ec7-7150-4d8d-9bf2-18e47e7f8447 · inbound

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation cites this paper.

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 213

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:43:49.502913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T05:10:44.608959Z digest=sha256:fb4ba2086eee41922002d2b75346aa5d22dee01c8598e0695d6bfade092f5cb7

Observation ea694479-978c-4604-9b46-9c67c71b600d · inbound

SketchVLM: Vision language models can annotate images to explain thoughts and guide users cites this paper.

SketchVLM: Vision language models can annotate images to explain thoughts and guide users V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-09T21:33:27.942692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T21:32:08.503584Z digest=sha256:76ec0f6adaa9139ec3de3765f3a662571cf509ae799e160e7772fce82973b513

Observation d44ab193-c2c8-4f94-9ecf-b6b415bbd7b6 · inbound

Improving Vision-language Models with Perception-centric Process Reward Models cites this paper.

Improving Vision-language Models with Perception-centric Process Reward Models V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:17.052295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T04:33:36.634359Z digest=sha256:781636cab5348c7a30b0eb1e439c39a6d22ca39c5bd0d12c2f5f3c38cdcec2b1

Observation 4d1a14b4-c409-4d27-8502-8d53df71a759 · inbound

GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning cites this paper.

GazeVLM: Active Vision via Internal Attention Control for Multimodal Reasoning V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:54.650794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T02:03:52.566413Z digest=sha256:4b7afca9eca22d3ee14f81c9ebafacfa9c041f2aaf3b6b8681f601d8781e7b30

Observation ccad97df-3074-475e-b73f-7de3b8ff1d3e · inbound

What's Holding Back Latent Visual Reasoning? cites this paper.

What's Holding Back Latent Visual Reasoning? V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:03:15.473753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:59:14.134917Z digest=sha256:0e2bbf4ff724fff4e71893ed1df333013f216f3986f517e577d7a9c3d2e068da

Observation e89677b3-b613-488c-9467-a4394c408300 · inbound

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth cites this paper.

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T10:53:13.123588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T10:52:32.238241Z digest=sha256:f9bece1506bee5863a4fd74c87be670bbae82271d6dea7dbe46c409507a69640

Observation a867bb2b-1d00-4b86-9775-0b198c98424d · inbound

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth cites this paper.

Starve to Perceive: Taming Lazy Perception in VLMs with Constrained Visual Bandwidth V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T16:34:34.452430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T16:34:34.452430Z digest=sha256:988514f2dae4f3863f8b5f3bbe655064063c7b12fc8d57bb7efec688e501fffe

Observation d5408f8a-eebe-4ad9-a88a-24da1b98dd3f · inbound

Self-Prophetic Decoding to Unlock Visual Search in LVLMs cites this paper.

Self-Prophetic Decoding to Unlock Visual Search in LVLMs V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T13:03:26.884105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T12:53:40.783281Z digest=sha256:ed720f9c519ecfb7552422469b8f72c6016614e252ad7a0f4e116e845667fb8e

Observation 22009e46-56ed-46c6-ab6a-b1cf6c9a5d9a · inbound

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents cites this paper.

ToolGate: Token-Efficient Pre-Call Control for Tool-Augmented Vision-Language Agents V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:46:28.939546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T10:38:41.735204Z digest=sha256:b3c76372a5e100e731f1b7f90b6fdbce15c1bb1401d7e5b8adb7b4cef43c1765

Observation 79c455c0-6e7c-4123-848c-947c636d6fd2 · inbound

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention cites this paper.

MOSS-Video-Preview: Toward Real-Time Video Understanding via Cross-Attention V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:26:17.951269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T15:22:31.310003Z digest=sha256:d0d0c50cd53b334ad19710da04034aea3b7c56b8007e437faeba15693f304fbe

Observation 02a06411-4ac5-46f5-9891-ef32db2a3b91 · inbound

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning cites this paper.

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:49:02.271066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T21:47:51.437284Z digest=sha256:c1ee067c5ca3f9f48b71748c93a386be5ead4ef93eccc36ea0e2db377d7bb7ff

Observation 8497a1cf-e8ba-4d0d-a171-6b6a1eee502d · inbound

Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG cites this paper.

Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:49:38.554498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T12:54:06.319322Z digest=sha256:30a8451d72f14e2610a76ca99d314fc2cb43189992185566587578476d81cd09

Observation 3b1d8caa-5deb-486b-8c29-6611a5ad14b9 · inbound

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse cites this paper.

Kamera: Unified Position-Invariant Multimodal KV Cache for Training-Free Reuse V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T12:09:48.891429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T07:13:19.593093Z digest=sha256:74d8f0a30422bc5d5a976246d18347ad7c5af49d0549afd3389790b2b3f67ebf

Observation e17be238-b3bf-460a-8f63-3713f499ddcd · inbound

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding cites this paper.

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-01T16:48:36.519393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:48:36.519393Z digest=sha256:0f7d4b5b6e309c1a18d6d9981eb95e761c869d7ab3df74259acdc70bc22047ed

Observation 54c1de77-47be-4752-be6a-b354061bc3c3 · inbound

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation cites this paper.

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-31T02:56:58.290582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T02:56:58.290582Z digest=sha256:0d8e736373470a81ed013f28854a13f8a0f303cadfbab094626f7c0d06725c9c

Observation 3c555d85-f35f-4299-851a-8fcf9e14a703 · inbound

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs cites this paper.

What Transfers from Text to Vision? Capability Scaling Laws and Transfer Dynamics for VLMs V*: Guided Visual Search as a Core Mechanism in Multimodal LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T02:46:49.732380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:46:49.732380Z digest=sha256:0897ff72ec707fc34e3d2d8acc775f948ff4280a07846a9573fa71f0a035fc29