Pith. sign in

Paper Citation Record · LEDGER

ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2411.16044.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16044 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:42:57.953040Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.313006Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6b4bfc76-4bee-4a75-ac32-64f54fd4d0b2 · inbound

Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning cites this paper.

Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:36:25.241161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-18T13:33:12.508639Z digest=sha256:b7561b21673596c0ace9ed53bb3e81b3f30931ababae67899d5ae22a9a67d690

Observation 6ab14e41-1ab1-49c4-9ed4-d75081c3c7b3 · inbound

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling cites this paper.

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:14:22.830823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T21:11:47.804851Z digest=sha256:11574f4dad575a1e04448f52f01d75abfdff26e1a02be4dd690ce19ac3673f1c

Observation 1b62e254-892a-4ea8-b532-a8c72a3dbdde · inbound

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling cites this paper.

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:57.953040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:42:57.953040Z digest=sha256:4ae393455b4ea8963ec2def5ea7a67344e77b53ed028ba978df5cca5a027ce10

Observation 9175071c-5681-4c59-a890-d355ebf7a666 · inbound

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search cites this paper.

Beyond Thinking: Imagining in 360$^\circ$ for Humanoid Visual Search ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:18.308575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-12T02:58:46.728868Z digest=sha256:438ef4102c198f797907efa0d9e67d34d9a4cf0ed867aa47fc8368af7a35df43

Observation ab1ae90f-18d1-4f8f-81a6-e3c2bc8047a4 · inbound

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning cites this paper.

Are Tools Always Beneficial? Learning to Invoke Tools Adaptively for Dual-Mode Multimodal LLM Reasoning ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:18:05.286120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-20T06:16:47.650748Z digest=sha256:284599bf5fa13e4f245ba7a91cd5f898e959c441dc65c4b69f8975a2f7927f1c

Observation 9ea350d7-4329-4003-a14e-9256b27b3f8b · inbound

Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG cites this paper.

Look Before You Zoom: Adaptive Routing for the Resolution-Context Trade-off in Visual RAG ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:49:38.564472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T12:54:06.319322Z digest=sha256:271f06ecf95443ccf536968fe23220a7988ae3d1ef318971a0784156d4e5269b

Observation a16f3f0b-d4b0-43d7-b71e-54abef94fb48 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Reference 136

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.314897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:7cba696e90ab56979bf7989299a45efc607b39e601897dbec9aaa365def8ab48

Observation 35fac300-f5e9-49d0-a713-41c44bf56be7 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Reference 282

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:797d8957254867353adc627981ed631b8b9df0ceeef88ce2596302dedb18be00

Observation e912b465-8b2e-418b-a622-9c4be1ac5a3d · inbound

WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding cites this paper.

WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T14:09:30.395518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:09:30.395518Z digest=sha256:7772cd5935c46a1c8c66118ee152bb0d5aae282c5d3566c8cd22e9f06d22493c

Observation 1b0f9ef0-4caf-448d-951d-91256e927aa7 · inbound

AdaTurn: Budget-Aware Test-Time Scaling for Active Visual Perception Agents cites this paper.

AdaTurn: Budget-Aware Test-Time Scaling for Active Visual Perception Agents ZoomEye: Enhancing Multimodal LLMs with Human-Like Zooming Capabilities through Tree-Based Image Exploration

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T01:50:55.420352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:50:55.420352Z digest=sha256:eb85e1070753e36c3d63d26b0c9897fab14ce4d9edf3d17272bde01945c9f097