Pith. sign in

Paper Citation Record · LEDGER

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2501.05767.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05767 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T18:15:05.561502Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.026823Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 85eec8d3-62e5-4e27-9f0f-024909f90687 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.209391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:d05e0b8b8d51956f1fee617c3bcfd52640c8c4188a3f4f5516398596b89c0ba0

Observation 9940576a-4182-421d-b421-9bfacf7d6405 · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:52:08.087721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:939c48c978580254169dd839075f1d2632ec5babe8144fcea195811419db362d

Observation 63fd2520-a011-4839-b7eb-25fdabc0b8cb · inbound

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension cites this paper.

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T18:15:05.561502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:15:05.561502Z digest=sha256:229d7708950f5a3833af2a9e5180913778748aa8fdac7f080082a3d86639fc73

Observation 98d23dba-4d73-42f7-b1e8-54a3b0e3bf51 · inbound

Training Multi-Image Vision Agents via End2End Reinforcement Learning cites this paper.

Training Multi-Image Vision Agents via End2End Reinforcement Learning Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:01:24.328661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T00:59:28.618477Z digest=sha256:e9caafe9d50178336efee47fe008f337c41b935d3e1081bfc7024e6d8876f85c

Observation a417a851-0566-48ec-9742-9f2b06906b2f · inbound

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding cites this paper.

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.325270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-08T12:26:01.568507Z digest=sha256:35895dc63cbe7ab6e7ed1d01b935f1322754204c2be52c684a7754b18c2ef465

Observation 2da49c6c-3e9a-4b2f-929b-abefa8092792 · inbound

Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning cites this paper.

Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:54:00.616287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T22:53:55.407461Z digest=sha256:47f6090ea58e4aec8bc07baa19066f40ff099de32855d660d4c82decdd37f936

Observation ba7be77d-951f-473d-bd51-a39bc7e96394 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.029033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:b73c8604d5b5fe6b9d43b882ee0f776e6a528cb862e9deb8a432bd9e5ba5296d

Observation 9aaaefa8-de82-4fd4-af0d-07d8ec5d2d9a · inbound

DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues cites this paper.

DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:19:50.963105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T05:18:01.931929Z digest=sha256:140c7eb907ef9d6c636552c7d4b9377a3a394a7862a1b00a810b31bdb6723f97