Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:40.716524Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 12 inbound Pith citation observations for arXiv:2501.05767.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:40.716524Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:01:28.455317Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T15:09:55.026823Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b30b4aec-5cdc-478f-a674-c70b7fe61afe · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cca933e-cffd-48ff-a901-e7d10c9a0c2c · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Multi view image surveillance and tracking
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b1c67119-a184-40f4-829f-73ce6c65083c · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models InternLM2 Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c3b2190-4ce9-47b0-a0d7-b9853fab6462 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Position-Enhanced Visual Instruction Tuning for Multimodal Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00bcfd2e-6a96-43f3-98ff-5d91c0ccd2c1 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dc47472-2d1e-449e-bd59-8cc2b0e2ff1c · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Imagenet: A large-scale hierarchical image database
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 60d252de-423d-47a6-b69a-11144dd6f1e5 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Imagination improves Multimodal Translation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fde9781-d160-4f1f-a4bd-0c8bf1502efa · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Lasot: A high-quality benchmark for large-scale single object tracking
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 883fad47-9807-4929-b45c-d6b1ae557098 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5a77d2f-59d6-410d-9198-5b1d8242d2c0 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Mme: A comprehen- sive evaluation benchmark for multimodal large language models, 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e016a139-3425-4f18-bda0-1e84d9647eef · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Blink: Multimodal large language mod- els can see but not perceive
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2d0554e3-53a9-4c3f-a3e4-9f1a6ea74ed5 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Ego4d: Around the world in 3,000 hours of egocentric video
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1138ccf4-0253-4ac8-90dd-b597bc4ccca3 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01ccf045-043a-4818-bdcd-84f6255c5eec · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Got-10k: A large high-diversity benchmark for generic object tracking in the wild
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3534631f-0115-4469-895c-735fdf42a91b · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Editing Models with Task Arithmetic
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06bd7538-097a-486c-b90a-f7143064a9c0 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Distilling Translations with Visual Awareness
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ed272f2c-1cbd-4ec6-a8fc-a8f4f6cdef4d · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Learning to Describe Differences Between Pairs of Similar Images
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 410fa98c-719c-44ef-ac9e-3f9771177c89 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d45ef472-ecbd-4580-88e2-d8349dac5d33 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3b7f030-d9f4-4756-bfb5-539380dd7744 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Referitgame: Referring to objects in pho- tographs of natural scenes
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 65a8985c-0b1f-4ee5-a539-8c7daf1c662d · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b5e9e031-0165-48f8-832b-3d3e9d062d03 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Gonzalez, Hao Zhang, and Ion Stoica
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cd547a0-359c-4e75-9f9e-cf09f0146ca2 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Seed-bench: Benchmark- ing multimodal large language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a808c9b6-1b81-496f-8207-a2cd1df8271f · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 941e3834-430a-4454-93d8-e918e571fb39 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17aa2a2b-b946-41e0-b321-f816fb3d129e · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Ground- inggpt: Language enhanced multi-modal grounding model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d09c57db-89ed-448f-b9ac-663e846cc7cc · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Microsoft coco: Common objects in context
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ef4b137d-6cb2-4f60-8e99-e537413b3477 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MIBench: Evaluating Multimodal Large Language Models over Multiple Images
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40107a42-142f-48c8-b58d-708475ffccea · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95752fd5-6da9-4e5c-a2a0-1234acb3658f · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31c2891e-4b05-4f44-8863-ec6f40c1da27 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 742f7802-8d86-4b36-8045-ea83bea9271e · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MOT16: A Benchmark for Multi-Object Tracking
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f587323b-1d15-467b-88f0-0068b35d64a1 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Trackingnet: A large-scale dataset and benchmark for object tracking in the wild
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b84aec49-bb5a-4504-ac30-4a3b0fc599c4 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Robust change captioning
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 619e72cd-00b9-4e6f-82aa-38fe3c3b4f0e · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad02df87-c198-482c-9495-3ea34ba06f9a · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Glamm: Pixel grounding large multimodal model
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d1afe61f-7315-4a17-8f8c-20df3bb1af93 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models High-resolution image synthesis with latent diffusion models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f3269a34-7e77-4bc9-a885-2f59e438e7e0 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Objects365: A large-scale, high-quality dataset for object detection
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ade021cd-7e6f-41ed-95f7-4adba5eae456 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ecfef9e5-880e-4904-957b-2da1627fd922 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1f1c0e6-9337-4ff9-9ed0-b5fe29a87c28 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51560bf4-0335-4d73-864d-e466776232d1 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d2bb530-79c7-4b0e-ba37-9d90cdd58618 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f724ee3e-9150-46fb-baec-f086dd201a47 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models V?: Guided visual search as a core mechanism in multimodal llms
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 154ba1ff-9d01-4503-94dc-3f7098f5277a · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a6118e7-d5cb-4af7-b94c-e384b4ff413b · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Qwen2 Technical Report
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13b99598-df00-44b9-8e97-4beceb65f314 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models ReAct: Synergizing Reasoning and Acting in Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 229b63c0-920d-4bf5-8e54-ef2870373f59 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75e807b4-543c-4b27-8c86-056d1dee2615 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f755047f-3c75-4c69-b5b8-fe4b0da45ea8 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54bc825f-9406-402d-9d7a-3aa8c46e3c24 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Ferret: Refer and Ground Anything Anywhere at Any Granularity
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 407a7a7f-e09c-4a35-870f-65152987203c · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da96d517-cdc6-4713-b2fd-604f638a4924 · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Magicbrush: A manually annotated dataset for instruction- guided image editing
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e611620f-9456-4039-991f-91af93396a1a · outbound
Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85eec8d3-62e5-4e27-9f0f-024909f90687 · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
Reference 135
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0e364c4d-9564-4082-bdab-640e6b79ba5c · inbound
Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e55e4bdb-d8b0-437a-9897-698b0016b6da · inbound
MedSG-Bench: A Benchmark for Medical Image Sequences Grounding Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48ecc93c-134f-4f3b-bb96-9ec944c1ce57 · inbound
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2274912a-62dd-498c-9c8a-40ea3cd393f8 · inbound
PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9940576a-4182-421d-b421-9bfacf7d6405 · inbound
Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 63fd2520-a011-4839-b7eb-25fdabc0b8cb · inbound
RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98d23dba-4d73-42f7-b1e8-54a3b0e3bf51 · inbound
Training Multi-Image Vision Agents via End2End Reinforcement Learning Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a417a851-0566-48ec-9742-9f2b06906b2f · inbound
CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2da49c6c-3e9a-4b2f-929b-abefa8092792 · inbound
Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ba7be77d-951f-473d-bd51-a39bc7e96394 · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
Reference 159
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9aaaefa8-de82-4fd4-af0d-07d8ec5d2d9a · inbound
DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.