Pith. sign in

Paper Citation Record · LEDGER

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

As of 19 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 12 inbound Pith citation observations for arXiv:2501.05767.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.05767 v3

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:10:40.716524Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:01:28.455317Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.026823Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved33
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b30b4aec-5cdc-478f-a674-c70b7fe61afe · outbound

This paper cites MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MathQA: Towards Interpretable Math Word Problem Solving with Operation-Based Formalisms

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.463158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.463158Z digest=sha256:6f65a252158bd00a64bfbaf361663c00da0a76a6ad037fbfb7256a023d9c9930

Observation 8cca933e-cffd-48ff-a901-e7d10c9a0c2c · outbound

This paper cites Multi view image surveillance and tracking.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Multi view image surveillance and tracking

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:42.069985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.468599Z digest=sha256:6b7a9bd0d82b79afc798d46169e905151e23ccc2ea660e7293fe30155cac7d8b

Observation b1c67119-a184-40f4-829f-73ce6c65083c · outbound

This paper cites InternLM2 Technical Report.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models InternLM2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.473625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.473625Z digest=sha256:430954323bea0cc655a72fffbf67b69b42c703c5a3babef33b3cb963a6f7623d

Observation 5c3b2190-4ce9-47b0-a0d7-b9853fab6462 · outbound

This paper cites Position-Enhanced Visual Instruction Tuning for Multimodal Large Language Models.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Position-Enhanced Visual Instruction Tuning for Multimodal Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.479999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.479999Z digest=sha256:cf254456037391a318e0414adbd91b0c530f0090a4d6345e4325ba89e7004392

Observation 00bcfd2e-6a96-43f3-98ff-5d91c0ccd2c1 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.485233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.485233Z digest=sha256:cd037bfc0ffc5810658593a039b0e4faed935f822bb9f5e8f4626f7bd04fe248

Observation 7dc47472-2d1e-449e-bd59-8cc2b0e2ff1c · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Imagenet: A large-scale hierarchical image database

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:42.035747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.491417Z digest=sha256:61678b34b34f96a36839c3a8fd558477a69f51062701836fcd2b7666ab66ab73

Observation 60d252de-423d-47a6-b69a-11144dd6f1e5 · outbound

This paper cites Imagination improves Multimodal Translation.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Imagination improves Multimodal Translation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.495802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.495802Z digest=sha256:d1315569c9f6e828cadf9f3d9ec8dd66e2c207db269320e9abf7b1dde2180089

Observation 1fde9781-d160-4f1f-a4bd-0c8bf1502efa · outbound

This paper cites Lasot: A high-quality benchmark for large-scale single object tracking.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Lasot: A high-quality benchmark for large-scale single object tracking

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.924766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.501146Z digest=sha256:dbfe43d06ecbaee48ecc7c7295c89b5bb33eee93558d401d627071319c264f45

Observation 883fad47-9807-4929-b45c-d6b1ae557098 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.506566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.506566Z digest=sha256:31bf25c0cd0473d2aa3af43359c6e5f42210fc63471c53a67eef1ea911451fd0

Observation f5a77d2f-59d6-410d-9198-5b1d8242d2c0 · outbound

This paper cites Mme: A comprehen- sive evaluation benchmark for multimodal large language models, 2024.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Mme: A comprehen- sive evaluation benchmark for multimodal large language models, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.886987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.511930Z digest=sha256:8e7f4e948e3eaaab5f1ffeecb15f5e913d6537550f0935a759e2ed3618338fc0

Observation e016a139-3425-4f18-bda0-1e84d9647eef · outbound

This paper cites Blink: Multimodal large language mod- els can see but not perceive.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Blink: Multimodal large language mod- els can see but not perceive

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.840776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.516576Z digest=sha256:d01a205afddd86be592a617616fa55174a525e4b3f51c66026e229d67b4c4a1e

Observation 2d0554e3-53a9-4c3f-a3e4-9f1a6ea74ed5 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Ego4d: Around the world in 3,000 hours of egocentric video

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.805142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.520798Z digest=sha256:b4ec708cea054bc4bb135a91326c2954ac24cea316c9803367741c71cdd13ec6

Observation 1138ccf4-0253-4ac8-90dd-b597bc4ccca3 · outbound

This paper cites mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.524366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.524366Z digest=sha256:3c6b9d510eeca2bcf8d98109dfd1fc02a39db8c2e4b69d41f3e778ad9ce532bb

Observation 01ccf045-043a-4818-bdcd-84f6255c5eec · outbound

This paper cites Got-10k: A large high-diversity benchmark for generic object tracking in the wild.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Got-10k: A large high-diversity benchmark for generic object tracking in the wild

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.771087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.528549Z digest=sha256:f57c16e479a2b5e323328e2e89e4899ec60d95ff75fab976f15e4b992f7ce7d5

Observation 3534631f-0115-4469-895c-735fdf42a91b · outbound

This paper cites Editing Models with Task Arithmetic.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Editing Models with Task Arithmetic

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.534642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.534642Z digest=sha256:08081373e6c654ad133de772e88003df1ef633cd1fac5d063b40956bb02c43f2

Observation 06bd7538-097a-486c-b90a-f7143064a9c0 · outbound

This paper cites Distilling Translations with Visual Awareness.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Distilling Translations with Visual Awareness

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-10T21:10:41.186432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.539237Z digest=sha256:1f5aac9278265cc6e94be0e8247c65b8af82016bb044ca75c3c5fda3094056bd

Observation ed272f2c-1cbd-4ec6-a8fc-a8f4f6cdef4d · outbound

This paper cites Learning to Describe Differences Between Pairs of Similar Images.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Learning to Describe Differences Between Pairs of Similar Images

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.544396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.544396Z digest=sha256:70a2df744432a4050675a50a44ae562b4bba9428cbe028067a8d80ac3e8ca1dd

Observation 410fa98c-719c-44ef-ac9e-3f9771177c89 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.549054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.549054Z digest=sha256:41cbdefd597731651ae706304b84ff7910d8ca43a7d31ca73f6d99ed7dd5abcd

Observation d45ef472-ecbd-4580-88e2-d8349dac5d33 · outbound

This paper cites Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.553252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.553252Z digest=sha256:a845717de416d088b519b2f2a5ca37f819539cc41770c1880237806072039d8b

Observation c3b7f030-d9f4-4756-bfb5-539380dd7744 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.758615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.557396Z digest=sha256:9de9d88b748fed2080d58c1f9b58366f1cf7e4f09c2167c022c42ffa082f9e6e

Observation 65a8985c-0b1f-4ee5-a539-8c7daf1c662d · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.740625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.561453Z digest=sha256:4a7a8866d8ae6135dbe6091691fe76b62dd229af2687a7a9fb755f9054f506ad

Observation b5e9e031-0165-48f8-832b-3d3e9d062d03 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.565649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.565649Z digest=sha256:f03cdb755c2dd2f4a682377f78cd9f4aebbad97ffcbd97012ed1d7e38fc7a597

Observation 8cd547a0-359c-4e75-9f9e-cf09f0146ca2 · outbound

This paper cites Seed-bench: Benchmark- ing multimodal large language models.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Seed-bench: Benchmark- ing multimodal large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.712584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.570266Z digest=sha256:ada41ca3ce357bd333bc247df6fe57e1b5d41972e1a01ab23250b4598d2884af

Observation a808c9b6-1b81-496f-8207-a2cd1df8271f · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.574358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.574358Z digest=sha256:25393e417896c424fff318303c89ffdc616e269406964adee9cdd0f3718a0264

Observation 941e3834-430a-4454-93d8-e918e571fb39 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.578987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.578987Z digest=sha256:3d93fcabdf2b13982573db83ff475b66544b893de3278b9f74c5b693d8df04b9

Observation 17aa2a2b-b946-41e0-b321-f816fb3d129e · outbound

This paper cites Ground- inggpt: Language enhanced multi-modal grounding model.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Ground- inggpt: Language enhanced multi-modal grounding model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.697747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.583608Z digest=sha256:50ab2ec1b61f0e14d6d59f2287c9171f007b4066962903ba0a13d4489c21bbfc

Observation d09c57db-89ed-448f-b9ac-663e846cc7cc · outbound

This paper cites Microsoft coco: Common objects in context.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Microsoft coco: Common objects in context

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.684162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.587587Z digest=sha256:f9838e4ab617befc02bf999ac5dd5dbf29f8287d19bf637463a56764e9ec3af0

Observation ef4b137d-6cb2-4f60-8e99-e537413b3477 · outbound

This paper cites MIBench: Evaluating Multimodal Large Language Models over Multiple Images.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.591924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.591924Z digest=sha256:5514d55eaf027571f7a0d17d8fa17f59d50a1a3be185ec7dcca2e4f15e2323d9

Observation 40107a42-142f-48c8-b58d-708475ffccea · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.595916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.595916Z digest=sha256:c5ac81ae607ca693868a3cef8acb56ae3ebe661fa31e1666830e02773f943cb6

Observation 95752fd5-6da9-4e5c-a2a0-1234acb3658f · outbound

This paper cites IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models IconQA: A New Benchmark for Abstract Diagram Understanding and Visual Language Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.602513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.602513Z digest=sha256:faa5d53c31a89e60c75938b6280fe4cc5a7f2bb89c4e73f14142a91816ff1090

Observation 31c2891e-4b05-4f44-8863-ec6f40c1da27 · outbound

This paper cites MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.608363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.608363Z digest=sha256:605f609855448be2ac28485a7646f022a06e77be6ddf5c5b5fd8613e6e158074

Observation 742f7802-8d86-4b36-8045-ea83bea9271e · outbound

This paper cites MOT16: A Benchmark for Multi-Object Tracking.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MOT16: A Benchmark for Multi-Object Tracking

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.613039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.613039Z digest=sha256:46467e35d691806159878c354cd72a70b4605a0fdbbe6561f393a866beff54c2

Observation f587323b-1d15-467b-88f0-0068b35d64a1 · outbound

This paper cites Trackingnet: A large-scale dataset and benchmark for object tracking in the wild.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Trackingnet: A large-scale dataset and benchmark for object tracking in the wild

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.659659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.617653Z digest=sha256:0eaaa27168a3517f7299aefec84e96280be190aa7fef8df53f0315b5f167d4cd

Observation b84aec49-bb5a-4504-ac30-4a3b0fc599c4 · outbound

This paper cites Robust change captioning.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Robust change captioning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.644602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.623135Z digest=sha256:079924c20f2c460de5825c39c300deac982e2578c43ae3ab28cae4200b43b169

Observation 619e72cd-00b9-4e6f-82aa-38fe3c3b4f0e · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.628644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.628644Z digest=sha256:343315959b19b6084693d9646fdad61f0fb0761ad46b99b20c2fe267196ef8fd

Observation ad02df87-c198-482c-9495-3ea34ba06f9a · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Glamm: Pixel grounding large multimodal model

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.629159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.636730Z digest=sha256:50e154674a3ba12587c335ae0046de5129f1b3f806b31bb1c7d8e852d6c009db

Observation d1afe61f-7315-4a17-8f8c-20df3bb1af93 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models High-resolution image synthesis with latent diffusion models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.584743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.641452Z digest=sha256:8f50575f3720e977f33a26ca1a78bd1d8e17724ba9c08a880a1a979c7b07d761

Observation f3269a34-7e77-4bc9-a885-2f59e438e7e0 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Objects365: A large-scale, high-quality dataset for object detection

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.554795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.645941Z digest=sha256:1b9d62de00899097c6bf4ea6ac93e296810f2a6ae58872cd7ccbeefab287ee5e

Observation ade021cd-7e6f-41ed-95f7-4adba5eae456 · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.650037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.650037Z digest=sha256:ef08e3335935aeee0adf7c056660edd58b155b94b526938a3812b67658b06da6

Observation ecfef9e5-880e-4904-957b-2da1627fd922 · outbound

This paper cites ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models ONE-PEACE: Exploring One General Representation Model Toward Unlimited Modalities

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.653859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.653859Z digest=sha256:2dd14729b736fd5c344c1af9a6ef06671b2904893fed555384525cc0175638ff

Observation b1f1c0e6-9337-4ff9-9ed0-b5fe29a87c28 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.657929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.657929Z digest=sha256:e2166920267680ecc15a9ca8f3d5cf9502144c80bac02aa534818ed5678ceb24

Observation 51560bf4-0335-4d73-864d-e466776232d1 · outbound

This paper cites Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Driving into the future: Multiview visual forecasting and planning with world model for au- tonomous driving

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.661810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.661810Z digest=sha256:7145560288fe69511317812527bd761bd4d394f6641fcc2f5201fb85a4a2456b

Observation 1d2bb530-79c7-4b0e-ba37-9d90cdd58618 · outbound

This paper cites VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language Tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.665767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.665767Z digest=sha256:e2964597fc38f5795c6e2af21b62fceff3ded097ed6fec60916f7c10637bdfb2

Observation f724ee3e-9150-46fb-baec-f086dd201a47 · outbound

This paper cites V?: Guided visual search as a core mechanism in multimodal llms.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models V?: Guided visual search as a core mechanism in multimodal llms

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.477961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.670408Z digest=sha256:0a0d242c8d3d947f12c4a3c9835e8f2bbefcae474591a142856ee5dbfb4fbe6f

Observation 154ba1ff-9d01-4503-94dc-3f7098f5277a · outbound

This paper cites MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MC-Bench: A Benchmark for Multi-Context Visual Grounding in the Era of MLLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.674859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.674859Z digest=sha256:4a79f06a677b2784efb759718be77de24bf66aa387ede8e674007a2295f921c9

Observation 2a6118e7-d5cb-4af7-b94c-e384b4ff413b · outbound

This paper cites Qwen2 Technical Report.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Qwen2 Technical Report

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.679693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.679693Z digest=sha256:d756cbef19368d48674d5aa6657130416993a16afacc7c910348959c474afa40

Observation 13b99598-df00-44b9-8e97-4beceb65f314 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models ReAct: Synergizing Reasoning and Acting in Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.684403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.684403Z digest=sha256:a95266985389df56d79809b9fb0631a9857998b5b9dbb033cfdea34698616997

Observation 229b63c0-920d-4bf5-8e54-ef2870373f59 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.688619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.688619Z digest=sha256:6503bf8fe3dae80e4600fc2220b1b293c1df0dddd45b7888684abdafda52909a

Observation 75e807b4-543c-4b27-8c86-056d1dee2615 · outbound

This paper cites UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models UReader: Universal OCR-free Visually-situated Language Understanding with Multimodal Large Language Model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.692936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.692936Z digest=sha256:67636c7e555af70342846f973b47a48a919cac18aa8f38a2a6c5a8a47abb04c2

Observation f755047f-3c75-4c69-b5b8-fe4b0da45ea8 · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.698570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.698570Z digest=sha256:a19349424175685635dcb3f2f187ce0cc0383885ce60a38f424f9ef5b54a85a1

Observation 54bc825f-9406-402d-9d7a-3aa8c46e3c24 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.703754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.703754Z digest=sha256:8d9d01be3bc5cf2d2ff5679f36923148b594c86060a9abbfa48d7a582bb6664f

Observation 407a7a7f-e09c-4a35-870f-65152987203c · outbound

This paper cites Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.708082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.708082Z digest=sha256:e2194c97f86bcefb1bdea843984f4445480196cb67cf5696dd70bee7e36dc1ab

Observation da96d517-cdc6-4713-b2fd-604f638a4924 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction- guided image editing.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Magicbrush: A manually annotated dataset for instruction- guided image editing

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T21:10:41.416987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T21:10:40.712271Z digest=sha256:52b416cf3e74db49eb47f9fdafebb77f8cd41bbf4575bd75970c1bd05d02f54d

Observation e611620f-9456-4039-991f-91af93396a1a · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 54

Resolution
malformed identifier
no resolver link, observed 2026-08-10T21:10:40.716524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.716524Z digest=sha256:4c349a5be62bc0709345e37634590874d63cc0344378509f9303c33dcb9c7e87

Pith citing papers

Observation 85eec8d3-62e5-4e27-9f0f-024909f90687 · inbound

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey cites this paper.

Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 135

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:18:53.209391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T17:18:52.996467Z digest=sha256:016afb4fadb8d95f8bc808d7bf1b57e4570ff6bba46b9f74e668894fcb9ba51c

Observation 0e364c4d-9564-4082-bdab-640e6b79ba5c · inbound

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions cites this paper.

Unveiling the Potential of Vision-Language-Action Models with Open-Ended Multimodal Instructions Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:28.455317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:01:28.455317Z digest=sha256:9e673a9015803f76bd307e9d7b0274182c8d52af4d6b0ee8f55a3fda554cd6b3

Observation e55e4bdb-d8b0-437a-9897-698b0016b6da · inbound

MedSG-Bench: A Benchmark for Medical Image Sequences Grounding cites this paper.

MedSG-Bench: A Benchmark for Medical Image Sequences Grounding Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-15T20:50:47.809142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:50:47.809142Z digest=sha256:a3a40eacb9899ff617b7da95e6e1ca621f1cf40f19ca4ea6397f2a4d4405cec9

Observation 48ecc93c-134f-4f3b-bb96-9ec944c1ce57 · inbound

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning cites this paper.

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:52.218083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:52.218083Z digest=sha256:5f48a60e834f123c4255f0f0100a284bb93b30f6bcd296697cb87d038897199e

Observation 2274912a-62dd-498c-9c8a-40ea3cd393f8 · inbound

PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning cites this paper.

PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:13:34.937970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:13:34.937970Z digest=sha256:755afa2aaa812bf589aa534af16c9b312a8fa1c77bafd78e984027944cff76de

Observation 9940576a-4182-421d-b421-9bfacf7d6405 · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:52:08.087721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:2ca92f0b7a1c95a038bffb12330398d3608e0a16bb8078fda18b35f9a6801896

Observation 63fd2520-a011-4839-b7eb-25fdabc0b8cb · inbound

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension cites this paper.

RefBench-PRO: Perceptual and Reasoning Oriented Benchmark for Referring Expression Comprehension Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T18:15:05.561502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:15:05.561502Z digest=sha256:08781343333228d8a520404a5af8aeefac10806bdf93e3aabcd01a94ca126ade

Observation 98d23dba-4d73-42f7-b1e8-54a3b0e3bf51 · inbound

Training Multi-Image Vision Agents via End2End Reinforcement Learning cites this paper.

Training Multi-Image Vision Agents via End2End Reinforcement Learning Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:01:24.328661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T00:59:28.618477Z digest=sha256:0aa6d81a14d1b10d89d065890a1662b9f218d2c481fd59cb96eb52a09633f1ee

Observation a417a851-0566-48ec-9742-9f2b06906b2f · inbound

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding cites this paper.

CGC: Compositional Grounded Contrast for Fine-Grained Multi-Image Understanding Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.325270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T12:26:01.568507Z digest=sha256:4360e43dcac33fe0f455ff9959e6ddb471ca689daff432f5dbcb9123fe435ba0

Observation 2da49c6c-3e9a-4b2f-929b-abefa8092792 · inbound

Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning cites this paper.

Does Seeing More Mean Knowing More? Mono-Anchored Advantage Normalization for Multi-Source Visual Reasoning Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:54:00.616287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T22:53:55.407461Z digest=sha256:c823f5332b041c47e199e41cdad6be9246b29ffc0850c03968e0da4a78ac4ef1

Observation ba7be77d-951f-473d-bd51-a39bc7e96394 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 159

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.029033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:eb8bb8498ca2bd2aa658209f080e18f1d8ad6d33086539b73a3e7d221a9662c8

Observation 9aaaefa8-de82-4fd4-af0d-07d8ec5d2d9a · inbound

DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues cites this paper.

DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:19:50.963105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T05:18:01.931929Z digest=sha256:499bc16596cf980b136d9ab1a0454ca135a7187a20e38e0e57eeffb9b6894c62