Pith. sign in

Paper Citation Record · LEDGER

Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2312.17172.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.17172 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:36:43.975806Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7673fcf4-7502-4ced-8300-5b3a687e1033 · inbound

BLINK: Multimodal Large Language Models Can See but Not Perceive cites this paper.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.684545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:b3ac6e576ce31f7a5ef8f5027fa7ba042a1c3e7ea7fa3a8bb29d210bc33a5d06

Observation 5eecd7f3-3e5e-47de-9c20-db37a670b5de · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.134116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:af20fc07bf174e23947b8171210fc3c8e4b4820d357007943f6474f8f561ec3b

Observation 5d5281de-347b-4e2b-b2d5-9b7b6ef518cb · inbound

Chameleon: Mixed-Modal Early-Fusion Foundation Models cites this paper.

Chameleon: Mixed-Modal Early-Fusion Foundation Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:03:28.141894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T10:03:27.919346Z digest=sha256:0e9ec3f6ae8b2861d9fa44fd4d370c0c9a069e51608f53d6bdcf3598899c3fa5

Observation 4e97cf76-c5ff-429f-a534-2f068825d26a · inbound

Learning Spatial-Preserving Hierarchical Representations for Digital Pathology cites this paper.

Learning Spatial-Preserving Hierarchical Representations for Digital Pathology Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:23:39.798726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-24T00:20:31.729964Z digest=sha256:6799dd137e77bc737cb878cec34b5278614339a0d5f0909d383710dc9a000331

Observation 4d64edaf-409b-4112-821c-a87bb58ef97e · inbound

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding cites this paper.

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:09:30.405274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T01:09:30.360275Z digest=sha256:6ff8b9973791097e9cf41d63171f8c73de9b61038b439f1ea0406e4c3e6cd5cc

Observation 35db2fdd-d625-4dea-8c92-b57aef434d7b · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:20.629194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:2bfeed01b0c8640343c7c3f5b2e9d0a1bbc9773e57bdc6d8f7ffd32a098e9e41

Observation 13875ab8-b8a4-4864-9934-86a0932c7482 · inbound

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching cites this paper.

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:43.975806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:43.975806Z digest=sha256:92af5c8259af4302c56f52ff4237ef595df44fc9abaad633c2af4cb3543d5328

Observation cf89c0f2-b74a-44b5-9376-208df931765b · inbound

One Diffusion to Generate Them All cites this paper.

One Diffusion to Generate Them All Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T13:19:54.679704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:19:54.679704Z digest=sha256:64a4f441caa8ce5574553d19fd83ab633140d9a43ce979b0f7c812cfdc256e2f

Observation e8ef846f-071d-402c-b08e-c4830fcd9b60 · inbound

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation cites this paper.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.385354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.385354Z digest=sha256:af2d4a7c621e63efa50c162ff69c0030b860b64286449ebf60e0fe60d270f0f3

Observation 4bb96c3a-44ef-4fd2-8a33-8135c6256bc9 · inbound

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents cites this paper.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.116529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.116529Z digest=sha256:7bd72a608c6373741966f3c582e1160c32c05ba19600f1d60919706faefb7905

Observation 237b7a46-057a-4fcd-8767-6c3bab50669e · inbound

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants cites this paper.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.958991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.958991Z digest=sha256:91978f3df3a11095e8cfa8305f73dadf7e9555291c29b563464eafde1c0556aa

Observation ce4e3727-b7e0-45e7-b3ad-69a5353cf09e · inbound

Next Patch Prediction for Autoregressive Visual Generation cites this paper.

Next Patch Prediction for Autoregressive Visual Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.051512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.051512Z digest=sha256:10019799c12c5b4db19a818004103ea9e733a4c3ef87463beb2c02ce8883eabb

Observation 31682340-93b8-4c4c-83fc-ac2dd32318cc · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 276

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:02.354880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:02.354880Z digest=sha256:de49d434f07114feab73c65e7c18fab07bbac73292bb4701c274ae5d57ef5199

Observation 51954eab-9ef7-4b15-87e6-d813b9c6c99d · inbound

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design cites this paper.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.060895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.060895Z digest=sha256:a3f1f90c32c5a93dcd7c3520a64e304f7af0fd5e63dcf3bfacef0079865c3fa5

Observation 3bb90c00-50f0-47ca-b7e3-50a3fd2f7d87 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.170183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.170183Z digest=sha256:42821802f285893c0651f76dff6247518e950e19d74878458d52dfa89eb6a109

Observation 49637e00-5f91-43ca-929f-2b7bf399ef2e · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:16.073781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:78799a4d75252775419acc0de091cd552b01b60d47d0033247c2514f140a63f4

Observation 2e351de5-1f16-41ee-ba07-365d0832fabb · inbound

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan cites this paper.

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:49:51.545721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:49:51.545721Z digest=sha256:b32bab9ddfc1214e183472d81dd3e31a73bd89046c851e421ace4bc3ece04e31

Observation d8b909db-3d64-4012-8b9b-39ba0cbf3a0f · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.304736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:2f33b895202d338a8f648fa7eeb4a7ff35d66489c1eb80851d2ad27f4f15211e

Observation b5181804-620d-4b7b-a402-dd4c8fec86f9 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.290588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:e35cbccb449352d92918f6edc38bf49c539ef961497edd50a8a644f5b17b4109

Observation e23884c3-5ce4-479e-8c1d-6ff67692997b · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:4cc5eb803f94b159238197ff67dbf699e513fee931852b43fbed25fb9c603e8a