Pith. sign in

Paper Citation Record · LEDGER

Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2312.17172.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.17172 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:36:43.975806Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

3
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7673fcf4-7502-4ced-8300-5b3a687e1033 · inbound

BLINK: Multimodal Large Language Models Can See but Not Perceive cites this paper.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.684545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:f95796d3694ddb92ea9499a8f5ce24fc31895033c7470471eebd102c55682875

Observation 5eecd7f3-3e5e-47de-9c20-db37a670b5de · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.134116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:38d57a54a8eb0f84dbf9738c00a49e6b4b67a0fc881511688b3782ce8aa06dd0

Observation 5d5281de-347b-4e2b-b2d5-9b7b6ef518cb · inbound

Chameleon: Mixed-Modal Early-Fusion Foundation Models cites this paper.

Chameleon: Mixed-Modal Early-Fusion Foundation Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T10:03:28.141894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T10:03:27.919346Z digest=sha256:1870fc0d8442c9f878ed1f32cfb784c0ab1b9dc46cb064f3121421247a206151

Observation 4e97cf76-c5ff-429f-a534-2f068825d26a · inbound

Learning Spatial-Preserving Hierarchical Representations for Digital Pathology cites this paper.

Learning Spatial-Preserving Hierarchical Representations for Digital Pathology Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-24T00:23:39.798726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-24T00:20:31.729964Z digest=sha256:ce37ca9dd131ea07e63deaa11d20df8a20de4ea7020b8e35382bdd112ea32c51

Observation 4d64edaf-409b-4112-821c-a87bb58ef97e · inbound

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding cites this paper.

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:09:30.405274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T01:09:30.360275Z digest=sha256:157e5faab1d1f2274b1c619446921fcb6fbb90dba8781dd28bd7a269663f8e80

Observation 35db2fdd-d625-4dea-8c92-b57aef434d7b · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:20.629194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:c9c13ec09c46af7672c0944fe598db1a4bf4ae72bbf426d310c4cb9577f835c9

Observation 13875ab8-b8a4-4864-9934-86a0932c7482 · inbound

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching cites this paper.

BlendServe: Optimizing Offline Inference for Auto-regressive Large Models with Resource-aware Batching Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T13:36:43.975806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:36:43.975806Z digest=sha256:589c3754393b345c64ff2cfec0173a4b37be24791838be00efcec645d84d98f5

Observation cf89c0f2-b74a-44b5-9376-208df931765b · inbound

One Diffusion to Generate Them All cites this paper.

One Diffusion to Generate Them All Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T13:19:54.679704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:19:54.679704Z digest=sha256:6f2d4129a8c92d3f79c6ec04c0b7be31d387bb24fad8dd0ac7d9d036f2b7f0a0

Observation e8ef846f-071d-402c-b08e-c4830fcd9b60 · inbound

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation cites this paper.

Divot: Diffusion Powers Video Tokenizer for Comprehension and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:08.385354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:08.385354Z digest=sha256:0ca7c22f59dd662c345e9b3f53c3a66ec6465edbd3cbf00bc5602d321527f489

Observation 4bb96c3a-44ef-4fd2-8a33-8135c6256bc9 · inbound

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents cites this paper.

GROOT-2: Weakly Supervised Multi-Modal Instruction Following Agents Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T20:41:45.116529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T20:41:45.116529Z digest=sha256:246877be98854daaf9f1a394654f0da16020eee40fd9b068679506b025e7ff27

Observation 237b7a46-057a-4fcd-8767-6c3bab50669e · inbound

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants cites this paper.

MedMax: Mixed-Modal Instruction Tuning for Training Biomedical Assistants Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T13:53:57.958991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T13:53:57.958991Z digest=sha256:69009cba29c3465be8f7d17b1c562a4d0df55d24f775f7ac240d1a66ecacacab

Observation ce4e3727-b7e0-45e7-b3ad-69a5353cf09e · inbound

Next Patch Prediction for Autoregressive Visual Generation cites this paper.

Next Patch Prediction for Autoregressive Visual Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T11:37:42.051512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:37:42.051512Z digest=sha256:42d983b9ad155898a52dd764b32fcb2715cf9f3c7755da902697e4a079b7769b

Observation 31682340-93b8-4c4c-83fc-ac2dd32318cc · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 276

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:02.354880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:02.354880Z digest=sha256:51ff37702eda911544c54ab8138ec003cd73e67d6bf5569090aa6bb9083cf890

Observation 51954eab-9ef7-4b15-87e6-d813b9c6c99d · inbound

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design cites this paper.

Valley2: Exploring Multimodal Models with Scalable Vision-Language Design Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:19.060895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:19.060895Z digest=sha256:a036ff51a2d1956292dc74288b1cf7dc517c76b5f1e5945a13a53d85dd1248c1

Observation 3bb90c00-50f0-47ca-b7e3-50a3fd2f7d87 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.170183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.170183Z digest=sha256:4b619f0695e50035108a7ba3e8697283d6b25acf63cf38da3e3b5f695b9947c6

Observation 49637e00-5f91-43ca-929f-2b7bf399ef2e · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:16.073781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:f10ae26874f13f29eb0d52d5ae3bde1dcbc315933521397bfea9a71749f04747

Observation 2e351de5-1f16-41ee-ba07-365d0832fabb · inbound

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan cites this paper.

Multi-TW: Benchmarking Multimodal Models on Traditional Chinese Question Answering in Taiwan Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T05:49:51.545721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:49:51.545721Z digest=sha256:5dcbf8989e996b6b8bfd743b7795af79eacefa6b3adaf1f7e2ac790d5851f0c1

Observation d8b909db-3d64-4012-8b9b-39ba0cbf3a0f · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.304736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:f8bd1357d3c19beaf4dd8b8bd56c3e06402d1752eda682cf21fdfce8f74a2dc8

Observation b5181804-620d-4b7b-a402-dd4c8fec86f9 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.290588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:b4f0f2b77d59e248c68169543beeeef6f9e9e0e4ab88ee4bc5e3ff6544e35956

Observation e23884c3-5ce4-479e-8c1d-6ff67692997b · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 92

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:038d0d7e5ec2973aab1e21a9e964043f3fbd20a849b54d5078160a7c0a594c1e