Pith. sign in

Paper Citation Record · LEDGER

Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2206.08916.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2206.08916 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 28 of 28 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:44:58.818749Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:44.037310Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a1f6a5bc-6a6b-4eff-9270-32a7c46fc751 · inbound

PaLI: A Jointly-Scaled Multilingual Language-Image Model cites this paper.

PaLI: A Jointly-Scaled Multilingual Language-Image Model Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 198

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:29:06.215627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-16T09:29:05.956863Z digest=sha256:0d040b94208553c16516d7fa94bb24759a1cb9477ada3072772cc28da45c2608

Observation 5af69b53-b45f-4708-8630-98950a2d779f · inbound

Objaverse-XL: A Universe of 10M+ 3D Objects cites this paper.

Objaverse-XL: A Universe of 10M+ 3D Objects Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:02:11.638947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T13:02:11.512409Z digest=sha256:08d7c54a9f0e3447438f3e4ed85f937ffbcfeee735c5e74d89250e4634298e37

Observation ae304f91-df2e-44eb-85c5-d36eacc2c498 · inbound

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration cites this paper.

mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:18:51.723093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T03:18:51.582340Z digest=sha256:c643737bfa0418303d1d7ecc635fbf52982fd4ccb1723d82f4c6886813a8205e

Observation 57d8c058-f261-4e64-bae7-caa00386527f · inbound

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation cites this paper.

Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:09:16.941469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T22:09:16.622717Z digest=sha256:094791fc45e425145a1b8ba6062ae82db2e2a2bc9adb04d6c6f8e346af799159

Observation 18821cef-6e58-4eb4-af69-e4f3a8c919e3 · inbound

SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living cites this paper.

SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T04:44:58.818749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:44:58.818749Z digest=sha256:61b9a66f197951e60baae22bf1e53c922968151afbd5ae7ac54bb874ac7f8f8e

Observation 8a54fbf0-55d9-4f9f-b408-f2d3202d4874 · inbound

Vision-Language Models for Edge Networks: A Comprehensive Survey cites this paper.

Vision-Language Models for Edge Networks: A Comprehensive Survey Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T12:20:08.262482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:20:08.262482Z digest=sha256:8e08d32256a05335500d5405ec692614fc079ba39afe92130401a3d4eeec69b6

Observation a974054f-a8f6-4b30-9fba-5ae051fc1008 · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.204865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.204865Z digest=sha256:6249210b8166931020797738e35580d2550017fa90f1ae321d6ba486afadb0b0

Observation e685d759-ba8f-47d8-b3a2-b88bacd1eddc · inbound

LlamaSeg: Image Segmentation via Autoregressive Mask Generation cites this paper.

LlamaSeg: Image Segmentation via Autoregressive Mask Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:55.397888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:55.397888Z digest=sha256:7c0ded33c9df7064e029513fc5ca52269a1b2bcc7415b45b3801f5f74562b9a1

Observation 083b6baa-5d7f-478c-88d3-a32dcd1b8eb4 · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:01.991113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:01.991113Z digest=sha256:c5d46e9d24c790d63ef4928a4b1178f00610c6ca563b4c1c90f8e61c13197118

Observation 529958fe-3f15-4769-8f18-a9bcdd368ede · inbound

Is Extending Modality The Right Path Towards Omni-Modality? cites this paper.

Is Extending Modality The Right Path Towards Omni-Modality? Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:36:40.825255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:36:40.825255Z digest=sha256:707d34ddbf10b1b2a0fdd1606287a0ba64e6156bbefa3ee53a8e4e5fdb9d943a

Observation 0330e435-b54c-45b0-8640-3fb3636ff3be · inbound

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens cites this paper.

FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:26.091512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:26.091512Z digest=sha256:b50372f9f3a9847aaa0b9ebccedfb259ec566ab77b94778518b5e32b70adb35f

Observation 8129eeaa-5041-402e-90f5-2c55f99f7835 · inbound

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks cites this paper.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.505625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.505625Z digest=sha256:a15fe782a7f63a158a587b747e6a40c34a74d83f1748f18bf01135ab7d6964b6

Observation 446505ae-d8ab-4536-8a30-77cdf3a0a24c · inbound

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation cites this paper.

MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:02.310361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:26:02.310361Z digest=sha256:f01b289b02cfa212e2c28fb2bd1ffd9680b16d04c6484e4b81dedcb092bb78b1

Observation 71529ebe-e13d-4b3c-9d1c-94d7a2f340da · inbound

Vision Generalist Model: A Survey cites this paper.

Vision Generalist Model: A Survey Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-07T04:43:51.884460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:43:51.884460Z digest=sha256:88427baadb4a93b6854940b8550cb7d6fcaca61f5cef06c243df1af35fad3afa

Observation 40c626f2-5ee3-4f25-95ef-a2428b0320ff · inbound

Unified Multimodal Understanding via Byte-Pair Visual Encoding cites this paper.

Unified Multimodal Understanding via Byte-Pair Visual Encoding Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:41:51.615891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:41:51.615891Z digest=sha256:d6855d464f9a012a7e4a5ebc28aa014bb1a01211cf574123249de698902c1544

Observation 5c500d54-2f44-47a8-bf42-e6aab6e177ed · inbound

Is Visual in-Context Learning for Compositional Medical Tasks within Reach? cites this paper.

Is Visual in-Context Learning for Compositional Medical Tasks within Reach? Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T21:12:05.931398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:12:05.931398Z digest=sha256:6cd2cd571d935ec33a844226f4ba8539f59ff53ca3270c591922771ed96e8a6b

Observation e546804f-2ced-4318-aeb4-ad53f3e54008 · inbound

Grounding Intelligence in Movement cites this paper.

Grounding Intelligence in Movement Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:25:27.623864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:25:27.623864Z digest=sha256:11fa82e6a756975262c790ad2c251463e412cb42fc610ba88b9888084c709f25

Observation 3d48944d-23b5-4973-ace7-88db099cecdf · inbound

Open-set Cross Modal Generalization via Multimodal Unified Representation cites this paper.

Open-set Cross Modal Generalization via Multimodal Unified Representation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:29.595479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:29.595479Z digest=sha256:64a996e18bb8ecda4619c680934b72f38b921e2910759178fd8668d304227045

Observation 86f3f575-5e5f-445e-8920-db1c4aa906ec · inbound

MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding cites this paper.

MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:18:43.931214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T00:14:56.268778Z digest=sha256:8380f02a07af38abcd8e8c866347b6aa138f46daf9dafdb6a39ab0bbcde0df49

Observation 4414eaf0-8606-4f5f-90d1-088fd1508bb6 · inbound

On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training cites this paper.

On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-03T07:53:48.871339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:53:48.871339Z digest=sha256:172135191f219348499d52943ff2084852b5f247362f432e50e23a2a884c1cc4

Observation de1818f7-d813-403e-b2d3-770f40479d4f · inbound

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens cites this paper.

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:02:19.639749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:01:11.880003Z digest=sha256:4463f02d93e26eebf1188d2dffb4065cf21007bd0ad06d993ef477b7f14567a5

Observation 12a1438e-57bd-4ec8-b060-cf3f501a0141 · inbound

Modular Sensory Stream for Integrating Physical Feedback in Vision-Language-Action Models cites this paper.

Modular Sensory Stream for Integrating Physical Feedback in Vision-Language-Action Models Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:46:09.717580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T08:12:14.321695Z digest=sha256:ec85797e79aeef94637c82b49bc400171374b6eeace01c9b84d7b1253535835e

Observation baf255d8-f9fd-46ca-8ed8-7e50cc722152 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 114

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:44.038647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:06cfe5474e7c200552376a5c456103b697cd2b1787541307585d8c2103b2b296

Observation bde73f9d-46ea-42af-ba62-e49c096c564c · inbound

JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation cites this paper.

JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:28:38.368478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T16:22:31.982484Z digest=sha256:d0d590d2abaed1bec180a344316f53d1d88dc10ecae7a0e8b28f7463dcbcf53d

Observation 564e6f96-593b-4fac-a887-9ee827fb2ff0 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 91

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:93637c7213ce2f38899668dff3ad810b1cdb42fe79cfd6b747e08410c7669f86

Observation 44a4d8c8-6f4d-4fcb-8ebf-b5d8d4230cd6 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 129

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:1f150c6a5996b67b08222d3758b05e9ca297ed38832a974ab2acd3f8c88bb44c

Observation f98a7a04-3648-44f8-889e-ca36a41f7dc5 · inbound

Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features cites this paper.

Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T23:40:55.041093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:40:55.041093Z digest=sha256:ef65c8a135620238aa94d5f7ab5f52ae57e23ddc150eabc6d0c9d52365f246d6

Observation 5792015c-7b09-4844-ad46-4aa3ae0fcff8 · inbound

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling cites this paper.

UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-31T22:25:12.401226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T22:25:12.401226Z digest=sha256:d90dd1ee12846b5135d588e63d093ee350109c356ec09b957e3f6fc5d34e43fb