Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 28 inbound Pith citation observations for arXiv:2206.08916.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T04:44:58.818749Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T17:18:44.037310Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation a1f6a5bc-6a6b-4eff-9270-32a7c46fc751 · inbound
PaLI: A Jointly-Scaled Multilingual Language-Image Model Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 198
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5af69b53-b45f-4708-8630-98950a2d779f · inbound
Objaverse-XL: A Universe of 10M+ 3D Objects Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ae304f91-df2e-44eb-85c5-d36eacc2c498 · inbound
mPLUG-Owl2: Revolutionizing Multi-modal Large Language Model with Modality Collaboration Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 57d8c058-f261-4e64-bae7-caa00386527f · inbound
Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 18821cef-6e58-4eb4-af69-e4f3a8c919e3 · inbound
SKI Models: Skeleton Induced Vision-Language Embeddings for Understanding Activities of Daily Living Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a54fbf0-55d9-4f9f-b408-f2d3202d4874 · inbound
Vision-Language Models for Edge Networks: A Comprehensive Survey Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a974054f-a8f6-4b30-9fba-5ae051fc1008 · inbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e685d759-ba8f-47d8-b3a2-b88bacd1eddc · inbound
LlamaSeg: Image Segmentation via Autoregressive Mask Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 083b6baa-5d7f-478c-88d3-a32dcd1b8eb4 · inbound
FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 529958fe-3f15-4769-8f18-a9bcdd368ede · inbound
Is Extending Modality The Right Path Towards Omni-Modality? Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0330e435-b54c-45b0-8640-3fb3636ff3be · inbound
FuseLIP: Multimodal Embeddings via Early Fusion of Discrete Tokens Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8129eeaa-5041-402e-90f5-2c55f99f7835 · inbound
MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 446505ae-d8ab-4536-8a30-77cdf3a0a24c · inbound
MADFormer: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71529ebe-e13d-4b3c-9d1c-94d7a2f340da · inbound
Vision Generalist Model: A Survey Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 116
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40c626f2-5ee3-4f25-95ef-a2428b0320ff · inbound
Unified Multimodal Understanding via Byte-Pair Visual Encoding Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c500d54-2f44-47a8-bf42-e6aab6e177ed · inbound
Is Visual in-Context Learning for Compositional Medical Tasks within Reach? Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e546804f-2ced-4318-aeb4-ad53f3e54008 · inbound
Grounding Intelligence in Movement Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d48944d-23b5-4973-ace7-88db099cecdf · inbound
Open-set Cross Modal Generalization via Multimodal Unified Representation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86f3f575-5e5f-445e-8920-db1c4aa906ec · inbound
MedGRPO: Multi-Task Reinforcement Learning for Heterogeneous Medical Video Understanding Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4414eaf0-8606-4f5f-90d1-088fd1508bb6 · inbound
On the Role of Depth in Surgical Vision Foundation Models: An Empirical Study of RGB-D Pre-training Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 2004
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de1818f7-d813-403e-b2d3-770f40479d4f · inbound
LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 12a1438e-57bd-4ec8-b060-cf3f501a0141 · inbound
Modular Sensory Stream for Integrating Physical Feedback in Vision-Language-Action Models Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation baf255d8-f9fd-46ca-8ed8-7e50cc722152 · inbound
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bde73f9d-46ea-42af-ba62-e49c096c564c · inbound
JointHOI: Jointly Generating Contact Maps Enhances Hand Object Interaction Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 564e6f96-593b-4fac-a887-9ee827fb2ff0 · inbound
MentalThink: Shaping Thoughts in Mental SVG World Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a4d8c8-6f4d-4fcb-8ebf-b5d8d4230cd6 · inbound
Qwen-Audio-VAE Technical Report Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 129
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f98a7a04-3648-44f8-889e-ca36a41f7dc5 · inbound
Bitcoin Price Direction Prediction via Regime-Aware Multi-Modal Fusion of Social Sentiment and Technical Features Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5792015c-7b09-4844-ad46-4aa3ae0fcff8 · inbound
UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.