Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:16:21.519000Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 5 inbound Pith citation observations for arXiv:2506.02975.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:16:21.519000Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-01T05:57:54.653504Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T14:38:28.862952Z
70 of 70 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a13f4005-cb86-45a2-a717-72beb0583d9a · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79264a30-30f2-4191-9e39-c409ab5016f2 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f3a5d2b-7a8e-4831-b203-41a29c957e94 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: IEEE International Conference on Computer Vision (2021)
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afc1382b-fb08-467a-9ef1-24cacf15ea8d · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fa59cb0f-beef-4c50-a460-976dbf4e9bdf · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7be0a79-3ea8-415c-aefb-68673c5eedbf · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b32ebac2-456b-43ad-be0c-e6e28bd0992e · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5fc3d08-eb27-4049-9d88-beff913c26ce · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 44b6df66-8a4a-4019-963a-265169d7be78 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: CVPR (2024)
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e38fb6ad-af18-409c-8cd7-4b9980fb91f2 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e6550da-6808-44b6-b291-a2efcd82aa39 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: NeurIPS (2023)
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9713abf4-f732-4f4c-9f1b-cdedf42b3fbf · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Unveiling Encoder-Free Vision-Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83d4e62d-1993-41be-9d4e-f41215657048 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation DreamLLM: Synergistic Multimodal Comprehension and Creation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25abc879-463d-47f7-9273-9f48f651a92b · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a391fc56-3afb-4c82-85ca-037994c666d1 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation The Llama 3 Herd of Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d56cd85d-821c-4c2a-ab6e-384756e61da4 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0976f8d8-0e68-4cb6-92a5-4d4f8be446cc · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE conference on computer vision and pattern recognition
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b7e2b992-0987-43d9-850d-72b70be437e9 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea726f6-3011-4dcc-9cf2-46f5dfc35b3e · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6dc66f1-3ea0-4975-8326-04d175955325 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Advances in Neural Information Processing Systems35, 8633–8646 (2022)
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bea976c4-af4f-48e1-bc3d-914e52d7c605 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8237ec16-8a61-4426-9b6b-9ed4d31c10f6 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 52fc6de1-7fce-4994-95ac-e888d8261ee7 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aa920d3a-5679-4c4f-a214-4711302fe5e6 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation arXiv preprint arXiv:2402.03161 (2024)
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 092f870b-41a7-4c8f-9949-328d1eafa47d · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c0e42125-a9d0-4ca7-9443-6c089dcdadfd · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation https://klingai.com/ (2024)
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 32e5afa7-020c-484e-bbc9-a13bd6dfb8be · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 180e0478-7161-47f9-94a3-1f6d0ac71469 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cea17f4b-ab73-427e-bddc-12dcdbe5bc59 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2eb70775-e30e-4e55-9c11-d5faa1b38a0e · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: European Conference on Computer Vision
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7642fe84-a521-4794-a55e-d0ab25927c10 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Evaluating Object Hallucination in Large Vision-Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca173804-0844-4389-9719-ce8ec513b164 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Open-Sora Plan: Open-Source Large Video Generation Model
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7bbd53c-99a3-4dcc-8578-8cf3b081f682 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2489a54f-585c-4fc0-9818-7f6660f54ee1 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b18aed56-303a-4aa4-8578-05fca8648344 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation arXiv preprint (2024)
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a476b683-fe59-4762-85af-ee9bcdb361f8 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: CVPR (2024) 13
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0e461a0e-472b-44d6-a8d5-62d4c0cc5a83 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Unresolved cited work
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a13691e4-453e-487d-ba56-7438da341eeb · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: NeurIPS (2024)
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b2d9467b-2076-4c4a-8f5c-6afb0dd7a594 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Decoupled Weight Decay Regularization
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 168f2578-0999-4167-a20c-7dc6c21f0002 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Advances in Neural Information Processing Systems 36, 46212–46244 (2023)
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 03947037-f2dd-4076-9ce9-0b9e7098a129 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 789c8f16-7dce-45b8-aa24-df7044f9d3f0 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Advances in neural information processing systems35, 27730–27744 (2022)
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1a4ad03-ae42-422d-9f49-35f5d75f7aad · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF International Conference on Computer Vision
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c4ba381-1719-471b-b545-5aeae0ccb0cb · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation https://pika.art/home/ (2023)
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7bdd302f-5511-4610-8466-56f2b6827031 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation deba6181-a615-42f0-841b-d0ff73d2e289 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f80f9974-b276-4976-a819-7aea019853a1 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation https://runwayml.com/research/introducing-gen-3-alpha/ (2024)
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6f6a0211-05a8-4d32-92f6-9f6446e0342f · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Advances in Neural Information Processing Systems36 (2024)
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3aae1503-de1f-4f03-ba0e-9a56b5dba7b6 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Denoising Diffusion Implicit Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56294852-2b48-4862-9837-13a766e2e86e · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Advances in Neural Information Processing Systems36(2024)
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 407cd4e8-149f-4c23-a819-f973fd0cf18f · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 05833d42-5a6a-400a-88d8-e531f700e712 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ce96e01-6308-4b66-bd01-e4a105341fdd · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Emu3: Next-Token Prediction is All You Need
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09b87698-aeaa-4141-b4d2-80968a4d8457 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 461ee106-6d4a-4934-8cdb-427e0f7fba8f · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc9f6bf1-9567-4db4-b3a9-4b3b70ef090d · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1cae4a7-b6dd-41d1-b017-3f8e3b2bbce5 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation OmniGen: Unified Image Generation
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55738e0d-c946-4f45-88b5-97084329bb43 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1b66d188-aedf-42e4-9b63-f192e703d310 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebe99975-196b-4298-881d-08125ddea5b5 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Advances in Neural Information Processing Systems37, 75329–75354 (2024)
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ba7731e0-5a26-4d91-a2bf-54d9b040e728 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b3c9276-4fe7-490c-8ff4-69da8e1198d5 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Qwen2 Technical Report
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cda4bc13-1d3f-49b5-a92e-af949630442a · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Qwen2.5 Technical Report
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c9854a8-3074-4a3b-9afa-01ab99df80f4 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43715fe0-ce7c-4853-afe7-a83dde80beaf · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65eca0ba-72b5-4989-9cff-db240952f1ab · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d2090dbf-d07f-4626-8914-5e8024d52412 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 379e6b75-2222-4fe3-b221-6e2c9c7a5844 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05d6bf8d-edad-4a8c-ae4b-ed5e8185b22f · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fb1e717-4b06-4838-a2d0-2521d0f7bfa3 · outbound
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 965001ce-327d-438c-9039-a5e6a168f8e6 · inbound
Show-o2: Improved Native Unified Multimodal Models HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
Reference 124
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 457ff10f-9cbf-418a-bfb6-e055bc7015f5 · inbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
Reference 133
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0ccd6c77-5403-4c12-897d-a4c838967cf7 · inbound
Lance: Unified Multimodal Modeling by Multi-Task Synergy HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
Reference 134
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b7a9c49b-0821-4690-9c9f-c2d70ad77567 · inbound
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
Reference 206
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ed9aaa69-983e-4bc4-91c5-f14b6ab72a28 · inbound
Bridging Video Understanding and Generation in a Unified Framework HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.