Pith. sign in

Paper Citation Record · LEDGER

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation

As of 12 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 5 inbound Pith citation observations for arXiv:2506.02975.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02975 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:16:21.519000Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T05:57:54.653504Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T14:38:28.862952Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved48
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a13f4005-cb86-45a2-a717-72beb0583d9a · outbound

This paper cites GPT-4 Technical Report.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:14.506946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:14.506946Z digest=sha256:597e384d6de140ac7cb0d7e7bcfb43e613a1c65eb5cd98e82e01da8747a0f62e

Observation 79264a30-30f2-4191-9e39-c409ab5016f2 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:14.572299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:14.572299Z digest=sha256:e380432a8ed0546b4510f462631b2d06a26edd5cf0ae88628850ca9390d549a0

Observation 1f3a5d2b-7a8e-4831-b203-41a29c957e94 · outbound

This paper cites In: IEEE International Conference on Computer Vision (2021).

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: IEEE International Conference on Computer Vision (2021)

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:14.659832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:14.659832Z digest=sha256:c7220be6cb9979931bc6821fb9accdc19e420f4ff33f388e43ae1eb719fbf29d

Observation afc1382b-fb08-467a-9ef1-24cacf15ea8d · outbound

This paper cites an unresolved cited work.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:16:27.013945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:14.784115Z digest=sha256:57608aaae525cea9255cfaab80c6586f754607967c8fef6ff987caa6b14d6c5c

Observation fa59cb0f-beef-4c50-a460-976dbf4e9bdf · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:14.922831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:14.922831Z digest=sha256:3055cf716b2803e235e8c0a12890d2556e01f78204627e98c30608a0c8345a27

Observation b7be0a79-3ea8-415c-aefb-68673c5eedbf · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:15.006996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:15.006996Z digest=sha256:9414582de9176cebef39e541ada1b19b28b8c6c86eb46865d135b1ad780a960b

Observation b32ebac2-456b-43ad-be0c-e6e28bd0992e · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:15.089684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:15.089684Z digest=sha256:0d614a0d394681ecd6c994fc48d2c0e2f306e46d176c9b5c85bbb5df4ca92ce6

Observation d5fc3d08-eb27-4049-9d88-beff913c26ce · outbound

This paper cites an unresolved cited work.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:16:26.856003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:15.228633Z digest=sha256:d0dce0543bccd81f71fd8e2a80da4891308209f8afcb1463735fe49c28decfc2

Observation 44b6df66-8a4a-4019-963a-265169d7be78 · outbound

This paper cites In: CVPR (2024).

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: CVPR (2024)

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:15.334960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:15.334960Z digest=sha256:88907bbd3299d35fb7d3f09b30d16a8b45a7994630aa1ff6d5b05754b4bcf8e1

Observation e38fb6ad-af18-409c-8cd7-4b9980fb91f2 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:15.447091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:15.447091Z digest=sha256:c4ce9b9cc8878ede952fa73ad043522d5bcada58415c27055b413040e0572cbb

Observation 0e6550da-6808-44b6-b291-a2efcd82aa39 · outbound

This paper cites In: NeurIPS (2023).

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: NeurIPS (2023)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:26.696605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:15.578839Z digest=sha256:020240d68db15a6cfdf182d3ce9638baac3dfe5667fcc5f95ac0f2704e96323c

Observation 9713abf4-f732-4f4c-9f1b-cdedf42b3fbf · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Unveiling Encoder-Free Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:15.700721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:15.700721Z digest=sha256:08a5df8a51c19a5d19a194b5663a8fa1e0052c21ff5663adb657eba9ebeb07ea

Observation 83d4e62d-1993-41be-9d4e-f41215657048 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:15.792420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:15.792420Z digest=sha256:be9e4d6bf574602937df0b8325c0fa001e4c14d2c8b4d0d1d00d1515a3bf0f64

Observation 25abc879-463d-47f7-9273-9f48f651a92b · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:16.069504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:16.069504Z digest=sha256:6d9921353cbce24779a276a03ab65814be32743937232d062d4cceea74b4f500

Observation a391fc56-3afb-4c82-85ca-037994c666d1 · outbound

This paper cites The Llama 3 Herd of Models.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation The Llama 3 Herd of Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:16.243885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:16.243885Z digest=sha256:0d068c15cf223754f2c7d4605bf10529091ae30308e1c3e89ef61653e65dfd25

Observation d56cd85d-821c-4c2a-ab6e-384756e61da4 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:16.395393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:16.395393Z digest=sha256:83ece00768482a1bd6e5edc55f4ec6ae0615aef45b0b82f507bb38f1daec2882

Observation 0976f8d8-0e68-4cb6-92a5-4d4f8be446cc · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:26.505291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:16.536886Z digest=sha256:5f6be1063eedaa44bc0e28a786a66e3172f9733411fc691ea5d6f6be03144b01

Observation b7e2b992-0987-43d9-850d-72b70be437e9 · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:16.636718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:16.636718Z digest=sha256:d5153a3d31322c2c7b60e9769affddfcd4b4d820dce7be71021e6c94f6ca0cbe

Observation 8ea726f6-3011-4dcc-9cf2-46f5dfc35b3e · outbound

This paper cites Latent Video Diffusion Models for High-Fidelity Long Video Generation.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Latent Video Diffusion Models for High-Fidelity Long Video Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:16.736485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:16.736485Z digest=sha256:fd4ffa7e56c954ab993960ce4ff13cb33ccd938da8a5fb2cff2cdffa3bf78c39

Observation a6dc66f1-3ea0-4975-8326-04d175955325 · outbound

This paper cites Advances in Neural Information Processing Systems35, 8633–8646 (2022).

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Advances in Neural Information Processing Systems35, 8633–8646 (2022)

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:16.864103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:16.864103Z digest=sha256:eab90b22b09c6ec2c963b15b16d954afb2d2df120a50e02b7c0f86a5d4a801fd

Observation bea976c4-af4f-48e1-bc3d-914e52d7c605 · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:17.002270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:17.002270Z digest=sha256:b9e7abed3611087b18ad101b2a87dd33fb416f8c014cb1bb2b20dbf9e0b0bf6c

Observation 8237ec16-8a61-4426-9b6b-9ed4d31c10f6 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:26.290481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:17.079127Z digest=sha256:4c57a55f09e1537376b0267331499dc9419b3697d2dc2b8826096c7c4958e90c

Observation 52fc6de1-7fce-4994-95ac-e888d8261ee7 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:26.068827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:17.176950Z digest=sha256:22c56b2e2ad460e6b0afd81246ffc910c311c471a4f9cd75c581da5dbcaf7700

Observation aa920d3a-5679-4c4f-a214-4711302fe5e6 · outbound

This paper cites arXiv preprint arXiv:2402.03161 (2024).

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation arXiv preprint arXiv:2402.03161 (2024)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:17.282827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:17.282827Z digest=sha256:326eb56c33b992ff73f854d903f8090e91635e31e57a26e5fd6f91ba5ffafa94

Observation 092f870b-41a7-4c8f-9949-328d1eafa47d · outbound

This paper cites In: Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Computer Vision–ECCV 2016: 14th European Conference, Amsterdam, The Netherlands, October 11–14, 2016, Proceedings, Part IV 14

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:25.902895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:17.420701Z digest=sha256:baec84ea93943c8777edb071d589d73543650e43c41241ecaffa2bb26666daa5

Observation c0e42125-a9d0-4ca7-9443-6c089dcdadfd · outbound

This paper cites https://klingai.com/ (2024).

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation https://klingai.com/ (2024)

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:25.683493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:17.575158Z digest=sha256:35513b5c640dce97b2c4dc9d528a577cd80a4e79b89110a38750a9a854edce29

Observation 32e5afa7-020c-484e-bbc9-a13bd6dfb8be · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:17.717095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:17.717095Z digest=sha256:8624ce524f64195c2d370d31b1cf73566552c72f1bff1249f7fc695dfc22aef5

Observation 180e0478-7161-47f9-94a3-1f6d0ac71469 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:17.940410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:17.940410Z digest=sha256:d8c69c4a7f7b107c22aa3925b9e8eea461eb1982f95a45aa48c504e52a5104ae

Observation cea17f4b-ab73-427e-bddc-12dcdbe5bc59 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:25.494808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:18.034666Z digest=sha256:7071d8e016d0035442336defc1a341dd60abf279d46919731fe8b1a53a3c38d0

Observation 2eb70775-e30e-4e55-9c11-d5faa1b38a0e · outbound

This paper cites In: European Conference on Computer Vision.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: European Conference on Computer Vision

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:25.351624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:18.131456Z digest=sha256:96ec05dd6b2c58f220c77ac756b88ded3c06573f6512eb78a89f986189c89bdd

Observation 7642fe84-a521-4794-a55e-d0ab25927c10 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Evaluating Object Hallucination in Large Vision-Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:18.221076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:18.221076Z digest=sha256:5d90bfde77f8ac61a47aacd312e5224aa3adebd76241f6edafeeb88fa9a3dec3

Observation ca173804-0844-4389-9719-ce8ec513b164 · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Open-Sora Plan: Open-Source Large Video Generation Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:18.284392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:18.284392Z digest=sha256:fbd9e9987e4bd2052fd5ba524698a27e3c9f65b10cb29b7a6ab66100e91eb3f8

Observation e7bbd53c-99a3-4dcc-8578-8cf3b081f682 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:18.380392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:18.380392Z digest=sha256:9f9404ff5227e3242a9eb9708d2ae782abcc07f2f57b38dc0ffe66986be2a2f2

Observation 2489a54f-585c-4fc0-9818-7f6660f54ee1 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:25.163273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:18.454646Z digest=sha256:1731f9591a5fda75b65d0ff174be0e47eb4190ce7e691ac34693fad12590ec21

Observation b18aed56-303a-4aa4-8578-05fca8648344 · outbound

This paper cites arXiv preprint (2024).

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation arXiv preprint (2024)

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:24.939398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:18.559489Z digest=sha256:3085c8cc6ce02013c1c6a87d1548928352b3d5953ccc888c52d5cb6a18bd1f66

Observation a476b683-fe59-4762-85af-ee9bcdb361f8 · outbound

This paper cites In: CVPR (2024) 13.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: CVPR (2024) 13

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:24.743289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:18.629428Z digest=sha256:46e43f9e2ea94758ada31d5f3b67a54029b7ff55a927e289cc18a1d1b38a8b7e

Observation 0e461a0e-472b-44d6-a8d5-62d4c0cc5a83 · outbound

This paper cites an unresolved cited work.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:16:24.547264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:18.714499Z digest=sha256:8e78ec2b8c9a594cc781079f50cd4adc64e621b24c5fedb9042d3f1943d8e0ef

Observation a13691e4-453e-487d-ba56-7438da341eeb · outbound

This paper cites In: NeurIPS (2024).

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: NeurIPS (2024)

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:24.317407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:18.841664Z digest=sha256:fe5e28662fc75967b92d306cf9278c52903b083ed264f1d66fce4da984339a92

Observation b2d9467b-2076-4c4a-8f5c-6afb0dd7a594 · outbound

This paper cites Decoupled Weight Decay Regularization.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Decoupled Weight Decay Regularization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:18.935298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:18.935298Z digest=sha256:1755b3603087d9014e098bf8df0a417789133089d907d61a31c12081177cae3d

Observation 168f2578-0999-4167-a20c-7dc6c21f0002 · outbound

This paper cites Advances in Neural Information Processing Systems 36, 46212–46244 (2023).

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Advances in Neural Information Processing Systems 36, 46212–46244 (2023)

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:24.118268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:19.027439Z digest=sha256:6150b7475d8479b667f8e48416b27046862994c46c4b838074b50d2f14ff5319

Observation 03947037-f2dd-4076-9ce9-0b9e7098a129 · outbound

This paper cites an unresolved cited work.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:16:23.882413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:19.118945Z digest=sha256:daabf1fa9966af5349ae0aab78e34f135005992e8c9d107472fbec5edcd14371

Observation 789c8f16-7dce-45b8-aa24-df7044f9d3f0 · outbound

This paper cites Advances in neural information processing systems35, 27730–27744 (2022).

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Advances in neural information processing systems35, 27730–27744 (2022)

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:19.211614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:19.211614Z digest=sha256:a5f77f0f4decfde936fc6609e6bed170a5612522974db355c853104966293bd4

Observation c1a4ad03-ae42-422d-9f49-35f5d75f7aad · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:19.276511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:19.276511Z digest=sha256:9e07391169c9510fc586d965ade1eb5559309fc89ccbd241fd39b3c2d7e966f9

Observation 6c4ba381-1719-471b-b545-5aeae0ccb0cb · outbound

This paper cites https://pika.art/home/ (2023).

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation https://pika.art/home/ (2023)

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:23.695032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:19.360988Z digest=sha256:f294a5b7055a7d96d3e044285df1bdd31e6696cfb5f6a4bc8b7ce60af066fb41

Observation 7bdd302f-5511-4610-8466-56f2b6827031 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:19.466351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:19.466351Z digest=sha256:977de7b601ff16d6834625646bea1912bd65d2bcd2d30fa274466810b0dc91fe

Observation deba6181-a615-42f0-841b-d0ff73d2e289 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:19.588611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:19.588611Z digest=sha256:7ab3f8d703258497895026796ff9856eb306becf03bd2171b27f9d2436609b77

Observation f80f9974-b276-4976-a819-7aea019853a1 · outbound

This paper cites https://runwayml.com/research/introducing-gen-3-alpha/ (2024).

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation https://runwayml.com/research/introducing-gen-3-alpha/ (2024)

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:23.482110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:19.654917Z digest=sha256:77acc075d47cd29be279a8818b746030894961340bf5bfdd82595b639f8ec9c6

Observation 6f6a0211-05a8-4d32-92f6-9f6446e0342f · outbound

This paper cites Advances in Neural Information Processing Systems36 (2024).

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Advances in Neural Information Processing Systems36 (2024)

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:23.295844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:19.741402Z digest=sha256:475c8d2a8f6b07709cb2f9b25b8b4505869c77f919e9941d6b5c717353fff636

Observation 3aae1503-de1f-4f03-ba0e-9a56b5dba7b6 · outbound

This paper cites Denoising Diffusion Implicit Models.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Denoising Diffusion Implicit Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:19.828951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:19.828951Z digest=sha256:32c5c3dd9f6c8acddbc933023308a1219b35c1aee07b5c31546f144a460ab4f6

Observation 56294852-2b48-4862-9837-13a766e2e86e · outbound

This paper cites Advances in Neural Information Processing Systems36(2024).

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Advances in Neural Information Processing Systems36(2024)

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:23.074955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:19.939119Z digest=sha256:38d62d6778369fec3eef9349287c7c2019e43885c454579197bfe8225943094f

Observation 407cd4e8-149f-4c23-a819-f973fd0cf18f · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:22.863884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:20.033243Z digest=sha256:fcd8a80a6593cdaa7aed296afc7ed7f38d459917a7218f2cc0910efd754d3172

Observation 05833d42-5a6a-400a-88d8-e531f700e712 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.114731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.114731Z digest=sha256:b105640771ef4122d365fbaa5eb9dee4b6920f35f7aa93b187912aeafa1ad146

Observation 4ce96e01-6308-4b66-bd01-e4a105341fdd · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Emu3: Next-Token Prediction is All You Need

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.206272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.206272Z digest=sha256:fc0e2c47c82eae69881f26d437dac795594d5f0ddb74bac3d1e3592ea9b17205

Observation 09b87698-aeaa-4141-b4d2-80968a4d8457 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.271828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.271828Z digest=sha256:c7e18df547255f1796b2fb1efdfe41b6f2864e98a5ac9a966fb2ce888ca5fe12

Observation 461ee106-6d4a-4934-8cdb-427e0f7fba8f · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation NExT-GPT: Any-to-Any Multimodal LLM

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.351944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.351944Z digest=sha256:59c5ac430a3a8c0e25cc1e6affe25e833b54efe7f20cc3f2443c86fc7176d169

Observation dc9f6bf1-9567-4db4-b3a9-4b3b70ef090d · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.415168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.415168Z digest=sha256:1e30cadfed32b47b89a6dcab72d29c32fa5d4769d87429cbb91ea6362fc99571

Observation a1cae4a7-b6dd-41d1-b017-3f8e3b2bbce5 · outbound

This paper cites OmniGen: Unified Image Generation.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation OmniGen: Unified Image Generation

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.505845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.505845Z digest=sha256:2a7b9b24a2ffc580df3724c8d9b680c0d051122b46d3993feb76ee38c6f3dc0b

Observation 55738e0d-c946-4f45-88b5-97084329bb43 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:22.678176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:20.562045Z digest=sha256:3c83445e10ceaa59264d0ea60420184ba6c6200977f1aa163e5bcabfbbb24c37

Observation 1b66d188-aedf-42e4-9b63-f192e703d310 · outbound

This paper cites MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.624099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.624099Z digest=sha256:be02bd56c92397db2e5bc79e816fdd501eac3bd17ca572f427d587b550bb2e6d

Observation ebe99975-196b-4298-881d-08125ddea5b5 · outbound

This paper cites Advances in Neural Information Processing Systems37, 75329–75354 (2024).

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Advances in Neural Information Processing Systems37, 75329–75354 (2024)

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:22.536374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:20.685023Z digest=sha256:a7758fa7282872d9af6ce9a4691869f1e0214378046b7ec64529f7181412abb2

Observation ba7731e0-5a26-4d91-a2bf-54d9b040e728 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.794156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.794156Z digest=sha256:ddcdbf225ddfa182b6cb70fd8303fa9be626df2fc68bee831612f7af01670e50

Observation 8b3c9276-4fe7-490c-8ff4-69da8e1198d5 · outbound

This paper cites Qwen2 Technical Report.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Qwen2 Technical Report

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.859051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.859051Z digest=sha256:f0698c8d502f021031d577c8051729e053df195f80fbec50012b34df84aa12e9

Observation cda4bc13-1d3f-49b5-a92e-af949630442a · outbound

This paper cites Qwen2.5 Technical Report.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Qwen2.5 Technical Report

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:20.936715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:20.936715Z digest=sha256:15018eccac39e049dd90c33e415d4aa587fcf2a225d80c02a786a5aef8074717

Observation 9c9854a8-3074-4a3b-9afa-01ab99df80f4 · outbound

This paper cites HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation HaploVL: A Single-Transformer Baseline for Multi-Modal Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:21.023602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:21.023602Z digest=sha256:c104102cc882a6ddb68ee86c42bbeed2fea3801bbe7496a41a5475e123099c0f

Observation 43715fe0-ce7c-4853-afe7-a83dde80beaf · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:21.110375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:21.110375Z digest=sha256:22afe233293ab5132d8cdfe6ea6f6152468b39524d696f46a86044bfeb37c25a

Observation 65eca0ba-72b5-4989-9cff-db240952f1ab · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:22.348993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:21.194750Z digest=sha256:fcd3cffbbe89e45f67aa5e4873c6d28b8fa536c59ac5dfab3b4814430cbd433c

Observation d2090dbf-d07f-4626-8914-5e8024d52412 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:16:22.168992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-07T11:16:21.267984Z digest=sha256:f468873a16aa81522f864a79010ba729605bb9b28c0c63ab3352e79f4182add1

Observation 379e6b75-2222-4fe3-b221-6e2c9c7a5844 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:21.341077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:21.341077Z digest=sha256:739dbc4c05ea8cb89defebdabc5158025ab9eef5c01ed2e24a67f2d6fe5445bb

Observation 05d6bf8d-edad-4a8c-ae4b-ed5e8185b22f · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:21.439643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:21.439643Z digest=sha256:f6caf2655d7acc6f68d76c10f31ee0664219cac41bfc6ee4346bd7f28f978719

Observation 7fb1e717-4b06-4838-a2d0-2521d0f7bfa3 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:21.519000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:21.519000Z digest=sha256:dc50fef71db542da6a047284a758082ce00a339b8e7bf6a084990218266bbc09

Pith citing papers

Observation 965001ce-327d-438c-9039-a5e6a168f8e6 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation

Reference 124

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.840730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:88097b525488c7d0be9b9df76b3a24bdf64bf7c5682b8b8153665bdca4323616

Observation 457ff10f-9cbf-418a-bfb6-e055bc7015f5 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation

Reference 133

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:48:14.984218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T11:46:52.658984Z digest=sha256:5d7c9ec803193b365eed2254babba30bfaee87de99842ca7adfd43de29a6c052

Observation 0ccd6c77-5403-4c12-897d-a4c838967cf7 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation

Reference 134

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.484062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:0303e268d28464acbeec972b14a818a6f2c8bbd16f8fc999e2da7ee6a298fccc

Observation b7a9c49b-0821-4690-9c9f-c2d70ad77567 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation

Reference 206

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T14:38:28.864384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:d9d077ac745adfdf2013033ed8f4cceef595d352828ea0a5c02f87ba99aa8236

Observation ed9aaa69-983e-4bc4-91c5-f14b6ab72a28 · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-01T10:05:40.316917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:6fbf36af4ad567cbb93ff3f410db0156a32dc053d79056984fa700e6466ccc56