Pith. sign in

Paper Citation Record · LEDGER

UniVideo: Unified Understanding, Generation, and Editing for Videos

As of 21 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 36 inbound Pith citation observations for arXiv:2510.08377.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.08377 v4

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:49:50.259288Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 36 of 36 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:08:56.520608Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:29:44.530433Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 73f4a28d-5da8-46e6-90dc-209baa4763cc · outbound

This paper cites Qwen2.5-VL Technical Report.

UniVideo: Unified Understanding, Generation, and Editing for Videos Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:47.214037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:47.214037Z digest=sha256:bff8e96cb7208d767208bccba7a5c10dcd65efa0439c9a0f77ee2962629d8f47

Observation 291f494f-cd31-4f1c-af3a-e10e1a0ffbf9 · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

UniVideo: Unified Understanding, Generation, and Editing for Videos DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:47.633545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:47.633545Z digest=sha256:05a8dce528f22366221ef2f3ecc9f2162d375ffc268ccc57737695652e6d3d87

Observation f2594ed8-bc9d-4eef-85ea-42ed4a43db50 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

UniVideo: Unified Understanding, Generation, and Editing for Videos SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:47.713548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:47.713548Z digest=sha256:839464849b088ab6ee838ddcfaac22b134cbc2151a7920fe0cb7b1b7e1c8dee2

Observation 8fb516ce-e28f-4d8c-9908-fd9be1e60df7 · outbound

This paper cites MetaMorph: Learning Universal Controllers with Transformers.

UniVideo: Unified Understanding, Generation, and Editing for Videos MetaMorph: Learning Universal Controllers with Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:47.817748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:47.817748Z digest=sha256:c77e24ddedd44d95109771422a0191b6c0b7d06c92c8e049896dc36de931db15

Observation 72a86b75-9ad0-468f-9b00-5473d56c1d02 · outbound

This paper cites ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning.

UniVideo: Unified Understanding, Generation, and Editing for Videos ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:47.891881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:47.891881Z digest=sha256:9a99e1c62406b77030231c6262be8624f98e3faea5fc5983d4a88e2a41f4a1bb

Observation 1bb0791c-c04b-43ee-a89f-c91a7479c3a2 · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

UniVideo: Unified Understanding, Generation, and Editing for Videos VACE: All-in-One Video Creation and Editing

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:47.960559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:47.960559Z digest=sha256:7738d2e19bada81fccae372e6e7685f7350e7878abd6193daf723ac5e7a5aa78

Observation 43531f67-cbdb-47f5-8b93-7649a5243a81 · outbound

This paper cites FullDiT: Multi-Task Video Generative Foundation Model with Full Attention.

UniVideo: Unified Understanding, Generation, and Editing for Videos FullDiT: Multi-Task Video Generative Foundation Model with Full Attention

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.042362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.042362Z digest=sha256:53abbb1f09fb12d402bd5d7653d96cce94f3f1925f4021f80f6cab1b0e96f6ae

Observation cbe1dc13-452b-4167-b290-dd39e60f4179 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

UniVideo: Unified Understanding, Generation, and Editing for Videos HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.115320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.115320Z digest=sha256:c93d09aba02bbc1e926c02362c9d5fee6127986854fa2fd3e3b33e6e94bb6108

Observation f1cd6c3b-2e50-46d2-969e-99bfb4bbf384 · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

UniVideo: Unified Understanding, Generation, and Editing for Videos AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.197996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.197996Z digest=sha256:b3017a3af378da0ecdf0a90dba872871ed6486083484dcadb8ef7128d9c50efb

Observation 5c5c57b3-eb6b-4eea-930f-77781188d7e9 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

UniVideo: Unified Understanding, Generation, and Editing for Videos FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.293926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.293926Z digest=sha256:a032d018c3db009f99bc6bbdcde6771b426021392237769d18d08aa0212346a6

Observation e33603d9-f1d7-4fba-99cd-1d3502c90ba1 · outbound

This paper cites Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation.

UniVideo: Unified Understanding, Generation, and Editing for Videos Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.370382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.370382Z digest=sha256:832da025166da7c205b9e8de8b02d9a2429bea91e6ba59934c8435dd16a2123a

Observation 22f64a6c-48c4-457a-96a0-ed4b53d86ad2 · outbound

This paper cites MagicEdit: High-Fidelity and Temporally Coherent Video Editing.

UniVideo: Unified Understanding, Generation, and Editing for Videos MagicEdit: High-Fidelity and Temporally Coherent Video Editing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.433935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.433935Z digest=sha256:fe906a9aef1c462e9d6ea405f50d67da35eef3c898c8e12c630453d67c318a44

Observation 0361f6c6-13a6-45b9-be31-c9d621a5fcdb · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

UniVideo: Unified Understanding, Generation, and Editing for Videos UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.516988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.516988Z digest=sha256:cfee5b52a46e08745913d92cbd7cc0a74da5a428be51d4b61acb6458f2d9ae19

Observation 91ed0832-9ef6-4f2a-a878-5ce4e3355301 · outbound

This paper cites Phantom: Subject-consistent video generation via cross-modal alignment.

UniVideo: Unified Understanding, Generation, and Editing for Videos Phantom: Subject-consistent video generation via cross-modal alignment

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.589583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.589583Z digest=sha256:91a65f84e77b86e2cebd0556fc89bcddc2fe7b057d552b760e30d1f406bf1e0f

Observation 8a4bf5bb-9ddc-4ecb-9c1e-0b0b518987e1 · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

UniVideo: Unified Understanding, Generation, and Editing for Videos Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.658757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.658757Z digest=sha256:be40bdec5a80046133fe1c917fd103276bad5d6d37c56deb918c07a3f2a7ceba

Observation 8dd57aff-87dc-49d6-8422-088ada56070b · outbound

This paper cites Transfer between Modalities with MetaQueries.

UniVideo: Unified Understanding, Generation, and Editing for Videos Transfer between Modalities with MetaQueries

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.716774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.716774Z digest=sha256:381cd0452bc5616c0e17e1824f38dd4bc14303c4e157dab1324dda8e5cf73969

Observation 224b0d5e-a92d-4b5c-ae89-95ef1590bc8b · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

UniVideo: Unified Understanding, Generation, and Editing for Videos SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.799944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.799944Z digest=sha256:1f2fca7773957db382bf0be72923215e6ea1691e69fd349ff12530adda574f26

Observation 6f57bb97-5184-4f30-b402-53c3a92da22a · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

UniVideo: Unified Understanding, Generation, and Editing for Videos Movie Gen: A Cast of Media Foundation Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.898580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.898580Z digest=sha256:d5bfedc8ba6f6b252ac8185ade1cc11eac8caf6b1af8e9dcde8ec11a7d05cb7b

Observation ef364fbd-f8fd-4b0c-bf37-103f5d74a52b · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

UniVideo: Unified Understanding, Generation, and Editing for Videos LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:49.019944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:49.019944Z digest=sha256:5ded32bd7b946e87aae9e06b1f88c4ca97a941f4b2f3b666be25cdc61abebf70

Observation 87f1b4b7-2b07-45c9-9548-2fccc12e8e81 · outbound

This paper cites OminiControl: Minimal and Universal Control for Diffusion Transformer.

UniVideo: Unified Understanding, Generation, and Editing for Videos OminiControl: Minimal and Universal Control for Diffusion Transformer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:49.107557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:49.107557Z digest=sha256:79c674d2a7fd7bf2c3971bf4b0a7c9c0dfb0c9bfe6ca4f1c2552701dfe5b66f3

Observation c385d80f-a4ea-44aa-a241-00729ebadf3d · outbound

This paper cites Omni-video: Democ- ratizing unified video understanding and generation.arXiv preprint arXiv:2507.06119,.

UniVideo: Unified Understanding, Generation, and Editing for Videos Omni-video: Democ- ratizing unified video understanding and generation.arXiv preprint arXiv:2507.06119,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:49.208736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:49.208736Z digest=sha256:9bc7f1aeda41bb8e16b8dec53714c65ab28c7adacc67e1be7c53d35db4a970be

Observation 68e86199-c3d0-4dec-afdc-12e88f8d9689 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

UniVideo: Unified Understanding, Generation, and Editing for Videos Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:49.292341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:49.292341Z digest=sha256:a4045be7eebae0600e7207ee86d77a89114a395179684e8d9fe856a0fbd619f1

Observation 2b00b1bd-ecf1-4dbd-8086-f5a51e1eb447 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

UniVideo: Unified Understanding, Generation, and Editing for Videos MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:49.376070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:49.376070Z digest=sha256:b8c352fd02503a3c06d2c6ff4f8f1d4236e7e65046da513bb4dcd5db5d1834c9

Observation 8e50bb36-f926-4665-9017-123eaa9d181f · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

UniVideo: Unified Understanding, Generation, and Editing for Videos Wan: Open and Advanced Large-Scale Video Generative Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:49.471148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:49.471148Z digest=sha256:43fb98add9b41ca3205530c7428340d1478a81ae6e2270d94ca2a467a6d3d9c9

Observation e19da07c-ce57-492d-ade3-cdc033495cb4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

UniVideo: Unified Understanding, Generation, and Editing for Videos Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:49.533691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:49.533691Z digest=sha256:92b17c6c092e11963a7e99524e744ccc6fa86b529601d5664a2f02ce9f969e37

Observation f9e00c27-378d-45e7-b361-a3fddda9c664 · outbound

This paper cites Qwen-Image Technical Report.

UniVideo: Unified Understanding, Generation, and Editing for Videos Qwen-Image Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:49.621745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:49.621745Z digest=sha256:93aaebe08d79b804aace1c315886c86e970a3645984a2e6bc2e8b6a8449b12c6

Observation d5e5cbc3-690f-4ec3-a6df-8f03e30e3d25 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

UniVideo: Unified Understanding, Generation, and Editing for Videos Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:49.687091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:49.687091Z digest=sha256:f0e08b17fa194d89359f04f415b41d96bbb542e26d28935b9e81a2cf96fb430e

Observation 3636b190-470a-4ac6-a97b-fc1713b41f8d · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

UniVideo: Unified Understanding, Generation, and Editing for Videos Show-o2: Improved Native Unified Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:49.772549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:49.772549Z digest=sha256:e7b35126798dc271410f1ef325561577210b8afb13a93ec8272826918e35a6e8

Observation 04ae6fef-ce96-4ed5-af21-10ba586ce2df · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

UniVideo: Unified Understanding, Generation, and Editing for Videos CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:49.864956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:49.864956Z digest=sha256:5cfa7483f87cf793458ae181f463e1fdb4e791b89a7df5a8068b6bca165e475c

Observation e81c6731-857f-4f5e-a24b-85227af32ce8 · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

UniVideo: Unified Understanding, Generation, and Editing for Videos ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:49.960605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:49.960605Z digest=sha256:1cf04668dc10f604201d0a5931d4235f5ddb5aacdb772e18cd94b97d50e1885c

Observation 8b1fcd1d-cedf-4410-a845-cbedbab1ccf7 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

UniVideo: Unified Understanding, Generation, and Editing for Videos Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:50.038738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:50.038738Z digest=sha256:9ef3069a8e50784e634a6202f072dc5e735f60da3b70797e004c9406ba5dd8ca

Observation 7da4950e-ded6-4e99-96ca-c6a2a9f1d7f7 · outbound

This paper cites While we do not observe task confusion, it sometimes fails to strictly follow editing instructions, occasionally over-editing unre- lated regions.

UniVideo: Unified Understanding, Generation, and Editing for Videos While we do not observe task confusion, it sometimes fails to strictly follow editing instructions, occasionally over-editing unre- lated regions

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:50.096466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:50.096466Z digest=sha256:49e912889493c42d244227758d8ca95974d87a5af135697c119cdc98047c26bc

Observation 27333d34-02c2-48aa-bf54-a360629ce48a · outbound

This paper cites We also source open source data such as OmniEdit(Wei et al., 18 Table 8: Model capabilities across understanding, generation, editing, and in-context generation.

UniVideo: Unified Understanding, Generation, and Editing for Videos We also source open source data such as OmniEdit(Wei et al., 18 Table 8: Model capabilities across understanding, generation, editing, and in-context generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:50.166756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:50.166756Z digest=sha256:0cf876363fbfe849dafc5835803620ed523a178ec48d7b4b1d0fb95c3ad33809

Observation bef9cec9-32b9-4f93-ae64-ad6939331c77 · outbound

This paper cites an unresolved cited work.

UniVideo: Unified Understanding, Generation, and Editing for Videos Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:50.259288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:50.259288Z digest=sha256:6d7513c1afb3fde89e9143232a83b9a2e1a73c21f00830c720b37490c8da5b3c

Observation cd706d27-0f10-46e7-a2ff-ee8f5b103b7a · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

UniVideo: Unified Understanding, Generation, and Editing for Videos SAM 2: Segment Anything in Images and Videos

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:48.965766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:48.965766Z digest=sha256:c5bafc41fb9592ddfe1d948da48eec98bab9ca0f7294f84dcdddf540d8994bf7

Observation 7622feaa-cad1-4b97-ad46-ebacadcfb9db · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

UniVideo: Unified Understanding, Generation, and Editing for Videos BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:47.484727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:47.484727Z digest=sha256:e2aa3052bd719fdd4f728ce38b0df553b1f0f3adf8111306f89a158a929fce4e

Observation 602babb6-035c-45da-8804-32abc1170651 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

UniVideo: Unified Understanding, Generation, and Editing for Videos VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:47.409544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:47.409544Z digest=sha256:2948a3566ca31aa902854f9dfb5c28336924a58496edb747523c1567a80fcebf

Observation ad60875f-2bd0-459c-bbc6-78c5792f2242 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

UniVideo: Unified Understanding, Generation, and Editing for Videos Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T10:49:47.297010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:49:47.297010Z digest=sha256:2cc0c7d230fd78e0eea872dd003399d30abf7885d2f67a4fc9ff989a9c657ab1

Pith citing papers

Observation 5b95dbdd-b7d7-40cd-8b07-15eb457e9171 · inbound

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation cites this paper.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:09.213780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:09.213780Z digest=sha256:fdcc4e8c6d6df70ae2f0cae94b04ec25313a02cd33ef80526436a968a80d4c0d

Observation aa5d514e-1975-4f61-8e4c-313360402fc6 · inbound

VideoCoF: Unified Video Editing with Temporal Reasoner cites this paper.

VideoCoF: Unified Video Editing with Temporal Reasoner UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T00:08:09.479706Z digest=sha256:4fd1f7730be678c2a1ccc0cf501c020de28104230de6cc623f9a0f4f8685358c

Observation 5b021933-8d1e-435d-9505-8fe570e0155a · inbound

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens cites this paper.

LLaMo: Scaling Pretrained Language Models for Unified Motion Understanding and Generation with Continuous Autoregressive Tokens UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T05:01:11.880003Z digest=sha256:21ee0e12465901a48339ff0a0e4b0b261902fda147040061d15e1c5ed0c38110

Observation 6f86792b-2400-422e-adda-9970a75c8ed7 · inbound

Under One Sun: Multi-Object Generative Perception of Materials and Illumination cites this paper.

Under One Sun: Multi-Object Generative Perception of Materials and Illumination UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-13T22:08:47.493022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:08:47.493022Z digest=sha256:3c21b4a7135787b8ba3e209f48f171208867e4304ca5d744ee67f9c5f005c199

Observation 3efc82ee-462b-4a27-9657-cc0fc8864b4a · inbound

Physics-Aware Video Instance Removal Benchmark cites this paper.

Physics-Aware Video Instance Removal Benchmark UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T19:27:33.716719Z digest=sha256:5e376287d4246f32888f06a4728679cd9a494f4fde71b664cef34cb6656e67b8

Observation 4bc9f2fb-742e-4e37-b893-0f5b9d27a9de · inbound

ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks cites this paper.

ImVideoEdit: Image-learning Video Editing via 2D Spatial Difference Attention Blocks UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:49:33.822264Z digest=sha256:259b3dde6e97fb314f5ea5b52785981e7c24bd2e3f760c0415fc2974d4df6ddb

Observation e9852fea-77fb-450e-b677-8fe93699ea63 · inbound

InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation cites this paper.

InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T17:39:59.791758Z digest=sha256:f513173b12e78a942313efc58e0403ed5d0e951dbbf771bcca0ae62045df9eff

Observation bb785f1a-bfac-4540-9be7-897b51354570 · inbound

Controllable Video Object Insertion via Multi-View Priors cites this paper.

Controllable Video Object Insertion via Multi-View Priors UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T11:44:17.033051Z digest=sha256:72aa9660a8143f5f640bbea2f03dbc56948abb2b9580f501a9bc52dade294dc1

Observation e2699323-abc0-4aa2-9de6-774b28e2d6e7 · inbound

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects cites this paper.

VEFX-Bench: A Holistic Benchmark for Generic Video Editing and Visual Effects UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T08:24:41.925800Z digest=sha256:a2d41aa65b47e5d7a1c46a6bfabec2030488f51cb139fe46078270b99972ba90

Observation f9122528-1e0e-48b7-b0d4-4eca101c9181 · inbound

How Far Are Video Models from True Multimodal Reasoning? cites this paper.

How Far Are Video Models from True Multimodal Reasoning? UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T02:44:52.920816Z digest=sha256:0a1bebc75397ca17d7d8fcbcbc5c349f617e83c45870f9b2a73b0bef204ed1dd

Observation 89729954-58f5-41a2-b937-434747ac3461 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T04:31:26.325118Z digest=sha256:d0269323cf35bc444670f565a1d553eeec083165f5f4eb9e6899c74afd8f02f0

Observation 334ecb48-2067-4346-91be-f1a6d78d489f · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:7f7849b583e6b7305297055f4aa06553c24aac9567f2381cc4357f84d89b3ffb

Observation 384b0a38-4bef-4938-b650-a1f5be80ec0c · inbound

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE cites this paper.

Mamoda2.5: Enhancing Unified Multimodal Model with DiT-MoE UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T18:26:58.696936Z digest=sha256:8910d24f421439a8c3373bbb1d6bfd625a2ee7dfbb69790faa2c0966e65c1a4b

Observation 8b34f157-24f5-4478-a288-0f55552c113f · inbound

LIVEditor-14B: Lightning Unified Video Editing via In-Context Sparse Attention cites this paper.

LIVEditor-14B: Lightning Unified Video Editing via In-Context Sparse Attention UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-07-01T00:24:30.573122Z digest=sha256:be668ae8f8a7baaae6b79006b67089aedd9101b609740bd1efdb9cc5281b36d5

Observation 71384b89-cd41-4808-8cb9-2dbaa7f9deb8 · inbound

Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance cites this paper.

Sparkle: Realizing Lively Instruction-Guided Video Background Replacement via Decoupled Guidance UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T12:34:34.315135Z digest=sha256:edd84f556f215316810214986950b0b8ac7f391787aa22a9aa51fc71598844da

Observation af13ddc4-6e8b-4a57-8904-a94ce9bbae4b · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T11:46:52.658984Z digest=sha256:a0f7b119f1f3e4c714aff27c030131fea3f76acc2571b6f6a833a4e8ec2c7b54

Observation f0a9f1bd-47a0-4652-9d7d-90ddbe70b3be · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 121

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:0a2af6e622d22cc6de6a56293404b76543fb65b193e38234f0dae156422a5bfa

Observation df27bdd1-8ddb-4d15-951a-846dd2d2b781 · inbound

Aurora: Unified Video Editing with a Tool-Using Agent cites this paper.

Aurora: Unified Video Editing with a Tool-Using Agent UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T10:47:05.038308Z digest=sha256:1a5bce452cd7445e28f85461f420520a9c681a8195126a102c23eb4df8679322

Observation 24ee577e-95a1-4ce1-b1b3-10ef721208b6 · inbound

What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing cites this paper.

What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T05:38:29.361039Z digest=sha256:19cd23f4d2a5cd2a2cd4ca699ec82eb86b555deb33e3958eeabeb56986bd2a31

Observation 8220ccec-a832-42d0-b2ed-492ddddc7d86 · inbound

What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing cites this paper.

What Semantics Survive the Connector? Diagnosing VLM-to-DiT Alignment in Video Editing UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T17:25:59.714252Z digest=sha256:9a1351f845952a41d507ff7120b0ef29234e9acea48ce794ac8f6fb4f72bcdb9

Observation 31d6847a-ac4d-4a11-8254-548505d81543 · inbound

Bernini: Latent Semantic Planning for Video Diffusion cites this paper.

Bernini: Latent Semantic Planning for Video Diffusion UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-22T06:39:47.124605Z digest=sha256:be7a5fa3dffdb6191f5d4f706d1746c759e0eec8645d054bfd5c91bba6c6a4e1

Observation 63d551c3-d4ae-49b3-a8ea-6d4d86cd5da0 · inbound

MotiMotion: Motion-Controlled Video Generation with Visual Reasoning cites this paper.

MotiMotion: Motion-Controlled Video Generation with Visual Reasoning UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-22T05:51:17.890788Z digest=sha256:065a0838fc04ea4c80aaf06574cdf13d51fab5a212fef8a0dd1e71ab8e978c1a

Observation 66a067ea-f1a5-4b27-9f52-812908c993ee · inbound

Smart-Insertion-V: Photorealistic Video Insertion via a Closed-Loop Feedback Dual-Stream Framework cites this paper.

Smart-Insertion-V: Photorealistic Video Insertion via a Closed-Loop Feedback Dual-Stream Framework UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:30:20.593882Z digest=sha256:2dd784aa9bdf998a2b4260c5f5fa9bcf6bad9301eb480cdd47211a9ccd9da5d1

Observation c19fe86d-b4ad-4bc7-97d4-a6e86dc794dd · inbound

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models cites this paper.

Lumos-Nexus: Efficient Frequency Bridging with Homogeneous Latent Space for Video Unified Models UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T22:56:21.783415Z digest=sha256:90e74d0ac5b4f38c5da63a1ada0faad61c9cd95d03fa1bb05ffa577e76ad6d7e

Observation 7645f489-9dd0-4229-9cf7-9e4552c6986b · inbound

AlbedoEdit: Unified Instance-Level Video Editing with Albedo Guidance cites this paper.

AlbedoEdit: Unified Instance-Level Video Editing with Albedo Guidance UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T15:55:23.583463Z digest=sha256:32807946c30b64e2eb4f131dcc6e6dac019786e2a65091b4ef5ffaf7d7b9ccde

Observation 6a1d2a9e-063f-4382-b861-e68e9b452ec6 · inbound

SteerVTE: Seamless Video Text Editing with Style and Glyph Control cites this paper.

SteerVTE: Seamless Video Text Editing with Style and Glyph Control UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T08:52:23.330736Z digest=sha256:094343a347e26a45cc245e13882825572cf00ddda280cab719100b6dfca7befc

Observation 30e503cf-3bfd-4b26-b6bb-7617b919afa7 · inbound

Bridging Video Understanding and Generation in a Unified Framework cites this paper.

Bridging Video Understanding and Generation in a Unified Framework UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-07T03:17:13.018030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-01T05:57:54.653504Z digest=sha256:049a8db1c6df649744c55a11824163caaf2fd69eb67e54e8c1e78137d3025e90

Observation d34f55a7-8b24-4822-8baf-0564b94102cd · inbound

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships cites this paper.

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:23.605545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:23.605545Z digest=sha256:5ce8dd441a114cc38f9949ed5bf8da57970e18efe5ca77d81f532974043654a1

Observation 25bf3166-fb60-4963-99e1-95f6468e73dd · inbound

Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence cites this paper.

Apple-$\pi$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T21:07:27.750672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:07:27.750672Z digest=sha256:fd74f6ecd4f351390e847b6481eb336e40a037efcce1d570a8faf2d793642a1c

Observation 0efbaabf-c1b7-4f11-ac4b-db74755c689b · inbound

VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation cites this paper.

VIPER: Visual In-Context Physics Reasoning for Physically Plausible Video Generation UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-30T21:29:22.037630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T21:29:22.037630Z digest=sha256:3412851696c367e8241d92cbe3c9dea8114740add07bc6b8d0cd55c9a3708196

Observation 2051de52-566c-451f-a9b9-cd78f1a01fc2 · inbound

FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control cites this paper.

FlexComposer: Unified Video Compositing from Images to Dynamic Footage with Flexible Trajectory Control UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T03:17:58.854818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:17:58.854818Z digest=sha256:f4c957e571926f23358e356cf6cf5cfab17b152abd6b775e83984ae9f34f55c6

Observation 712e81ca-a89f-4179-ab3c-039df3dce6cf · inbound

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion cites this paper.

JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T04:52:48.552579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:52:48.552579Z digest=sha256:0a50034c0b1d7f5b92abbe902d715ee5391932851eeddd7e2a034e523a0f84b4

Observation 21b72b64-7767-4efb-8901-05e75f21e5a1 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:28.155399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:28.155399Z digest=sha256:fcb28c1679db919ad0666d5d315d8411333e30364355865080a34e22d5c6b18a

Observation 8113f314-ad63-4368-9e81-7c10a5e28a52 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 133

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:56.520608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:56.520608Z digest=sha256:15f51668c37ea4fa0775004809a688709224beee2469a984b3287586de4e8d07

Observation 10f85d00-c887-4d7d-b835-3a774a43a6de · inbound

OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing cites this paper.

OmniEdit-Bench: A Comprehensive Benchmark for Instruction-based Video Editing UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T10:34:42.397646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:34:42.397646Z digest=sha256:c19f39caae5b26fdf01708843429efa718a2c10f6e630b662fa6485aa975aaa4

Observation 792a3c7b-e677-477f-9104-59b7fa1676d2 · inbound

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing cites this paper.

VideoArgus: Agentic Rubric-Grounded Unified Evaluation for Video Generation and Editing UniVideo: Unified Understanding, Generation, and Editing for Videos

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T12:16:44.816528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:16:44.816528Z digest=sha256:60899ff9b4a9756b301cdfeb1630613d4d25421b9022cb3380966b167c3e992c