Pith. sign in

Paper Citation Record · LEDGER

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation

As of 6 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 3 inbound Pith citation observations for arXiv:2604.11804.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.11804 v2

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:09:02.727887Z

measured 82 of 82 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:58:52.344771Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-11T12:41:04.415342Z

Reference resolution

79 of 79 outbound references displayed

  • verified exact43
  • verified fuzzy33
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7c0e0ab3-8b5c-46a1-9ab6-317818d6db75 · outbound

This paper cites wav2vec 2.0: A framework for self- supervised learning of speech representations.Advances in neural information processing systems, 33:12449–12460.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation wav2vec 2.0: A framework for self- supervised learning of speech representations.Advances in neural information processing systems, 33:12449–12460

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.119097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:90b7f79d73cebc74672d7c48fb56c1bb7bd6bce3bba1d9e0ef65a84536975b4d

Observation 0a11ef38-8f34-497f-bbe9-3778fef293b1 · outbound

This paper cites Pyscenedetect: Python and opencv-based scene cut/transition detection program & library.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Pyscenedetect: Python and opencv-based scene cut/transition detection program & library

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.122167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:2a88c00f375be568b46050885115669e558c7bf0d0de9f94207041ae59ff6f1e

Observation d78f62b0-a7ab-4b1e-a4f5-6106ce7fac0a · outbound

This paper cites VirtualModel: Generating Object-ID-retentive Human-object Interaction Image by Diffusion Model for E-commerce Marketing.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation VirtualModel: Generating Object-ID-retentive Human-object Interaction Image by Diffusion Model for E-commerce Marketing

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.859329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:f04618cd34843ab00b78aa790e1668de8360cf029986a2dc6ca37118a5199cd8

Observation c0aa2cf3-d9d9-4912-b478-6bb81d717bd9 · outbound

This paper cites HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HuMo: Human-Centric Video Generation via Collaborative Multi-Modal Conditioning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.846071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:8c6e3f52f18f1f7c855ccf8bc4adf9984792ab8d118cb923ec32822a4f1e8f4b

Observation 69fa8af0-a40b-4691-8988-975eaecdbabb · outbound

This paper cites Posteromni: Generalized artistic poster creation via task distillation and unified reward feedback.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Posteromni: Generalized artistic poster creation via task distillation and unified reward feedback

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.676680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:fa17ea6074ca4a8891b6c24b8811dfff74be2f20d76814e03a03c77cb290e374

Observation d234a1bb-a95d-4a21-870b-5da508d07eaa · outbound

This paper cites Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Seedance 1.5 pro: A Native Audio-Visual Joint Generation Foundation Model

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-16T01:35:37.953991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:d178bcf31e50b623fb422386ee1560dc4679715f043f8ced4dadf7b25d17be8a

Observation a70e2c08-2720-43c8-ab53-4ee090b879ef · outbound

This paper cites HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.831890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:8e5eac1f96eef2a6264954ff7029fe038cfc44ce0157515b87d766e723ac0ea2

Observation 1e0deade-a0fc-441c-a69f-1d71fc7c4483 · outbound

This paper cites Out of time: automated lip sync in the wild.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Out of time: automated lip sync in the wild

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.113205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:f32b16e83c0a8640cc07d73fcebc5ca432a51d5991035f5f41e6c867ab6710b1

Observation 5dda2b6c-5a65-4c0b-b65d-ebe7ba1e6547 · outbound

This paper cites Hallo4: High-fidelity dynamic portrait animation via direct preference optimization and temporal motion modulation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Hallo4: High-fidelity dynamic portrait animation via direct preference optimization and temporal motion modulation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.774839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:7082137a6ae2be1720a08843167c9344e36cb4c34b477d36fb9c3d627597f508

Observation b992870a-45a2-4560-9a3b-bdb6a46030d4 · outbound

This paper cites Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.053218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:bde65a025b7bd33f437eccc21409de2db6648ba03a018f48144f9b50f7d7c7e6

Observation 2474ed55-67f8-4df7-83a7-40a46e12e8f2 · outbound

This paper cites Cg-hoi: Contact-guided 3d human-object interaction generation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Cg-hoi: Contact-guided 3d human-object interaction generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.032473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:94945b537a9c354eceed249875769ded52e9c34283063e8c63cbb096f551cfaf

Observation 2f250b84-4147-482e-b721-6df31f58e8e8 · outbound

This paper cites Elevenlabs: The most realistic voice ai platform.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Elevenlabs: The most realistic voice ai platform

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.055705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:6ace3480fe4146a4db8501f2b4e4db24f4a679f72764c19c96f774129e1216fe

Observation 43e4b1f9-9535-4e37-900f-e14a34faf8b9 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Scaling rectified flow transformers for high-resolution image synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.061791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:2de3453ca7609ae68cbbd999409265e5656793874d63336001dd9e5dd075a044

Observation 8038e625-0a36-4e33-a68f-f695e43e8a77 · outbound

This paper cites Re-hold: Video hand object interaction reenactment via adaptive layout-instructed diffusion model.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Re-hold: Video hand object interaction reenactment via adaptive layout-instructed diffusion model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.053914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:bc7be6ceb431765c3fbb7dc428555bb315b84748479f87d3abf81e1a9e3d3c93

Observation 7be73ee3-1d8f-4aa0-a671-589d0b55a3ea · outbound

This paper cites Two-frame motion estimation based on polynomial expansion.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Two-frame motion estimation based on polynomial expansion

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.044227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:633947344b661aee8496fb2e91325fec1e39b71da425da79d74ab1a8ad8f81b5

Observation 6d430c7f-fd44-475e-80bd-e20e4a92d472 · outbound

This paper cites SkyReels-A2: Compose Anything in Video Diffusion Transformers.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation SkyReels-A2: Compose Anything in Video Diffusion Transformers

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.589537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:6b5a95b68e81b7b6d68b4c6aedc785c5948de68c236d425ebb4e014b3c967cb5

Observation d78da58c-5a2f-425c-a00f-78ab574cdf1f · outbound

This paper cites HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.606070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:a1db6eb42f25b8d79f189b522ad78a2201c7decfe4189b5ebad1f72bb70755b6

Observation 12c4009b-f34c-4184-9261-dbdc18e49691 · outbound

This paper cites OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:11:00.788856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:67dae847e3231174cae00359dbd9aa5638d5892caaed2a8418dc8684ebcd9c9e

Observation 5dde06dd-01e1-4411-83d8-9171573b25ad · outbound

This paper cites Nano banana.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Nano banana

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.116094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:ae6eb92469251f50ba6c1d1af9af528d562db4ca5160274c2ba09c36b21d4b54

Observation 16312be4-5a15-43be-8551-a33f0c4f803c · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.914673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:b48c8a25d548845ecc26c6541fdd078fff84e7371253014557b403eb562f9e47

Observation c1d8d2a0-a7d1-48f9-9bb7-bff428615352 · outbound

This paper cites HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HunyuanCustom: A Multimodal-Driven Architecture for Customized Video Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.806138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:324f63b98f8cd8ea7b4b0966215726da44a827dbd8ce2cd5d7df09e0c0babcc3

Observation 6863fa26-4df1-46ec-b2ed-2a160aa76563 · outbound

This paper cites Magicfight: Personalized martial arts combat video generation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Magicfight: Personalized martial arts combat video generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.061022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:d675c4644cb5f574cc7fdef7fbb18cd343a020c645594c0dba6a7b18407671c0

Observation db41c143-0e15-4ef0-96ed-cddfe57c7860 · outbound

This paper cites Dual-schedule inversion: Training-and tuning-free inversion for real image editing.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Dual-schedule inversion: Training-and tuning-free inversion for real image editing

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.070649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:28e48ff5324b079a73d8831476ef4db3ef7795128c1ef3d0d1fe88baf143b500

Observation c8844a78-64e8-4506-aee1-dfe9f0ec77da · outbound

This paper cites M4V: Multimodal Mamba for Efficient Text-to-Video Generation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation M4V: Multimodal Mamba for Efficient Text-to-Video Generation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.767995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:398dbd3565114a02c47c98069ad6ef4bf8c3d2554534c2dc11676aef78b1962a

Observation 106168a0-6b39-4c67-a652-01b39ba9afa4 · outbound

This paper cites JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-08-03T03:15:29.416338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:dbf4ee0266d026c49963301fa1d84e320f472a5834d90ce8062d4bc55b841377

Observation 5af15bff-4151-4342-8651-a88c127594c5 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Vbench: Comprehensive benchmark suite for video generative models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.041097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:85b098d774694578c0ed662ecfe690b21a832ac3dce4b9243ed18a1c3d311595

Observation 85627e22-73e7-4684-b5a1-bea08e89bdd5 · outbound

This paper cites HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HunyuanVideo-HOMA: Generic Human-Object Interaction in Multimodal Driven Human Animation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-09T01:19:36.380554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:639d603a154b605a4c3c232d9673e285e3848885ecc33f8d81abf36d54d10cf4

Observation b0e241ac-e508-4514-8400-ee23df8a4e38 · outbound

This paper cites GPT-4o System Card.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation GPT-4o System Card

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:06:05.898145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:fa79152e35401c2370ab42273461e535ac602dbd51ea7c141e0112a6d3fe664e

Observation 7646f2dc-3ccc-4346-b9bf-599f0d9694a9 · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:22.650116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:3fba991b0da6e6cf90ff8e3c91419b21d2f36008912ce6ec1f37ddeb055c9374

Observation bd328091-e011-475a-81cf-d0acb6ef7d7f · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.798707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:70f9289088124d546cb19e781c0897874c0fc70a075408c63488d1b1ebf031cd

Observation f3b4905b-9c3c-4370-862c-3e9719185e14 · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation VACE: All-in-One Video Creation and Editing

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:53:54.408026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:ac605bf1860f23cf51bcce9fb5375725130b6cab3324806c3d12188af36c4bba

Observation d7b726b8-ea76-40d6-baae-3f7019829382 · outbound

This paper cites Fulldit: Video generative foundation models with multimodal control via full attention.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Fulldit: Video generative foundation models with multimodal control via full attention

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.057330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:5300fa76ca74b3c52948cd3239fd993485f5a2d07b01bca4333f0d814689cbbc

Observation 789d9386-e5fa-4bdc-9bb3-9bbcd33da32b · outbound

This paper cites Kling-Omni Technical Report.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Kling-Omni Technical Report

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:00:58.771713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:dfa483995c6dcc9a512dc2841216c0d9dcc69e364d1a02c1d1e6ac36b23e66d4

Observation 8a522e79-6f7f-40c3-9027-d8ac000648e5 · outbound

This paper cites Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.818834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:730ac0289cb06f502b80735519d9d3b3b5a9f6885788c5693c0119e4aecc789e

Observation aedb031d-19e2-402c-90cd-b707372eda83 · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.648558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:e35d00e349c11e3f830129b980e8eb2cd3f13560ab0c6f822849dae9dbc34ad6

Observation b5d5222c-58c6-4e63-a675-b14c2f4d778a · outbound

This paper cites Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.099469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:1b36321a687bbba07632a494cd5a3eca473a987433e1f44e5ce661482085ee2f

Observation 224727c4-3375-483c-89c8-3fe684d959c7 · outbound

This paper cites Apoavatar: Expressive audio-driven avatar generation via refocused audio-pose priors.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Apoavatar: Expressive audio-driven avatar generation via refocused audio-pose priors

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.103213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:c780c925588d881dc5d6fdf06e83b2d443f45f38bea0905c96decdbab548f09b

Observation 7d92686b-176f-4cef-9db2-f4464c4cb97b · outbound

This paper cites JarvisEvo: Towards a self-evolving photo editing agent with synergistic editor-evaluator optimization.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation JarvisEvo: Towards a self-evolving photo editing agent with synergistic editor-evaluator optimization

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.699778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:1782e24157658b8ebb767caa652934d4e659021c3e632194297a296074183fa3

Observation 7db8cd8d-d960-41c3-a896-f37002162a86 · outbound

This paper cites Mofu: Scale-aware modulation and fourier fusion for multi-subject video generation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Mofu: Scale-aware modulation and fourier fusion for multi-subject video generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.110199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:babf4d0ce1b881a822a762ccab921d32f58bea1c000425188e0b7f39e4eed4ba

Observation 6fca9cd2-8638-4bca-ac43-ed6b6e80d1c0 · outbound

This paper cites Flow Matching for Generative Modeling.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Flow Matching for Generative Modeling

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:06:05.833272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:dccac88756d116fe8e45923b95c1bd92f3f451872bc2c9b8b26a11bf30fb9bb2

Observation d395414f-ee34-4f9b-b9c6-279fdbed163f · outbound

This paper cites Improving Video Generation with Human Feedback.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Improving Video Generation with Human Feedback

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T15:30:03.116566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:c58c7a56777fbbb304c3a441d510e4fd4d00759d06183319ef21d1f55d1072ef

Observation 78cd8348-c81b-4c6c-8519-b0c3160d0004 · outbound

This paper cites Phantom: Subject-consistent video generation via cross-modal alignment.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Phantom: Subject-consistent video generation via cross-modal alignment

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:06:05.893388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:643e61bec3aa3403f22cccd6f85fe1b1d22ee9dad5edf4204200fd516d79b975

Observation 40f71188-02fd-412c-b9af-d41ddf073a70 · outbound

This paper cites HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:06:05.849493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:843c347d1c0151f7d025d53537989829d880b45ea1a2e5d9afb1d53bbba8f2bf

Observation 780d9e9b-b15e-475f-885f-ff553ae129c0 · outbound

This paper cites Decoupled Weight Decay Regularization.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Decoupled Weight Decay Regularization

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:11:00.639353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:0deadd282fbf5d950029c7d9f8105a59248e3e71df48e4c076c8adf2a4c06143

Observation bd144a32-dee2-4f37-ac36-96724636dad6 · outbound

This paper cites Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:03:15.294735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:9b521baa3a737630bc5313937e4d694f00d5c023eed55bd2ea8083237e1c7c44

Observation 2b4c8b52-aec9-4833-8315-9f910e82535e · outbound

This paper cites Echomimicv2: Towards striking, simplified, and semi-body human animation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Echomimicv2: Towards striking, simplified, and semi-body human animation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.082206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:0d9f11f472a1619e97220df73c8ce60f9a99642f64c1c5cf2c057b402a1f57a2

Observation 81a642b5-b0b2-4134-a1c6-bdcf4de6de8c · outbound

This paper cites Hoi-diff: Text-driven synthesis of 3d human-object interactions using diffusion models.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Hoi-diff: Text-driven synthesis of 3d human-object interactions using diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.078501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:08ece8b4d5b9a74f986042bfdcfa29e5ff92f1016af9e1cee17e17f27a846b4b

Observation 57c6b94f-3749-46eb-a887-b090a8b5d12f · outbound

This paper cites Innoads-composer: Efficient condition composition for e-commerce poster generation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Innoads-composer: Efficient condition composition for e-commerce poster generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.816000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:95e799a308c11dac010c2f30549bf0436ca988b3508ec0e98facdd46d71903bb

Observation ed4aee16-b160-4111-aed8-b9ac8270b8e5 · outbound

This paper cites MagicDistillation: Weak-to-Strong Video Distillation for Large-Scale Few-Step Synthesis.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation MagicDistillation: Weak-to-Strong Video Distillation for Large-Scale Few-Step Synthesis

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.806264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:9bbe36c17ad4276522d9dee5e61e3d6cb1b6c169cf5664ff6847cb4e1fbac9e4

Observation 89f7c18e-ebbf-42a1-9d1a-8ae16292ae35 · outbound

This paper cites HERO: Hierarchical Extrapolation and Refresh for Efficient World Models.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HERO: Hierarchical Extrapolation and Refresh for Efficient World Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:43:07.893654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:ef7b2d6add7e5946d905c700945b6976bf907bc48a64f08e0fc4646cd0dcca21

Observation fbec8712-353d-4f58-bcd6-b1bc2dbf7421 · outbound

This paper cites Scenedecorator: Towards scene-oriented story generation with scene planning and scene consistency.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Scenedecorator: Towards scene-oriented story generation with scene planning and scene consistency

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.887041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:fc978c24f674a453577b8d8d19b5cb1aae1f452d13fbd5da2ac6237acc95cd9f

Observation 5fe34f3c-5a62-48a1-8191-9cf55f247d50 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.095811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:9cf3e9e06805fad483c2b8aa1ad6cc67b8c1480f84d502f1459f512440482954

Observation 5396e65f-354a-42dd-89de-e8a75dad615d · outbound

This paper cites Ominicontrol: Minimal and universal control for diffusion transformer.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Ominicontrol: Minimal and universal control for diffusion transformer

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.012286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:6a35a879729d4ca9b33c9275ad097f1b180e07dbe32b1ae31d2123bc6776ce73

Observation 79a901a0-f9ae-4edf-836e-c405c8ea8cc1 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-11T11:06:05.856685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:f39ad5aa1444c5998fcd0098eb21d0160cc9fa29bfe0026a5395b1a859ce0c2e

Observation eb1f977d-5453-4bb6-a48c-b10c9c130485 · outbound

This paper cites WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.928659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:7773deefd77820fd1ea39edff836e3e5f9c85ab1fe5b8529aca665a49a608b66

Observation 6d2e36d4-16ea-45da-b32b-a61bb3019548 · outbound

This paper cites Language model based text-to-audio generation: Anti-causally aligned collaborative residual transformers.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Language model based text-to-audio generation: Anti-causally aligned collaborative residual transformers

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.071692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:9cccc494c0c94eb66b8ce098ae83b961d93280b8081485580b9e409e2c1c20d6

Observation f9652e81-3c17-4d91-bb19-dae390acd5be · outbound

This paper cites DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation DreamActor-H1: High-Fidelity Human-Product Demonstration Video Generation via Motion-designed Diffusion Transformers

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.709409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:85a23103e7e29bcd6faefb0951b2c78b92f539389f08f0fed5d5c2924a749a10

Observation 0ea98015-ddf6-441c-bedd-44687b90c40e · outbound

This paper cites Fantasytalking: Realistic talking portrait generation via coherent motion synthesis.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Fantasytalking: Realistic talking portrait generation via coherent motion synthesis

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.024248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:38c32d3fe26bde49b1d6610c6366a1025c7ffd0d57c5e26c0501b4a15ce8872c

Observation 2becd6ce-b95d-4dc3-b932-8384b7748564 · outbound

This paper cites In- terActHuman: Multi-concept human animation with layout- aligned audio conditions.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation In- terActHuman: Multi-concept human animation with layout- aligned audio conditions

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.874895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:6fd5edd326e6667ef1841e610a4adb38ecd4861f6d803a340b9cd5be043aa531

Observation d741a0f5-5443-4135-8300-49bb3c0d2ecb · outbound

This paper cites Mocha: Towards movie-grade talking character generation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Mocha: Towards movie-grade talking character generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.086563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:e4d56120f6cf4c8c824869db7fa0943661cecae90bd73b03a9af2f7222fd3f91

Observation b9539934-7b00-4ba7-b8a8-243a28e219ed · outbound

This paper cites HunyuanVideo 1.5 Technical Report.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation HunyuanVideo 1.5 Technical Report

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:31:35.128127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:3b370b173f1bb14c9ffac9cc4ed70bb00ec426c3a7767d4763f47972aa35ed84

Observation 1e47092f-291f-4559-800e-6c313b8ad0c4 · outbound

This paper cites D3D-HOI: Dynamic 3D Human-Object Interactions from Videos.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation D3D-HOI: Dynamic 3D Human-Object Interactions from Videos

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.722409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:26f3dfcf6a7a0910f6b77ef6f3204b7beb6e5a26ba1cd835a025a5c591225c95

Observation bdb5a142-c123-418c-ad72-ad2b9f362a96 · outbound

This paper cites Magicanimate: Temporally consistent human image animation using diffusion model.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Magicanimate: Temporally consistent human image animation using diffusion model

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.106741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:3e9c3ec161e818d915cbd40d44fdd11d61858f10534661615efa3a331c280793

Observation e8344706-456b-48f0-bf70-7eb1f68e219e · outbound

This paper cites AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.565785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:1b1ce5ea22963b6def8812067aad82731558dbb6d0375ef6ed775c34d322351a

Observation 1de117d3-2386-4086-8726-75efef6dd041 · outbound

This paper cites Follow-your-pose v2: Multiple-condition guided character image animation for stable pose control.arXiv e-prints, pages arXiv–2406.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Follow-your-pose v2: Multiple-condition guided character image animation for stable pose control.arXiv e-prints, pages arXiv–2406

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.092349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:7663dd8e8bbe171bde68c4f80dccd4fc3cdfde65a7def7d4a167e07121872bd6

Observation ed8bc4fc-3e2a-4194-b89f-6285583a5ccd · outbound

This paper cites Hoi-swap: Swapping objects in videos with hand-object interaction awareness.Advances in Neural Information Processing Systems, 37:77132–77164.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Hoi-swap: Swapping objects in videos with hand-object interaction awareness.Advances in Neural Information Processing Systems, 37:77132–77164

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.016377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:17689c7e9eaec6381ad6fa25db405be4cb98e3dbf88ee6b59264ecb21d37771c

Observation 715c3ac9-04bc-4eb8-83ba-5f56206943a5 · outbound

This paper cites Effective whole-body pose estimation with two-stages distillation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Effective whole-body pose estimation with two-stages distillation

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.068047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:9e68e56aa9ee86900a4fbd70f2e252575f560706559218a3860847d11c05e913

Observation 6d23955a-9035-441c-8a17-9f761317c7aa · outbound

This paper cites Diffusion-guided reconstruction of everyday hand-object interaction clips.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Diffusion-guided reconstruction of everyday hand-object interaction clips

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.047248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:b4573884cfed0a7bb54a857000a9a887524461164745a2645c3f900e6f568e98

Observation d85924bc-fd09-4bc1-82c1-58bee8dda5da · outbound

This paper cites MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.823368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:72d2f5e4f628d47e67619879895135ce1803abfa3bacfa33de1e4af167bf3b58

Observation 47eeac02-5dbc-4b30-b9c4-37208fa942b8 · outbound

This paper cites OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.839751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:b32f240d6b4a1d125f6bdc9a078981e8f6a47591f6ca05bf213e235c93cf52a0

Observation e641314b-9792-4f9e-b913-99cd95c05601 · outbound

This paper cites Identity- preserving text-to-video generation by frequency decomposition.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Identity- preserving text-to-video generation by frequency decomposition

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.047787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:4cd9b3a800010c2986676b170a1037c21795c739e8d2c89c401df724fe89074c

Observation 2d4ab8e6-f913-4d1e-b7a5-a8b370588c47 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Adding conditional control to text-to-image diffusion models

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.064563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:73359960361c98bff650406203defbc285a6f781f87e3a29952a90ba3d04c503

Observation 199cb090-1855-4413-888f-f4b95dbb5ac4 · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.044544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:d21fd5bb8a9ecadaaf8cbd8e6d9b0be1f03fbf8749f5a420edd615f878724b3b

Observation f57c246d-ff62-4f4c-8b9f-9827500d268f · outbound

This paper cites Waver: Wave Your Way to Lifelike Video Generation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Waver: Wave Your Way to Lifelike Video Generation

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:11:00.740481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:522e9e013745cabb4699b774811fdeb6aab9f3bdbe1d2b2f67b3dfa4ed5c6eb9

Observation 1b6b86ba-0476-4aa6-afc4-89da3b0f88e7 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:15:20.153254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:b6d416163edbc3b2d27e5c7a8b00d3a449b9934116d74a0fb6f0958ce9801950

Observation 07d5f765-e915-4aa0-9739-a987e41c1b6b · outbound

This paper cites Anytalker: Scaling multi- person talking video generation with interactivity refine- ment.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Anytalker: Scaling multi- person talking video generation with interactivity refine- ment

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.783116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:1cec4838965a9f038b31be10500319b25b409459ae94458f3142e858bcafa55f

Observation 366a78d4-111e-4e45-b269-d6f9b701451c · outbound

This paper cites MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation MagicTailor: Component-Controllable Personalization in Text-to-Image Diffusion Models

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.796535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:b26f150827cc038e2a5607978a789c8055d3089e0e89376b328fb6db1cf79fe6

Observation 6777b833-b724-4837-9248-097df5455720 · outbound

This paper cites Identitystory: Taming your identity-preserving generator for human-centric story generation.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Identitystory: Taming your identity-preserving generator for human-centric story generation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T07:42:30.074870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:bd5ea504c42ad5e940c574a085a660e29de4bde4bc7303063fe864472c26964a

Observation 333a1358-7d21-46e6-a6ae-6c04614f6027 · outbound

This paper cites Scaling zero-shot reference-to-video generation.arXiv preprint arXiv:2512.06905.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Scaling zero-shot reference-to-video generation.arXiv preprint arXiv:2512.06905

Reference 79

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T11:06:05.826762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:9ecdd479b8081471405bf17ea4ce8c785847be053041ddbbb2122486e2df9188

Pith citing papers

Observation 588a0fcf-f17a-4baf-8d4d-98bc2e668000 · inbound

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation cites this paper.

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation

Reference 61

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T12:41:04.418452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T03:14:45.834520Z digest=sha256:edae877bd3bbd8440fef5811847994df562a4ba2359ff02d1c98f975416b55e8

Observation b9e24e88-7ce4-47a6-b35e-5728b371dc09 · inbound

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships cites this paper.

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:25.434955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:25.434955Z digest=sha256:27bdac8215ea564d05d3c5fc799d1b17aedd99019c5e118ea9883b674764f105

Observation adbe0d63-f762-49f6-80d5-4e0a436c0e7f · inbound

MoRoute: Dynamic Routing for In-Context Multimodal Video Generation cites this paper.

MoRoute: Dynamic Routing for In-Context Multimodal Video Generation OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T04:58:52.344771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:58:52.344771Z digest=sha256:4e4b9e404e630a05f29b11fa3b5473b2d8460d734744dd92ba1690e6c1bc0d1e