Pith. sign in

Paper Citation Record · LEDGER

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

As of 7 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 4 inbound Pith citation observations for arXiv:2506.12847.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12847 v1

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:41:07.918981Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:23:09.117957Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

86 of 86 outbound references displayed

  • verified exact0
  • verified fuzzy67
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d78cc45-a367-4c45-9b1d-ed589175f3ab · outbound

This paper cites Person image synthesis via de- noising diffusion model.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Person image synthesis via de- noising diffusion model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:55.849134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:55.849134Z digest=sha256:115546fe62ff4b6e3ca93c68d868231925ca7c8d302178cfe5a09a1a304980ca

Observation 87e4ccb9-97a3-4d1d-88c0-96a38758b037 · outbound

This paper cites Videopainter: Any- length video inpainting and editing with plug-and-play con- text control.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Videopainter: Any- length video inpainting and editing with plug-and-play con- text control

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:56.026603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:56.026603Z digest=sha256:9769566322e787ffa95f3eb44712b1f3b5d097e238e1e40c05fe1719edec9aac

Observation 3c87dd3b-a552-442f-a769-09586745fb9e · outbound

This paper cites Smpler-x: Scaling up expressive human pose and shape estimation.NeurIPS, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Smpler-x: Scaling up expressive human pose and shape estimation.NeurIPS, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:56.147395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:56.147395Z digest=sha256:cce7ac4e8f2eae18fe1205f9f08c458ee32955202e23d1a2b1cd286681d2624b

Observation b7b6547e-4d76-439d-8174-ab27619128e3 · outbound

This paper cites Pix2video: Video editing using image diffusion.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Pix2video: Video editing using image diffusion

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:56.276084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:56.276084Z digest=sha256:cf5244557b4c72f89b0644903d035f208433ce1d93a8bbd82e55052fbdc80d61

Observation 3c527140-e1f0-462f-936b-7748adb26f9f · outbound

This paper cites Magicpose: Realistic human poses and facial expressions retargeting with identity-aware diffusion.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Magicpose: Realistic human poses and facial expressions retargeting with identity-aware diffusion

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:56.407455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:56.407455Z digest=sha256:61c4278c4a4df6eff07b4833b7dc5780a84d399742d5c50b7554929a900ffe24

Observation d0cd001b-8d41-4423-92d4-90dcb07eaa0c · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:56.552104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:56.552104Z digest=sha256:4bb829ebdca2ff600b672bc404a26995e9b3a9923d9af574ae9c9941cdd1cf2d

Observation 8eec9d24-25f8-4151-9226-3683a7196b86 · outbound

This paper cites Control-a-video: Controllable text-to-video generation with diffusion models.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Control-a-video: Controllable text-to-video generation with diffusion models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:56.697928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:56.697928Z digest=sha256:9590fc67df8c689730b2b160cb65ee557e837dfcfa468ae08f98bf7d9dc806d9

Observation d4bfc890-7a8e-4d38-a338-c94ad7667d51 · outbound

This paper cites Cg-hoi: Contact-guided 3d human-object interaction generation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Cg-hoi: Contact-guided 3d human-object interaction generation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.736762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:56.763263Z digest=sha256:f6c38355b960220f3e8eed25bc52f5e024c5c5b4b0bf11d78fef60e61c6abb1a

Observation dd4570e3-0316-4bbc-8397-25a22af4e0c0 · outbound

This paper cites Structure and content-guided video synthesis with diffusion models.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Structure and content-guided video synthesis with diffusion models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.717139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:56.841966Z digest=sha256:299bfc5ee3486bfb054a864808cb6fbdc0a4220f28e2a9a1a15f756b534accc9

Observation 28603426-b44b-423b-825c-2a8a5bbee419 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:57.009521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:57.009521Z digest=sha256:cea86d41c7db4c925d6f4318c729ead44cc34b67fd52cee8d638adfd81a666e6

Observation 946c9b23-1f5d-4fc3-805c-9019e4e3eaf0 · outbound

This paper cites Re-hold: Video hand object interaction reen- actment via adaptive layout-instructed diffusion model.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Re-hold: Video hand object interaction reen- actment via adaptive layout-instructed diffusion model

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.682177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:57.115303Z digest=sha256:b8fbe16162e82d606fbd659e134da01f7f8a2aa19dd35539fff9e8b70244cd5c

Observation 444ee6dc-b79b-48c1-b1ee-f76fd6ff1325 · outbound

This paper cites Ccedit: Creative and controllable video editing via diffusion models.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Ccedit: Creative and controllable video editing via diffusion models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.665051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:57.204044Z digest=sha256:57dc8ed1990d2add77bb5e44ae3b703c6f00cf71a17ad133e6b94baf5ed9b92a

Observation 7ac5ae84-74b1-4c0d-8aae-4c50a97ad4c1 · outbound

This paper cites HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer HumanDiT: Pose-Guided Diffusion Transformer for Long-form Human Motion Video Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:57.359048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:57.359048Z digest=sha256:719f144fd890c89e8e06060144f528ebf58d9a4b46e6fc6063320ed870626c0b

Observation 375e4770-abf2-456f-8dc5-8c9e8339deaa · outbound

This paper cites Tokenflow: Consistent diffusion features for consistent video editing.arXiv, 2023.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Tokenflow: Consistent diffusion features for consistent video editing.arXiv, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.646670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:57.530289Z digest=sha256:ff8cf379ecd0ec570b8080b5b54735dda25431a419b850c0ee89b9fa46aa3acd

Observation 891ee0a2-a6b1-4f00-96c4-38ad8a68f67b · outbound

This paper cites Emu video: Factoriz- ing text-to-video generation by explicit image conditioning.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Emu video: Factoriz- ing text-to-video generation by explicit image conditioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.626921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:57.712846Z digest=sha256:b937a904f9b2d2509afa16fd528656ea7b6e5a76b46050faf724976c68d7c0b7

Observation a74efc51-aa1f-442b-b944-fa215e95add4 · outbound

This paper cites Videoswap: Customized video subject swapping with interactive seman- tic point correspondence.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Videoswap: Customized video subject swapping with interactive seman- tic point correspondence

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.606275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:57.888052Z digest=sha256:caf4d3daae6a58ebbcbec0915c8f5be1fbed751272c015242f3bf09c8c143ed9

Observation ba21eacd-fd29-4ff8-b905-621052d961e4 · outbound

This paper cites Talk-act: Enhance textural- awareness for 2d speaking avatar reenactment with diffusion model.SIGGRAPH Asia, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Talk-act: Enhance textural- awareness for 2d speaking avatar reenactment with diffusion model.SIGGRAPH Asia, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.588021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:58.033983Z digest=sha256:c4ea773ecb4ad4f92cb8d830d50c803115aa68ab3613799e56d9c311aa5b8a1d

Observation 8a25d911-f505-423f-8ea9-2127cd6a0025 · outbound

This paper cites Livepor- trait: Efficient portrait animation with stitching and retarget- ing control.arXiv, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Livepor- trait: Efficient portrait animation with stitching and retarget- ing control.arXiv, 2024

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.569441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:58.141288Z digest=sha256:34b5a7fc44ba46cde47221a2ce723c8521168fecb8319382d56065756ab68dbb

Observation 6c07f26e-efc0-4834-a2f8-88d5c99d8dd6 · outbound

This paper cites Resolving 3d human pose ambigui- ties with 3d scene constraints.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Resolving 3d human pose ambigui- ties with 3d scene constraints

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.551542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:58.247226Z digest=sha256:f2c1d0a5d0b5fc48a7229d250785efc315f3abb450bd2ec96dada073070a4b47

Observation 7768c4ab-da41-48e1-a27a-a94e565f3f9d · outbound

This paper cites Populating 3d scenes by learning human-scene interaction.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Populating 3d scenes by learning human-scene interaction

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.528089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:58.327537Z digest=sha256:8146f7f5217a85686cc2b0ba2e8b3d4a3f69932f822059c9c65e554172c254da

Observation f652b488-9eb5-4420-8e2a-b8c3201b3c65 · outbound

This paper cites Synthesizing physi- cal character-scene interactions.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Synthesizing physi- cal character-scene interactions

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.443031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:58.419397Z digest=sha256:241a69106aa321834512362c064c0a6dcf8d6a1ee54d5a3463c762f1f805f2c8

Observation ee101561-8130-4d4f-b99d-e6d58a8674e8 · outbound

This paper cites Hand-object interaction image generation.NeurIPS,.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Hand-object interaction image generation.NeurIPS,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.272435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:58.505914Z digest=sha256:bd606b6b2b293e91441e7db046b6585f1b85f33dc7c3b880bd2e0d7e74531492

Observation c58761c5-83cf-436a-97f4-c6afe99a70db · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Animate anyone: Consistent and controllable image-to-video synthesis for character animation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:18.062117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:58.587901Z digest=sha256:5ce6d961fd3b0573fddccf1a15704b1bcf7cea9b454310fc866c11a5ddbe118f

Observation 42031114-248d-4ed7-9e35-685b36367ebc · outbound

This paper cites Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:58.645841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:58.645841Z digest=sha256:9002872cacd17adf36fd8382b2b35023ea07ff739392deab410fd1025c587267

Observation b1c244ba-1ebc-459c-ab1f-38d2afd9f69a · outbound

This paper cites VACE: All-in-One Video Creation and Editing.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer VACE: All-in-One Video Creation and Editing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:58.747297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:58.747297Z digest=sha256:86498eccfee7fc3525d3a5ad6e4ca9c8e35143a7dd65d6bbeb5ac8d29af9d3cf

Observation 1b1427ab-33b7-4321-9ca6-48409fc65980 · outbound

This paper cites Dreampose: Fashion image-to-video synthesis via stable diffusion.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Dreampose: Fashion image-to-video synthesis via stable diffusion

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:17.923114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:58.861552Z digest=sha256:2a7dbf3f5572ed22a7258b94bca442c8abd54a82a544a19e98ec311142abf427

Observation b7ea3f57-8e41-466c-861b-fa63dce0ba4b · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:58.985574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:58.985574Z digest=sha256:b38afc02ae04b36afabaee4b790c6bd259972a976104b88ad7735805ea676fb4

Observation 3f2c233f-983d-4237-8e3d-570a41636637 · outbound

This paper cites Anyv2v: A plug-and-play framework for any video- to-video editing tasks.arXiv, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Anyv2v: A plug-and-play framework for any video- to-video editing tasks.arXiv, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:17.845182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:59.071513Z digest=sha256:d0785a04224d3102a2cb161ea53d42f0ce42a3ab1503a9764efb40987a829aa4

Observation 5bffd73d-dcee-4468-8fb5-52609dddfbf6 · outbound

This paper cites Flux.https://github.com/ black-forest-labs/flux, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Flux.https://github.com/ black-forest-labs/flux, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:17.650595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:59.129725Z digest=sha256:a9425c5bbe0a02b623cdf93d0e9e1bc79099d04a28a55a695b7cb31e7eb95602

Observation 46857b48-d2cc-43fd-8860-b9fb80935fb5 · outbound

This paper cites Lego: Learning egocentric ac- tion frame generation via visual instruction tuning.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Lego: Learning egocentric ac- tion frame generation via visual instruction tuning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:17.470230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:59.225938Z digest=sha256:e205256504aa98940a2ce34b2b92e4a7ad1a5032163c8836d4390698936f81e5

Observation 5ac1839c-76f6-4fac-92e1-f41c9f028dd3 · outbound

This paper cites Shape-aware text-driven lay- ered video editing.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Shape-aware text-driven lay- ered video editing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:17.322349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:59.323533Z digest=sha256:a3a6cbd5a21f51248469ca7afbd5adc68220f0afe0029472a1a0f12263748ca4

Observation 5f05962b-104c-477a-aaed-7848f3fd0afb · outbound

This paper cites Generative om- nimatte: Learning to decompose video into layers.arXiv,.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Generative om- nimatte: Learning to decompose video into layers.arXiv,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:17.230606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:59.406104Z digest=sha256:657cfca69670a4e05a1635b835536665fe04968a46ddf40e8e7a7af23983a3f8

Observation 20f1775c-f3c3-45ef-a6b6-5ea631e9b977 · outbound

This paper cites Vidtome: Video token merging for zero-shot video editing.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Vidtome: Video token merging for zero-shot video editing

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:17.036408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:59.468695Z digest=sha256:cf2e9d69bf72a517fb09cd989ec8c55aaf09f35947e5b3d1acfa8673a0bf39a9

Observation 133a9f74-e380-4fa8-84b3-bb139730b507 · outbound

This paper cites Video generation from text.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Video generation from text

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:16.739371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:59.569258Z digest=sha256:649006f530bbb9e1e8d1c6d148ca3a9af6686fa4198705c1a56ff5c2522f7135

Observation 4158dac0-cec8-4cd3-bf99-534b47953d49 · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:59.660776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:59.660776Z digest=sha256:5f5c80912d8a2c61165efbc37001e3b9fe93095546749e4d43f4c8ecb881317f

Observation 237fda55-aa45-4e5b-acbc-6a79a7f5f4bf · outbound

This paper cites Flowvid: Taming imperfect opti- cal flows for consistent video-to-video synthesis.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Flowvid: Taming imperfect opti- cal flows for consistent video-to-video synthesis

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:16.482657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:59.767048Z digest=sha256:97d5a262f130afb7531635c4f7299e1ac05c50e5d8707e3f770aa4c1d0c021eb

Observation ccd59e8f-5f78-471c-b09d-8b6dd3b0cb1e · outbound

This paper cites Video-p2p: Video editing with cross-attention control.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Video-p2p: Video editing with cross-attention control

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:16.274883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:59.852408Z digest=sha256:abcd9eb55f7a3ba6cd1141f15756a73522229d11752c4a711bb8bdb4bbda062b

Observation 3a6392ac-1bf7-4d7c-bc07-7e55bdc26d36 · outbound

This paper cites Iterative ensemble training with anti-gradient control for mitigating memorization in diffusion models.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Iterative ensemble training with anti-gradient control for mitigating memorization in diffusion models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:16.035930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:40:59.952831Z digest=sha256:401eda79e101c4b04bd789307b1921adda0a6b5c0664fc5082851fe0ca23752f

Observation 21c8291f-9cfd-47cd-a041-c7b69fe23212 · outbound

This paper cites Live speech por- traits: real-time photorealistic talking-head animation.TOG,.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Live speech por- traits: real-time photorealistic talking-head animation.TOG,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:15.816820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:00.025488Z digest=sha256:af82da518a0843f242bb7ce0373aecda227a4f137bf22c272d432366da26bd6f

Observation d59c705e-ab3b-49b6-9c01-69b88252813f · outbound

This paper cites Dreamactor-m1: Holistic, expressive and robust human image animation with hybrid guidance, 2025.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Dreamactor-m1: Holistic, expressive and robust human image animation with hybrid guidance, 2025

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:15.614113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:00.125115Z digest=sha256:4a8e480df2c7ab988065b2a95a989c3dee4a5e14135d9271b84b8a52457c21e5

Observation e418d60b-fe81-41ae-b4c5-e3e8adca9598 · outbound

This paper cites ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:00.228866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:00.228866Z digest=sha256:55a17d0eda8c6f7f377578e7ec6663ef9d8a293431b50665a7dd8134ae267529

Observation e5ab77f8-b365-4af5-b0b8-fbc15c25910f · outbound

This paper cites Mimo: Controllable character video synthesis with spatial decomposed modeling.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Mimo: Controllable character video synthesis with spatial decomposed modeling

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:15.304385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:00.689236Z digest=sha256:bb44b9438136872c5fce2b9c70e2b770bc299631c86a93a2052e6981293fc7c8

Observation 4d4c7e16-5d70-4c67-8d23-969a26fe7c26 · outbound

This paper cites Sora: Creating video from text.https:// openai.com/sora/, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Sora: Creating video from text.https:// openai.com/sora/, 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:15.006145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:02.570292Z digest=sha256:eb0af248ae0a3cdabb6a2e80573efa87c5bdc4c7010c56f68bc45c4d03211a70

Observation ee9e51ab-dfc4-4569-8929-8ce1c90d0b96 · outbound

This paper cites Reconstruct- ing hands in 3d with transformers.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Reconstruct- ing hands in 3d with transformers

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:14.877718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:02.794703Z digest=sha256:295f228eb7851ba64627525c533b9e25b2252a7206071346069e0047208ff04f

Observation 93e9b51d-433f-4654-a869-9ff43e1212b1 · outbound

This paper cites Vase: Object-centric appearance and shape manipulation of real videos.arXiv, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Vase: Object-centric appearance and shape manipulation of real videos.arXiv, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:14.710627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:04.395040Z digest=sha256:75e3db37eefdca15b64a329667ddbd0b2d3c4dcf9e5aaba5200544793479df66

Observation 04adb67f-2efb-4d33-af6a-4a884744d84c · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer A lip sync expert is all you need for speech to lip generation in the wild

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:14.536687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:04.643291Z digest=sha256:8b902eb9e6c3653effac3fb7fb8d2bd6ab82aa794814107c07737dc5edfb1423

Observation f8e562a0-d805-4eed-97e3-6bea44ee846b · outbound

This paper cites Fatezero: Fus- ing attentions for zero-shot text-based video editing.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Fatezero: Fus- ing attentions for zero-shot text-based video editing

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:14.425726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:04.745650Z digest=sha256:1b52bfa7d38417d932799d2be3dab4300a6522d34e4214db022baf11cba3c75f

Observation 6dfce370-292e-4c4c-805d-fa5807744f9f · outbound

This paper cites Stable diffu- sion 2 inpainting.https : / / huggingface.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Stable diffu- sion 2 inpainting.https : / / huggingface

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:14.274527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:04.868749Z digest=sha256:36eddd6110a4022a2a99da31925e02aa48196f19fe6b02dc00fbfe6ca8391a6f

Observation db8cc9c2-7702-4c71-abc2-e3b9ec5bb039 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer High-resolution image syn- thesis with latent diffusion models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:04.923402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:04.923402Z digest=sha256:473d88f0feb06725136b2b0958bc2cd3d76abf6d7eef12a8c9adf0c255171ce6

Observation 79bedfab-3e06-479d-90e1-5b4817113664 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer U-net: Convolutional networks for biomedical image segmentation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:04.929104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:04.929104Z digest=sha256:9a1ed8a026a9cda1d74a76894bfa9a438b7b98b9e49b6777f46b877c14a54432

Observation 6256689b-ee15-4038-8bda-8e49319670bb · outbound

This paper cites Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Human4DiT: 360-degree Human Video Generation with 4D Diffusion Transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:04.992398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:04.992398Z digest=sha256:0a8040528793cd3e81c8c1f1ea48add3172bd9b311854af376521168ef432040

Observation 1531149b-5832-43c2-a3da-32aad6c2706d · outbound

This paper cites Edit-a-video: Single video editing with object-aware consistency.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Edit-a-video: Single video editing with object-aware consistency

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:14.118624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:05.054638Z digest=sha256:c5763eb247ab790ac68c981ea52cea13eded4c1f7e5c55f4a82e39d852ecfd52

Observation 76a50cca-916c-46f7-8a5d-960b83b9e323 · outbound

This paper cites Make-a-video: Text-to-video generation without text-video data.arXiv, 2022.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Make-a-video: Text-to-video generation without text-video data.arXiv, 2022

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:13.967877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:05.127368Z digest=sha256:c759d2838d0d01c5ee30d6474f8362e08d8c4f769ee4244d0dcae2a1cdf1db32

Observation 4fcbaeb9-bbc8-4323-936a-98225a940114 · outbound

This paper cites Emo: Emote portrait alive-generating expressive portrait videos with audio2video diffusion model under weak conditions.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Emo: Emote portrait alive-generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:13.844013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:05.215658Z digest=sha256:85220e230ed01c1b6e6e79ad1240c09d0c17707902a4c9ed6175081e74de9205

Observation 1833d2a1-9209-449d-a166-9147a385e97e · outbound

This paper cites Attention is all you need.NeurIPS, 2017.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Attention is all you need.NeurIPS, 2017

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:13.725532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:05.282748Z digest=sha256:262f322f66e74b48da73f992484b91792dbeebea1e5e8a9e6721d7b11db124b5

Observation 84e4ab52-6e9c-41ae-ae7f-98caa501a480 · outbound

This paper cites Wan: Open and advanced large-scale video gen- erative models, 2025.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Wan: Open and advanced large-scale video gen- erative models, 2025

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:13.590511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:05.351828Z digest=sha256:04b3f2dcc0cddefcdbe8d3a5340338228eb36e53b27604a9c058a11caab56566

Observation fca04e0f-4ebc-4160-b035-cb2e592d659f · outbound

This paper cites Robust video portrait reenact- ment via personalized representation quantization.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Robust video portrait reenact- ment via personalized representation quantization

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:13.434166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:05.445126Z digest=sha256:181648db0950f8c1817041c626917f37def969f281f18183a25317bd2b3b2d37

Observation 37e5e87b-32f7-40a8-9e2f-4f006ad5689c · outbound

This paper cites Efficient video portrait reenactment via grid-based codebook.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Efficient video portrait reenactment via grid-based codebook

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:13.264151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:05.540143Z digest=sha256:7e7cb3b08f6000ae585958a21d6658f7bdb481b6bb5dcf0c0a7f39c66e7e89ef

Observation 54a117d9-1ed8-4d65-927d-c2d7e24f1d09 · outbound

This paper cites Diffusion in diffusion: Cyclic one-way diffusion for text-vision-conditioned generation.ICLR, 2023.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Diffusion in diffusion: Cyclic one-way diffusion for text-vision-conditioned generation.ICLR, 2023

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:13.133190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:05.609950Z digest=sha256:57f7b3cd609c7ec7fbdacb00c883c9c46c7aedfd9624599a9a9b7b9934b876af

Observation a34be035-95e6-4c89-b858-84db5d41886d · outbound

This paper cites Disco: Disentangled control for realistic human dance generation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Disco: Disentangled control for realistic human dance generation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:12.936533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:05.693536Z digest=sha256:a5d578401fa7061e0d35b9f00a19a0e3d5ec89bdd8df658beb4978770360cef7

Observation 3fbab669-8ebb-455d-8e72-fd981d2c52ff · outbound

This paper cites One-shot free-view neural talking-head synthesis for video conferenc- ing.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer One-shot free-view neural talking-head synthesis for video conferenc- ing

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:12.730327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:05.791022Z digest=sha256:93ed5d9cdfbbe2d20f85472efb359216dc1d01ded59bdb8bf6d5d5116bed6acd

Observation 9b960a2d-fb93-481d-845f-7143d44f64e8 · outbound

This paper cites Videocomposer: Compositional video synthesis with motion controllability.NeurIPS, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Videocomposer: Compositional video synthesis with motion controllability.NeurIPS, 2024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:12.534289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:05.885683Z digest=sha256:bb50d341bd55cabc97c614f892b82ad35f614c07ee3ec16c7a0ced1d989101ad

Observation 98c38943-a261-4899-8dcb-9822c011902e · outbound

This paper cites Unianimate-dit: Human image animation with large-scale video diffusion transformer.arXiv preprint arXiv:2504.11289, 2025.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Unianimate-dit: Human image animation with large-scale video diffusion transformer.arXiv preprint arXiv:2504.11289, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:05.956615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:05.956615Z digest=sha256:c1b6e291d7fcff6dd750721b4f74128c0dd5b60b154bbeae6ae80b68367eb6f2

Observation f9d141e7-259d-49fe-94be-f4cdaf127bd8 · outbound

This paper cites Latent image animator: Learning to animate im- ages via latent space navigation.arXiv, 2022.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Latent image animator: Learning to animate im- ages via latent space navigation.arXiv, 2022

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:12.358387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:06.058039Z digest=sha256:6bbadf454d94db4325540920bcc110c7af6f14b073f3b324d08abdefe0480e7e

Observation 88f8bc3d-fdd1-41cb-abf3-75b1f792f7bd · outbound

This paper cites Foundationpose: Unified 6d pose estimation and tracking of novel objects.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Foundationpose: Unified 6d pose estimation and tracking of novel objects

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:12.187867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:06.158274Z digest=sha256:ee8f990a25741b0b2c48425e0784700205beeaa1d7051c445d4f0aa288b8ebec

Observation 93e611b8-95ff-4542-a080-9dbed1fe7a5d · outbound

This paper cites Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Tune-a-video: One-shot tuning of image diffusion models for text-to-video generation

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:11.988082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:06.227531Z digest=sha256:5ee61f163ce15dd92b9dbe7500adffc93564969626c54ead488dca49ac57f0c7

Observation d884100c-e331-4ee7-b609-eb18e98646e3 · outbound

This paper cites Hallo: Hierarchical audio-driven visual synthesis for portrait image animation.arXiv, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Hallo: Hierarchical audio-driven visual synthesis for portrait image animation.arXiv, 2024

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:11.830962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:06.307583Z digest=sha256:105cdd7be5e299f1462159d7301e3ee6feb82abd66aaf3124f92e4033fcc5904

Observation 3ea10905-0ca3-4cd4-a13b-ac1908019815 · outbound

This paper cites Hoi-swap: Swapping objects in videos with hand-object in- teraction awareness.NeurIPS, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Hoi-swap: Swapping objects in videos with hand-object in- teraction awareness.NeurIPS, 2024

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:11.670573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:06.377089Z digest=sha256:fa64bb837d2eaa112c630c4483f578889d2fd2b9393efeab7272afeb69b266b2

Observation 180c060e-094b-4a4b-a15d-ad6630d416b7 · outbound

This paper cites Showmaker: Creating high-fidelity 2d human video via fine-grained diffusion mod- eling.NeurIPS, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Showmaker: Creating high-fidelity 2d human video via fine-grained diffusion mod- eling.NeurIPS, 2024

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:11.431258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:06.445388Z digest=sha256:eb92bebbe9f554d954962c6b2015d4620086cbedbbbf4a88bfd75917dd8c2111

Observation c6a856d5-2818-4056-95cd-7c779519471e · outbound

This paper cites Rerender a video: Zero-shot text-guided video-to-video translation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Rerender a video: Zero-shot text-guided video-to-video translation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:11.304565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:06.520664Z digest=sha256:ff7abdcb05575562301c18571a1fb016b0572c3a0d6c4917d1a7154b717d55d6

Observation ffe09eaf-ae67-474f-a342-7cac1ae98853 · outbound

This paper cites Effec- tive whole-body pose estimation with two-stages distillation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Effec- tive whole-body pose estimation with two-stages distillation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:11.126525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:06.594525Z digest=sha256:650932db7c9bafad86a9078255b8107479355ca322a12b64969fdc9c6571a58b

Observation dcc5f97e-bee4-4cb1-983c-c7613a11b563 · outbound

This paper cites Space-time diffusion features for zero-shot text-driven motion transfer.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Space-time diffusion features for zero-shot text-driven motion transfer

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:10.950453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:06.674585Z digest=sha256:79fb0afb7740daabe08ef38f30602b1c875d6b47bd88c73845faf30189d9d974

Observation b47d8aa3-a7b1-4180-8aa7-15ad68ed279f · outbound

This paper cites Diffusion-guided reconstruction of everyday hand- object interaction clips.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Diffusion-guided reconstruction of everyday hand- object interaction clips

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:10.782808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:06.773556Z digest=sha256:50a11d035407224ae81c4004a76deab154246ac0f2fc23c35b0bb0fa32ee38cc

Observation d42757ab-d2b5-405e-b095-279a5de21aab · outbound

This paper cites Affordance diffusion: Synthesizing hand-object inter- actions.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Affordance diffusion: Synthesizing hand-object inter- actions

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:10.615828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:06.846684Z digest=sha256:314e4e658a20d18fa50b89b4012a77adcc02e1d461c610589d4496802ec22f92

Observation 2a5871f5-458d-4221-823f-9bf05474adcf · outbound

This paper cites Moonshot: To- wards controllable video generation and editing with multi- modal conditions.arXiv, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Moonshot: To- wards controllable video generation and editing with multi- modal conditions.arXiv, 2024

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:10.436269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:06.915945Z digest=sha256:db0274b4971d21ee0d519802f1b0f3985eca2565e8d144678c92be09df230831

Observation 4ffb5d07-6954-4cfe-9af1-3d347d2e6729 · outbound

This paper cites Graspxl: Generating grasping motions for di- verse objects at scale.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Graspxl: Generating grasping motions for di- verse objects at scale

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:10.220306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:06.985295Z digest=sha256:768feac25e865ff8295ccb06ac7c5a219ce39952ed2708438f3e4765f07eb143

Observation 18df7205-b067-46c6-b67d-2504aa2089dd · outbound

This paper cites Hoidiffusion: Generating realistic 3d hand-object interaction data.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Hoidiffusion: Generating realistic 3d hand-object interaction data

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:10.041631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:07.056644Z digest=sha256:8ae2f4e9cb1da74aa5f39cf48b9bdbe3f505db5a0700fa3e1002f22c9fc3a25d

Observation ee2636fd-962e-40aa-b2af-a097141604fa · outbound

This paper cites Place: Proximity learning of articulation and con- tact in 3d environments.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Place: Proximity learning of articulation and con- tact in 3d environments

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:09.852017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:07.125844Z digest=sha256:98bc8ac4a0543059a5731c811e4c46e5729439700512338b650e3dad2b01d591

Observation 05fe2724-05d6-488f-83ef-168f9d183dec · outbound

This paper cites Generating 3d people in scenes with- out people.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Generating 3d people in scenes with- out people

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:09.673946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:07.207170Z digest=sha256:a763d276b7591be0d737555fe5a0d4f998ae42a85fc2a95faedee34e502cf252

Observation 072c37cb-e248-4725-ba45-2e8b6c1da1bc · outbound

This paper cites MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:07.314086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:07.314086Z digest=sha256:a457212c399400511479641744597fb557fe34d706356ed5ca6997136d952366

Observation 789bac7d-8272-4bab-8e76-4cdd1d12f094 · outbound

This paper cites Avid: Any-length video inpainting with diffusion model.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Avid: Any-length video inpainting with diffusion model

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:09.464243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:07.430241Z digest=sha256:3d3a74ac63b3af593308610449f46dbb08eee296f5da2be49925efb74b05e9a9

Observation 3cbbf123-e7e0-49fa-8dac-c73266f3dae7 · outbound

This paper cites Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:09.247269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:07.513431Z digest=sha256:0b1f686c4442aa265e76e749851e923782413b711f8d3488c45f220773c838ef

Observation ae7da5e1-6bca-4c2c-9f02-5fcb4cafb4b8 · outbound

This paper cites Realisdance: Equip controllable character ani- mation with realistic hands.arXiv, 2024.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Realisdance: Equip controllable character ani- mation with realistic hands.arXiv, 2024

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:09.007311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:07.594471Z digest=sha256:e02eedbd963db0c6eec48a7e5a4a689a7b45b967e01382c6f99dc628377b4cc2

Observation a4e8ed7d-5a25-419c-ac8b-949aafbc6019 · outbound

This paper cites Discrete contrastive diffusion for cross-modal music and image generation.ICLR, 2022.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Discrete contrastive diffusion for cross-modal music and image generation.ICLR, 2022

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:08.827167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:07.671543Z digest=sha256:95b7fc6ea678badc0f499ff25ce5856521ea010e723629713cef5213cfa17227

Observation 9d0c6260-f658-4548-b34f-d47e83a80825 · outbound

This paper cites Cococo: Improving text-guided video inpainting for better consistency, controllability and compatibility.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Cococo: Improving text-guided video inpainting for better consistency, controllability and compatibility

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:08.662632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:07.779574Z digest=sha256:f5348fdb9e0bb337f60f3fe45bf6cc4ce5aff9f2d7ed9002cbd7c57f220508d2

Observation 500b1a36-b8f8-4578-88ea-446576a18f7f · outbound

This paper cites Cut-and-paste: Subject- driven video editing with attention control.Neural Networks,.

iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer Cut-and-paste: Subject- driven video editing with attention control.Neural Networks,

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:41:08.442439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:41:07.918981Z digest=sha256:5c5b09ae740372f8e0b3fab7c83eb838dcd3f0411b2a01f1c32a0bcfe1587f49

Pith citing papers

Observation 7addece3-57dc-4b59-ad74-400167705ffc · inbound

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation cites this paper.

StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T10:37:54.999164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:37:54.999164Z digest=sha256:4ce305c10d0c572dc3414dabe6da035adb0887c2df64d4719285d1d0433d7491

Observation 7100c40b-5eb5-4d71-b27d-1189285c9be0 · inbound

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment cites this paper.

AgentHOI: Multi-Agent Reasoning for Human-Object-Interaction Video Generation via Implicit Representation Alignment iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T05:27:18.894578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:27:18.894578Z digest=sha256:60798af91743eab758818351ead50201c28e402dbeebb861852677e8d40ad008

Observation f587cc87-3ff3-4bb0-8e72-3fd690d0778d · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

Reference 205

Resolution
unresolved
no resolver link, observed 2026-07-31T08:51:27.095194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T08:51:27.095194Z digest=sha256:71816e7c1e29048878a1e4f108285a49f431c0b69eebd8cccb2a3affee11d14d

Observation 30d0897e-a33f-44be-8a63-bbebb729aaee · inbound

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer cites this paper.

Hand-Object Interaction in the Age of Large Foundation Models:Reconstruction, Generation, and Embodied Transfer iDiT-HOI: Inpainting-based Hand Object Interaction Reenactment via Video Diffusion Transformer

Reference 187

Resolution
unresolved
no resolver link, observed 2026-08-04T01:23:09.117957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:23:09.117957Z digest=sha256:32e5e8f4cfbb70760b741d9df09d71864fba1fb9191ba7cefded7b884a14cbba