Pith. sign in

Paper Citation Record · LEDGER

GenHSI: Controllable Generation of Human-Scene Interaction Videos

As of 5 August 2026, this Paper Citation Record lists 100 of 109 outbound references and 2 inbound Pith citation observations for arXiv:2506.19840.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19840 v2

Coverage vector

measured 100 of 109 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T07:25:41.219753Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-16T23:30:14.969895Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-16T23:31:22.000482Z

Reference resolution

100 of 109 outbound references displayed

  • verified exact37
  • verified fuzzy52
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch6

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2d686a18-97ad-4803-ad5c-e21ec9abcca0 · outbound

This paper cites https://huggingface.

GenHSI: Controllable Generation of Human-Scene Interaction Videos https://huggingface

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.391897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:615ef6097d1ee94511be755296d5d1e46e46bc8bc7d1a34dc0fbc4de981e96d9

Observation 6aef5841-5d96-4d7f-bf92-0beee8b0ee9b · outbound

This paper cites https : / / klingai.

GenHSI: Controllable Generation of Human-Scene Interaction Videos https : / / klingai

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.383113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:5948fb38ce38c47ef42554d5c0533cc52c87fa21e8d7ceefd1a8193571465386

Observation ae8cea70-f067-4eee-a58e-543e57456759 · outbound

This paper cites https : / / klingai.

GenHSI: Controllable Generation of Human-Scene Interaction Videos https : / / klingai

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.370562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:64414a9de6e4668c0eb006107c7b8de3c3a0489d7be3de820f575588685e4711

Observation e94bc09f-f9df-4fa8-9b1d-fe8372b631f3 · outbound

This paper cites Circle: Capture in rich contextual environments.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Circle: Capture in rich contextual environments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.387795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:2632125ea080542073857dac8af8199c3e045379a3374f8d148a5b3d7f004dab

Observation 79e7acd1-7de7-44d2-a56f-9471e8513031 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:27:09.017753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:0dbfe3e10f45346472da23541e9af9b017222c8e3e33a303844b89faf50956e7

Observation 0b34be10-717e-434d-b030-c5ffc30ce52b · outbound

This paper cites Align your latents: High-resolution video syn- thesis with latent diffusion models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Align your latents: High-resolution video syn- thesis with latent diffusion models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.380862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:eb3b8ea1a7bdf92efe03197f58a9e52d6462c6c851f2d3c853ba644488f94366

Observation 8e7dfc2e-fee8-46e4-b10a-53b6c422f84d · outbound

This paper cites Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Go-with-the-Flow: Motion-Controllable Video Diffusion Models Using Real-Time Warped Noise

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.987121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:074cd88bf82909d66c8d945203bcf1ecc596907c821e0f12cbaad91cedec6c55

Observation 64831998-ba09-4e44-9bf8-46a485de7063 · outbound

This paper cites DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T07:27:08.888669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:d97b591e2a76cd7a6b21c46eaea73abcb9ead78c6468f2eddeec32bd6b717073

Observation 18ccce00-371a-4fde-b02b-54c13375c4f3 · outbound

This paper cites Wang, and Gordon Wet- zstein.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Wang, and Gordon Wet- zstein

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.372865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:abfd0dbdde23a463161e5a3d95cb719b1a541be45dcf6066dc4eea72dab0d081

Observation 01b88b08-d983-4942-ab93-41a85ebec5fc · outbound

This paper cites Gen- erating human motion in 3d scenes from text descriptions.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Gen- erating human motion in 3d scenes from text descriptions

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.432964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:4d326495f204cc604c0c0a21166ea37a84a385beb51b566dab353832cfa0dbd5

Observation d1813d2b-2b25-413a-98d9-b96938e3bb6f · outbound

This paper cites Videocrafter2: Overcoming data limitations for high- quality video diffusion models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Videocrafter2: Overcoming data limitations for high- quality video diffusion models

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.434985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:158f5c84224ceb97c0367c1b02c3bafa3438cae79d583c02f201bf90d98b752d

Observation d34ff940-edb5-4b82-ba0b-7606b5f9523b · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

GenHSI: Controllable Generation of Human-Scene Interaction Videos PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T07:27:09.031158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:058290861444f47da58e18b3e00a2ce44501e2414349356dd697a78fee1d6e86

Observation 8980f562-7d33-4e49-909f-b8303dd47069 · outbound

This paper cites FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos FoundHand: Large-Scale Domain-Specific Learning for Controllable Hand Image Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.950681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:1aabb0daaa47dc132b3de6393c174af2a9315d0833a7d6f97844caaf45e2a1f1

Observation a41d63e8-a50f-4d6d-b081-8964572aa162 · outbound

This paper cites Multi-subject Open-set Personalization in Video Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Multi-subject Open-set Personalization in Video Generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.962058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:abe7022d21a526862b2e8d58d795eb8f81c25afddae766cd2fafa9799faa15a0

Observation 95f81163-0c62-48c3-b556-f280a720f4a5 · outbound

This paper cites DreamCinema: Cinematic Transfer with Free Camera and 3D Character.

GenHSI: Controllable Generation of Human-Scene Interaction Videos DreamCinema: Cinematic Transfer with Free Camera and 3D Character

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.924583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:b30a74da2fe5b4ad1ed16b84da1012191a181182ab1defa2bc4f73c5da51a74c

Observation 6d017889-c582-4163-9e44-3fa9aad3cf04 · outbound

This paper cites Style in- jection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Style in- jection in diffusion: A training-free approach for adapting large-scale diffusion models for style transfer

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.407942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:4b0ec7043caeb4ed19814b5e6ab243ffd3b7440934f325b15de2d011d29d43b7

Observation ea15e6f4-ff2a-49bb-a810-e591ee266e7d · outbound

This paper cites LaserHuman: Language-guided Scene-aware Human Motion Generation in Free Environment.

GenHSI: Controllable Generation of Human-Scene Interaction Videos LaserHuman: Language-guided Scene-aware Human Motion Generation in Free Environment

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.928143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:8d0e6c64580c240b25c9562c9805241ddd840fb3e41676e646830510c1d92c36

Observation c5853e52-9dc2-49d6-9516-9bf42fe391bc · outbound

This paper cites Dragvideo: Interactive drag-style video editing.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Dragvideo: Interactive drag-style video editing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.415913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:8da21cf27c37794a2e635b38487b8edd4a0be183894d29a888dd27d1c962662f

Observation 4735d37b-4b3c-40b7-ac2b-de5f38da1b2b · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.460875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:40d31a5a569bada18f1b246f8a84c4fbec685946fa24b46ff7c79f0aad0637e7

Observation 38402460-8972-4a89-b6c1-797260cc1575 · outbound

This paper cites an unresolved cited work.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-19T07:27:09.402967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:fb0a3e3e0dc7cae38ed21988863091e1e7ac552ddb04a6f98237f755abe4ebe1

Observation 44297a95-be3f-42aa-a795-984e0ec19c5f · outbound

This paper cites Motioncharacter: Identity-preserving and motion controllable human video generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Motioncharacter: Identity-preserving and motion controllable human video generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.943609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:848187383833232b0dac1f076437e1f33a783292afb15f79414eb4592003117d

Observation 5833cf9c-ea0e-4167-b952-ddf3442a58a5 · outbound

This paper cites DreaMoving: A Human Video Generation Framework based on Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos DreaMoving: A Human Video Generation Framework based on Diffusion Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.903461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:f850f9ffe089af03b68495a024152a5ff5dc8a7beca752d1b576f63b276705ce

Observation d4a87c14-ff9c-4fb5-992a-95da6d00b502 · outbound

This paper cites Hu- mandit: Pose-guided diffusion transformer for long-form human motion video generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Hu- mandit: Pose-guided diffusion transformer for long-form human motion video generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.424764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:3f19aa7dd66c33a85bb9a6644289c34f0cd396f6349505bc0deb7c78275772ac

Observation bde0305e-a6aa-4411-b26e-5ab1ff159543 · outbound

This paper cites Preserve your own cor- relation: A noise prior for video diffusion models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Preserve your own cor- relation: A noise prior for video diffusion models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.395903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:a59d4e762bf238439a0b55b05883b0dde482e63b63af94f5e4fb0b71a61b6ef7

Observation 1fe104f2-9d30-490c-a21c-dbc76fe11627 · outbound

This paper cites Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Diffusion as Shader: 3D-aware Video Diffusion for Versatile Video Generation Control

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T07:27:08.870018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:9c23cc57d596ee10979e0b47070a56ecabc8da8a05268244a3824d9012b8a8d2

Observation 6de8377e-6e7e-471a-99e9-7365c9a8aed0 · outbound

This paper cites I2v-adapter: A general image-to-video adapter for diffusion models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos I2v-adapter: A general image-to-video adapter for diffusion models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.470482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:75ea2a13afc8f401d200cc579dc43c77ee08b88544f82771ecf64fe39930275d

Observation 57fbd5a5-c7e7-4861-98c4-c4c292d21c2f · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

GenHSI: Controllable Generation of Human-Scene Interaction Videos AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T07:27:08.982745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:7a55c2531569fa850314a6ba76a8648d16b66248a22b2b2e4d1a83b5b697d884

Observation ec21d92d-03ff-446c-978f-9944d80f78c4 · outbound

This paper cites Human poseitioning system (hps): 3d human pose estimation and self-localization in large scenes from body-mounted sensors.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Human poseitioning system (hps): 3d human pose estimation and self-localization in large scenes from body-mounted sensors

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.448646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:a51e4a8406727de14a8994be0343fde0e4a5eaa43ec42795c00b76222c8f5cdf

Observation 6eb35fcf-5690-483d-b0ad-2b267534a11b · outbound

This paper cites LTX-Video: Realtime Video Latent Diffusion.

GenHSI: Controllable Generation of Human-Scene Interaction Videos LTX-Video: Realtime Video Latent Diffusion

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:27:08.881798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:6ce8f49935914a6617dc4a7598dee81b77e4a56218a123a21bdd1c0af44e5c6b

Observation f880363c-9778-4dc1-a132-03e48c0fac0c · outbound

This paper cites Resolving 3d human pose ambiguities with 3d scene constraints.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Resolving 3d human pose ambiguities with 3d scene constraints

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.442835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:a39b8b0c3295a82830d80ea0e01ec8d840b49f7d49ad87124f04b6e7c24b341d

Observation 29faf151-4e34-443f-a81c-c327a84bd195 · outbound

This paper cites Cameractrl: En- abling camera control for text-to-video generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Cameractrl: En- abling camera control for text-to-video generation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.438862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:bd82974250b5e888b177512ecbf793968258a6c239fadea77552d78c05b3f30c

Observation 0ee1e086-8eb5-41b6-bc25-094c13d7ab0e · outbound

This paper cites Denoising dif- fusion probabilistic models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Denoising dif- fusion probabilistic models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.352457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:2df633bd447caee7afbd6074b24e3ca0cdc47a3d17504171d573689be8612dd7

Observation 72adb8c4-cf57-42a5-950b-9c94279a883e · outbound

This paper cites Video dif- fusion models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Video dif- fusion models

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.378506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:831067f1bf64094d567d8a3e307c42b8647b61b67c54df2710f19d0f64c273ab

Observation d8dfb9d3-1e8f-49aa-b4a1-0c68a32fd565 · outbound

This paper cites StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration.

GenHSI: Controllable Generation of Human-Scene Interaction Videos StoryAgent: Customized Storytelling Video Generation via Multi-Agent Collaboration

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T07:27:09.047774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:e8922e4c38bb0181da6f32eccb54f279622546d1a9fe0c383e150c6067802a7b

Observation d8d95c5e-a9f6-46e1-8879-6bd7835672af · outbound

This paper cites Move-in-2D: 2D-Conditioned Human Motion Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Move-in-2D: 2D-Conditioned Human Motion Generation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.026206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:641fe24808961c57984f00b1b511bd4b3a5e06254c7f9a862c17ead7e068406f

Observation 0a703d01-3746-4922-a8fa-52d7f51a99b3 · outbound

This paper cites Diffusion- based generation, optimization, and planning in 3d scenes.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Diffusion- based generation, optimization, and planning in 3d scenes

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.400573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:2eb0c53ee61a01a6b5241be9475f0692ce5cd8cec6eae810dc146a47a78b45cf

Observation ab80ed95-0cf1-4260-bb39-cf4bdd53d804 · outbound

This paper cites Owl-1: Omni World Model for Consistent Long Video Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Owl-1: Omni World Model for Consistent Long Video Generation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.914073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:3e79b99c56dbce4956db239ffc4ef96c92fe95d91abf08a7ba7240f71863dfdf

Observation 5aef4116-47de-4fd1-8ffb-eab323b57762 · outbound

This paper cites VBench: Comprehensive benchmark suite for video generative mod- els.

GenHSI: Controllable Generation of Human-Scene Interaction Videos VBench: Comprehensive benchmark suite for video generative mod- els

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.338032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:609aa7bf94eb2282e6b1d4b2365b543bcf1fe0befc35e5b3d8688eb46d82d581

Observation 8c7ea417-d225-4159-9b03-a54064538929 · outbound

This paper cites Peekaboo: Interactive video generation via masked- diffusion.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Peekaboo: Interactive video generation via masked- diffusion

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.355593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:25b8c753e9ef6ccb0d2fd021223d4ab33aca5cc0e5b4abe017fa48d9ac108374

Observation 1abb434e-c6e7-4f68-9e21-90ea0fc7fec9 · outbound

This paper cites Scaling up dynamic human-scene interaction mod- eling.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Scaling up dynamic human-scene interaction mod- eling

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.349819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:192328f14daf204a7861727226fb6557e7cc37ef3c3007ba51ffa3eb09487a00

Observation d75ab0ff-7002-442f-8baa-d0364867cf24 · outbound

This paper cites Story-adapter: A training-free iterative framework for long story visualization.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Story-adapter: A training-free iterative framework for long story visualization

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.969859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:247d508558828c9cc9b27c18c97469c3dc6d692bf45068c07ff12485724a6c71

Observation 7fb8259d-5cb8-4452-a262-6136b00325dd · outbound

This paper cites an unresolved cited work.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-05-19T07:27:09.347205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:cc023377019ff54bab7a3b8e2675b7adb0a775587f3b41c11ae101add7c0ddd1

Observation 070fbb13-03a5-4e37-aef3-d79af46ed906 · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.

GenHSI: Controllable Generation of Human-Scene Interaction Videos 3d gaussian splatting for real-time radiance field rendering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.430996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:fb526e1c91d3cada5e4916cfe390bbde0ca82463649d28987c42499e778149e4

Observation 0536bf5d-27ba-4e47-b7e2-bf5ad2f02654 · outbound

This paper cites DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.899329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:d11bb3ca13be8d87b3ef30911ce4e9d5e65810524943cb55dc3b6b98c256ec09

Observation 03e928e1-e6fa-47f9-b330-463f389b43dc · outbound

This paper cites Segment anything.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Segment anything

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.444712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:1577ee368844c5675184f8bfb80c54600616563a5433e479905b24ea5a59bfdb

Observation e4bc4dcd-92d7-45d4-b02a-4937a5878675 · outbound

This paper cites Putting people in their place: Affordance-aware hu- man insertion into scenes.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Putting people in their place: Affordance-aware hu- man insertion into scenes

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.440806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:f48dabc0f5ad4c2e7fecf32edd0cdbdac3d288ee6a981a85c06c4e0a4a3dbaa4

Observation 296e44be-638a-4080-b07d-a60c071a34bc · outbound

This paper cites an unresolved cited work.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-19T07:27:09.335703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:c030cbef3402a19be715555e18c5103071c6c6958fc48a54e7b526f3f0ab65b8

Observation d12c0e47-faf6-40eb-b8fd-ec341830cc80 · outbound

This paper cites ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.005765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:3dcd3338292af6007226f1b102341683604ecaf1e12ce13c84eb03e2bad547a3

Observation 8535c6c0-ccc4-4ca8-92e7-d29bd5452c65 · outbound

This paper cites Hybrik: A hybrid analytical-neural inverse kinematics solution for 3d human pose and shape estimation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Hybrik: A hybrid analytical-neural inverse kinematics solution for 3d human pose and shape estimation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.342596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:dd9330907d140812e0ecddc8ed1dc02ef779120b88f75c8eab3eef7f667f927d

Observation 1849a30c-4223-4bbd-8554-76fc0c4a368f · outbound

This paper cites HybrIK-X: Hybrid Analytical-Neural Inverse Kinematics for Whole-body Mesh Recovery.

GenHSI: Controllable Generation of Human-Scene Interaction Videos HybrIK-X: Hybrid Analytical-Neural Inverse Kinematics for Whole-body Mesh Recovery

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.966145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:b93b4c9586f56c1a9887e1105c711e40e172a3e04f0f43007e095f41c7ea2f22

Observation 4a4af4a0-1bae-4dbf-8d40-ea9a1363cc3b · outbound

This paper cites Genzi: Zero-shot 3d human-scene interaction generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Genzi: Zero-shot 3d human-scene interaction generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.364710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:c92eff5b40d516bacab8933d71fbda372b7cdf09294b9dac7d8a04aacab35ff8

Observation f588f632-89b7-4004-80a6-2383245b6ef4 · outbound

This paper cites Animatable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Animatable gaussians: Learning pose-dependent gaussian maps for high-fidelity human avatar modeling

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.458737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:c4025e5326dd4cc6e46f586f997bb3ddb10b188a3a0b903407cd0000bf080d57

Observation 115bb29e-130f-4e5b-bd3f-2b358d905c90 · outbound

This paper cites Intergen: Diffusion-based multi-human motion generation under complex interactions.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Intergen: Diffusion-based multi-human motion generation under complex interactions

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.456774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:a2852b777e2188c0d810ec71d93e4f15d03cb5180dbbe11625a6756858dfc4d9

Observation 34ca8adf-9162-44e3-8901-b29d53285c1a · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Open-Sora Plan: Open-Source Large Video Generation Model

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:27:09.022234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:4c06c015818ec21708c56348d7052d027828f4191028d1e6607364b6f44c1a0b

Observation 189a6fe0-a962-4f45-b246-e815895e0b69 · outbound

This paper cites Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.462806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:7e175436feb2c99b5d2f82cf67eee4df5fb5bbeca574d2b391822bc77b246ae5

Observation 5f29e41d-f228-456c-88b1-cfec2d8fa30f · outbound

This paper cites Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Separate Motion from Appearance: Customizing Motion via Customizing Text-to-Video Diffusion Models

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.001999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:0ea2bd47df7586b4a8224de79c88f18c76a7026303296a0af85eebf80deac6cc

Observation 1be80dee-8757-4f06-a69e-df212fd69564 · outbound

This paper cites Phantom: Subject- consistent video generation via cross-modal alignment.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Phantom: Subject- consistent video generation via cross-modal alignment

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.452798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:a4fd78e9e0d555ce125325709b60691ddace0cdb99242a643a094e972e6bb65a

Observation 2d7f362f-8694-427e-b6e4-b59211684046 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.454779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:3089790254c97c4c0e20bceba5154805da81ce2724255755ad63703de024da89

Observation 39061465-73aa-4e5a-b3a5-af5efd170eed · outbound

This paper cites an unresolved cited work.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-05-19T07:27:09.358337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:3b320d9deb0932f31a58a29b51e38da814cffcca69e7558c5c76d06b0d7f0dc6

Observation 96cf5621-27a3-4696-bea5-df3c0801f1de · outbound

This paper cites Smpl: A skinned multi-person linear model.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Smpl: A skinned multi-person linear model

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.446574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:7518492b0d49d8e58ed7c360305f773d079d8ceb7e568e7cf8de5355bdba9409

Observation 0d4a097e-ebf6-431b-9307-90e7333825a0 · outbound

This paper cites Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:02:24.560574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:7a6e83bb6aeef2307af920bb78a1bdf56c5c9737fa625d1ffbf81b24700d38ab

Observation 1d152e8b-37c2-4b57-bcc9-ad345fa439e4 · outbound

This paper cites Trailblazer: Trajectory control for diffusion-based video generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Trailblazer: Trajectory control for diffusion-based video generation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.464690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:f83433a843eebffb3b14310ffcef008fb5c52e64dfacf74aa975880232312240

Observation 16e914aa-bf57-4099-b717-a41c5f3bbeab · outbound

This paper cites Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Cinemo: Consistent and Controllable Image Animation with Motion Diffusion Models

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.974069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:b292a01e7735831b7d5e9c78d716152398daaaeee638b99c356c9512d2b367f8

Observation 4bfec488-8c34-49e8-b044-e6fbfcdd13ee · outbound

This paper cites GenHeld: Generating and Editing Handheld Objects.

GenHSI: Controllable Generation of Human-Scene Interaction Videos GenHeld: Generating and Editing Handheld Objects

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.990992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:aefad644714ef9540a7d32baf78005bb476ca3d637f6e118cd496bf7a15c2782

Observation 300ce8fe-0613-4839-9536-e790d8bb5566 · outbound

This paper cites Chatgpt-4o.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Chatgpt-4o

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.466611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:57b773a40f8f2d22925ca023de5b60625455714ee3df89ce5afc3bea57d6c390

Observation c4e4a262-85db-4ee2-8300-f1c4f750e3ca · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

GenHSI: Controllable Generation of Human-Scene Interaction Videos DINOv2: Learning Robust Visual Features without Supervision

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:27:08.873264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:288cee1a9c841362d1cc008412e8e0d012501216cdeb6cbc9c0255be0051b149

Observation 0aa829ac-d6d5-4bb9-a7ca-3fedc685819d · outbound

This paper cites Text2place: Affordance-aware text guided human placement.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Text2place: Affordance-aware text guided human placement

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.389822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:1805d7f9405bf43f296ba69d76ce54b409d28f0f37fe8d6035898257600cfd6a

Observation 71e92500-4790-482e-923f-fd7e378ded66 · outbound

This paper cites Expressive body capture: 3d hands, face, and body from a single image.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Expressive body capture: 3d hands, face, and body from a single image

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.450669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:c364c3aa5344cfa377d7c23a55639022f0ca37a2afd3f4ce9f0e0a20a8426b20

Observation b0877edb-f0a7-4c78-947a-deec8120059e · outbound

This paper cites an unresolved cited work.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-05-19T07:27:09.393722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:4e08868e025b42fa036888fbdd590a6c90ab4ccc2c68f296868f5d51a9206587

Observation 5591942f-fc0b-4541-b148-251f5b228a7d · outbound

This paper cites HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos HOI-Diff: Text-Driven Synthesis of 3D Human-Object Interactions using Diffusion Models

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.035768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:903af527a2dbd0eac572c9882d26a4e311a871cdf74af574c127a01243b520e1

Observation bf2fba20-6a2a-449f-b352-c2ea3df051f5 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos High-resolution image synthesis with latent diffusion models

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.385354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:e78c958e8248e10fc9225f33f566b813a7bb45c519d2e8608518c2f5a2c518e3

Observation e902d4ca-924e-4236-aabc-db5fc03a42e6 · outbound

This paper cites Dreambooth: Fine 11 tuning text-to-image diffusion models for subject-driven generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Dreambooth: Fine 11 tuning text-to-image diffusion models for subject-driven generation

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.375906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:fae9d5b4b807fe7fc2bacfdaac8020f12d67e02ce12aa072484b7d6b7eb53018

Observation 22bc4474-9940-4e22-83ba-eb914b642d16 · outbound

This paper cites Magic Insert: Style-Aware Drag-and-Drop.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Magic Insert: Style-Aware Drag-and-Drop

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.009377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:d47bd605d731ddca77059d0bff64e447ac6f0ae359fbad220fcfe03f17cc49b6

Observation 81b99bce-39b2-49de-8bae-4f071099b6c4 · outbound

This paper cites GeoDiffuser: Geometry-Based Image Editing with Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos GeoDiffuser: Geometry-Based Image Editing with Diffusion Models

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.043662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:1e084e527a3eeae70ab318ac031068e4bc98d39b2c37a39101710b7f971b4153

Observation 7bacf2e7-fe99-4746-8a9e-19a8c7e2727c · outbound

This paper cites Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Motion-I2V: Consistent and Controllable Image-to-Video Generation with Explicit Motion Modeling

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.917643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:2af4ab22457790d83efbb0f454294b905caea670fae66c266c636ca45f631766

Observation 7efac2ba-12ec-46e9-b146-7442b8da72ea · outbound

This paper cites Dragdiffusion: Harnessing diffusion models for interac- tive point-based image editing.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Dragdiffusion: Harnessing diffusion models for interac- tive point-based image editing

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.367844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:9194c12c82d4439903742310fb24a950b85f1a0402b38fa95c679850721401cd

Observation f7b39e52-d76b-492b-9d47-b0adf1258b4a · outbound

This paper cites Denoising Diffusion Implicit Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Denoising Diffusion Implicit Models

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:27:08.920941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:fb17bfa99d55d71648a046a35d75b0af98482a8a5f48d578d5d22afc57bbc2a4

Observation 4620c837-2869-44c6-af38-f5cd775c838c · outbound

This paper cites Sound to visual scene gener- ation by audio-to-visual latent alignment.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Sound to visual scene gener- ation by audio-to-visual latent alignment

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.340214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:af50f96a3589e4a9c7a1137b7af90d469ab9ee48ab8fd7c6d81b1eebeab50b25

Observation e1b32098-9779-423d-82a5-4c6ce7bf9df8 · outbound

This paper cites Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Add-it: Training-Free Object Insertion in Images With Pretrained Diffusion Models

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T07:27:09.013480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:f7bcb946ab97995ec5e37a928fcb8757f71c8e70e10cbc8a0a541a9dafaa12c2

Observation d2d1e529-6b84-4fb0-92c2-28dede2602f8 · outbound

This paper cites LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity.

GenHSI: Controllable Generation of Human-Scene Interaction Videos LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.877210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:9f83026b694dedb99f6117d48aea53660309fe804c934a0a7c36348d5272c8c4

Observation c3f32616-caa4-40b8-a3aa-2f27f420d4f9 · outbound

This paper cites Motion Inversion for Video Customization.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Motion Inversion for Video Customization

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.946885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:ea5d5f11dfa2630a4c0813c44ccfb28e6849f404b73a765f2bf9b004965c50d2

Observation 925e6259-f10e-4e01-9422-6020090b9be7 · outbound

This paper cites MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision.

GenHSI: Controllable Generation of Human-Scene Interaction Videos MoGe: Unlocking Accurate Monocular Geometry Estimation for Open-Domain Images with Optimal Training Supervision

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.895161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:f7f20162cc9b9d8a0ababc9ab8f1b995c74c24a28128c53fe274f0a426b6f9e3

Observation 50ce6ba8-6eed-4fe4-a2ae-a9a976fec222 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.958137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:17341919540f7ff0a496f473bcffe0855613a5306d366a7723b3b5ab8901c096

Observation 27d965a2-5d41-472a-8f55-7d87a2629c64 · outbound

This paper cites Humanise: Language-conditioned hu- man motion generation in 3d scenes.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Humanise: Language-conditioned hu- man motion generation in 3d scenes

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.420514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:7889559b6e99c8bb0257c392f622f3bbdf208cf6eaa32e7f4bf1981ca26b1f31

Observation c249ca28-d251-4837-b1b9-5beb53033a29 · outbound

This paper cites Motionctrl: A unified and flexible motion controller for video generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Motionctrl: A unified and flexible motion controller for video generation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.468603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:0f375291511c5601645a259c1fc87d3c0ea0821673ade959d78c629482e8862d

Observation f9bb6f49-3570-46f7-8380-9ba05f467d6f · outbound

This paper cites Move as you say interact as you can: Language-guided human motion generation with scene af- fordance.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Move as you say interact as you can: Language-guided human motion generation with scene af- fordance

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.426869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:c72a703506657f0fe5536cb2d60b9d4e98f0150fdab0fd08ee850ae4135c7db1

Observation e9760be3-30a2-4784-9777-75f367f9a1bb · outbound

This paper cites Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Orient Anything: Learning Robust Object Orientation Estimation from Rendering 3D Models

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.891900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:26a48f67df461a1de8caa975dc1634f8baa9478b8cd4f4c77d8cf46a4e87d866

Observation 88e19441-c22e-44dc-858e-1a9cb996cf0d · outbound

This paper cites Dreamvideo: Composing your dream videos with customized subject and motion.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Dreamvideo: Composing your dream videos with customized subject and motion

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.429085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:53d39a5cd43cdaa456af0a0026ec4ecf7deae6e54e33ce8125c63aaab74f0588

Observation 64cbe6db-137f-44c0-9de4-4844d727dc9f · outbound

This paper cites Detectron2.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Detectron2

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.436871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:a30644cf9f39d9435eeb392872e1af4b2185ff36b311f7221617cb0cb137b163

Observation 43fc0743-f735-4ecd-87f2-75b0dc7b0efb · outbound

This paper cites Mind the Time: Temporally-Controlled Multi-Event Video Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Mind the Time: Temporally-Controlled Multi-Event Video Generation

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.907092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:7c0f2e91fe2b86309ebe1ddb0eb3b33dcca7589d6953b818cb5b5d7ee1da8ae7

Observation 8290e019-f9b5-49cc-afa5-b322e83843f7 · outbound

This paper cites Structured 3D Latents for Scalable and Versatile 3D Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Structured 3D Latents for Scalable and Versatile 3D Generation

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:27:08.978037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:8a70a0977a4bc90b9cf7fc6c865f28ad75a32c67c084e08e1328cceffb2ea20a

Observation 49fb1186-151a-4121-86a4-0fe4be1d0868 · outbound

This paper cites VideoAuteur: Towards Long Narrative Video Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos VideoAuteur: Towards Long Narrative Video Generation

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.885606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:8994eb0323d24a8834668218b09fded9ba77666af5afbbda5eae82c375be6b7f

Observation 43bde310-09d2-4678-88eb-0cfe5a0906bf · outbound

This paper cites DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors.

GenHSI: Controllable Generation of Human-Scene Interaction Videos DynamiCrafter: Animating Open-domain Images with Video Diffusion Priors

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.040021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:d210346dee1479c2a5dd94fa9a0a72f63ad1c16e99bf44a7b26c11da54993856

Observation ebf16d2e-4af2-4012-8a34-6ea564663f14 · outbound

This paper cites Dynamicrafter: Animating open-domain images with video diffusion priors.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Dynamicrafter: Animating open-domain images with video diffusion priors

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.422556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:af22ae8052adfa6e38353c173877bf6224579eae04b0be87c7a14c9043044b78

Observation 37457f85-25e3-4d6e-956d-77a39d7f0d06 · outbound

This paper cites AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation.

GenHSI: Controllable Generation of Human-Scene Interaction Videos AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:08.931807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:3921f3fb67c4d712781b368758b351e1370f63c61038524e2d5c4e2852af2561

Observation 88d53b2d-cad6-438f-a4b7-61efa9a77901 · outbound

This paper cites Person in place: Generating associa- tive skeleton-guidance maps for human-object interaction image editing.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Person in place: Generating associa- tive skeleton-guidance maps for human-object interaction image editing

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.410778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:aa779994e34ec54bbd9903ab53febd40d421ea204c2521d0fded2231b216b773

Observation fb632a70-f0a9-4d66-b20e-c4f0989dbe53 · outbound

This paper cites Generating human interaction motions in scenes with text control.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Generating human interaction motions in scenes with text control

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.418152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:1c9458bcc0c7f3bbb17eb9be7ce6f3815fd1fd8ac768c6f187273bf7aaf82c8a

Observation b675014c-3183-4d0e-9b43-feb2333148a4 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Adding conditional control to text-to-image diffusion models

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.413583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:13f945336495c27d8a38cdabc7901241c3848a5f7ea71d3d7545d876bb478115

Observation 04faad48-d1f5-4bff-8960-98b4b22362f1 · outbound

This paper cites Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Training-Free Motion-Guided Video Generation with Enhanced Temporal Consistency Using Motion Consistency Loss

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:27:09.055448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:20880b39ef1563a5200bb51328b98490cc9c265ab0915e21dbd5a9d415d7a1a9

Observation 171c67e9-8747-4036-b9ea-b0410b11eb38 · outbound

This paper cites Generating 3d people in scenes with- out people.

GenHSI: Controllable Generation of Human-Scene Interaction Videos Generating 3d people in scenes with- out people

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T07:27:09.405656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T07:25:41.219753Z digest=sha256:e2261cfd58c7235060df786dad48cf3a30f45358a68faafa15bb830e6bd1c399

Pith citing papers

Observation 4a07035f-929a-442a-9a64-d6dbb5889d01 · inbound

VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification cites this paper.

VHOI: Controllable Video Generation of Human-Object Interactions from Sparse Trajectories via Motion Densification GenHSI: Controllable Generation of Human-Scene Interaction Videos

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-16T23:31:22.002931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T23:30:14.969895Z digest=sha256:d99a579f2679c1b8c4c907ffa99396647e9155f7457d1f9f9c1536369aa836ee

Observation fce67bb8-d112-45d1-9e60-fdf9e3d2b5ba · inbound

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos cites this paper.

CoMoVi: Co-Generation of 3D Human Motions and Realistic Videos GenHSI: Controllable Generation of Human-Scene Interaction Videos

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-16T13:47:57.383574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T13:43:26.460480Z digest=sha256:9e785c7d39e707749a3f480f6da247605b7bb45814dc29fc68d2e0aa11086191