Pith. sign in

Paper Citation Record · LEDGER

Populate-A-Scene: Affordance-Aware Human Video Generation

As of 8 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2507.00334.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00334 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:23:12.762375Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d2c00502-cd33-45ef-9bfe-2ffe5bd25e96 · outbound

This paper cites Fouhey, Ivan Laptev, Josef Sivic, Abhinav Gupta, and Alexei A.

Populate-A-Scene: Affordance-Aware Human Video Generation Fouhey, Ivan Laptev, Josef Sivic, Abhinav Gupta, and Alexei A

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:13.063341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:23:12.687060Z digest=sha256:5351f5be7ebdd1c301b072362d10ba635e4c1aa7152ddf3e4f59d94e7892653d

Observation 49c61981-868b-4297-82cc-18e96dff7a58 · outbound

This paper cites Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models.

Populate-A-Scene: Affordance-Aware Human Video Generation Preserve Your Own Correlation: A Noise Prior for Video Diffusion Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.697903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.697903Z digest=sha256:4902cc6ebb95d0d7bba3598dd0649bb9cedefb140ea6edd5db92914295afd843

Observation c25ede42-3767-4ac1-8629-05fa90181633 · outbound

This paper cites Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance.

Populate-A-Scene: Affordance-Aware Human Video Generation Animate Anyone 2: High-Fidelity Character Image Animation with Environment Affordance

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.708601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.708601Z digest=sha256:ebf3728025a2c6941c5b76c8b4ea2f73d5d92eec6f40a59e2a90f7a1af340232

Observation f64611c1-989b-4065-a6af-33dbb54fbe1c · outbound

This paper cites DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models.

Populate-A-Scene: Affordance-Aware Human Video Generation DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.712364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.712364Z digest=sha256:27fd06a1a52d07832c7393f3041b969716ac7721e0171380dc4daaaf66b3c5c3

Observation 62a50608-7b31-4e50-a1fd-98ee4c5747bd · outbound

This paper cites Flow Matching for Generative Modeling.

Populate-A-Scene: Affordance-Aware Human Video Generation Flow Matching for Generative Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.721211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.721211Z digest=sha256:f2e5b6231409bd204efa3efbb8a4e4e7b366293725c3bdfcb0a3d37a07df93f4

Observation 257b1446-f133-4fbe-9d3a-a6d3b4b6b768 · outbound

This paper cites Synthesizing Environment-Specific People in Photographs.

Populate-A-Scene: Affordance-Aware Human Video Generation Synthesizing Environment-Specific People in Photographs

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T21:23:12.941665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:23:12.728400Z digest=sha256:378b111d1749c7c3fb4adf13d7277e6ec44375a6a2293916f9d500a3e4585a7f

Observation 93484021-023b-4bf8-b5db-e9eec2f6ad75 · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

Populate-A-Scene: Affordance-Aware Human Video Generation Movie Gen: A Cast of Media Foundation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.731808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.731808Z digest=sha256:615f0503adbd63fe278684c5a2ec178c4fe2a82447b86947492737919f0920ae

Observation a2aed091-cead-4adc-91d6-3e620802ca67 · outbound

This paper cites ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation.

Populate-A-Scene: Affordance-Aware Human Video Generation ConsistI2V: Enhancing Visual Consistency for Image-to-Video Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.735107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.735107Z digest=sha256:4779e0878aab964443e9637260b18acfe7c9132dcd0a9690e18d65cb8651a8c1

Observation 94434dd0-0ad6-458c-94ec-7aad763aed84 · outbound

This paper cites InVi: Object Insertion In Videos Using Off-the-Shelf Diffusion Models.

Populate-A-Scene: Affordance-Aware Human Video Generation InVi: Object Insertion In Videos Using Off-the-Shelf Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.738455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.738455Z digest=sha256:fcc47261be8a4d778ab122763a06e5d20f1eb826f0c34e240afe33987cbd78ba

Observation 75d06842-c7b2-4987-8b31-20e8e90525cf · outbound

This paper cites Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman.

Populate-A-Scene: Affordance-Aware Human Video Generation Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.741893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.741893Z digest=sha256:a8fe1318d372d3ed78a766e1643c340129045b3217d94ba4b951a5eb0ca6a15f

Observation 4d6b4ed2-1c92-4488-b1ec-a1dd95fee2c0 · outbound

This paper cites UL2: Unifying Language Learning Paradigms.

Populate-A-Scene: Affordance-Aware Human Video Generation UL2: Unifying Language Learning Paradigms

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.745181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.745181Z digest=sha256:6540a51e4a7a838b26d84bdd6d957b1b33cfa12c27c22b4f88f91e07b429ad4e

Observation 0ab42c9e-6dae-4de6-bf30-21f072092da2 · outbound

This paper cites DAT++: Spatially Dynamic Vision Transformer with Deformable Attention.

Populate-A-Scene: Affordance-Aware Human Video Generation DAT++: Spatially Dynamic Vision Transformer with Deformable Attention

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.752065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.752065Z digest=sha256:4ba643a8531743b44b2d47a7ad9a919467a99a486966bb5d0869fedbef038122

Observation 9cea3e29-550c-4d6b-8a8e-b62f6ee3c1b6 · outbound

This paper cites AMG: Avatar Motion Guided Video Generation.

Populate-A-Scene: Affordance-Aware Human Video Generation AMG: Avatar Motion Guided Video Generation

Reference 24

Resolution
malformed identifier
local_arxiv, observed 2026-08-06T21:23:12.814830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:23:12.756044Z digest=sha256:b99f10b68b1388a812b714bffee8e4309b8c9362eafb6e8f4f9519dab8b7e618

Observation afc3823d-ff8f-4963-a9ce-6039fa11997e · outbound

This paper cites Make Pixels Dance: High-Dynamic Video Generation.

Populate-A-Scene: Affordance-Aware Human Video Generation Make Pixels Dance: High-Dynamic Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.759230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.759230Z digest=sha256:ec3a3cd971732398ea4e15c898873ae3db241440d622f2adfbb815cbfff402ba

Observation ceb6a3c0-c822-45b2-a3fc-e5bdec43865e · outbound

This paper cites Shenhao Zhu, Junming Leo Chen, Zuozhuo Dai, Yinghui Xu, Xun Cao, Yao Yao, Hao Zhu, and Siyu Zhu.

Populate-A-Scene: Affordance-Aware Human Video Generation Shenhao Zhu, Junming Leo Chen, Zuozhuo Dai, Yinghui Xu, Xun Cao, Yao Yao, Hao Zhu, and Siyu Zhu

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:23:13.054043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:23:12.762375Z digest=sha256:e1c805c1ceaf197f813e6d7988c12cd85606b6fda5d0514ede6cda87e0d33cca

Observation e35b201a-a990-4add-a86a-7cc744d679ba · outbound

This paper cites AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks.

Populate-A-Scene: Affordance-Aware Human Video Generation AnyV2V: A Tuning-Free Framework For Any Video-to-Video Editing Tasks

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.716870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.716870Z digest=sha256:62514768b234c472422812c663ad49d265f4a7d70d8dda6bef79a031a71b8bfd

Observation 1909a206-97fe-472a-ac8f-167bc2ca06ff · outbound

This paper cites Photorealistic Video Generation with Diffusion Models.

Populate-A-Scene: Affordance-Aware Human Video Generation Photorealistic Video Generation with Diffusion Models

Reference 2011

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.701371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.701371Z digest=sha256:49c1367055572036af1df29159853c3f2b4d66fe96e11d2c7f777f876f722b2a

Observation 620ed033-fd70-4614-88da-2ba92398777f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Populate-A-Scene: Affordance-Aware Human Video Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.690412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.690412Z digest=sha256:f91671c05b06c86db05bf0064a2638c6884fb5dd6e30db911f81f4bef4872b28

Observation b7dd9eae-fa44-4463-9bda-7f23d84e4b93 · outbound

This paper cites ModelScope Text-to-Video Technical Report.

Populate-A-Scene: Affordance-Aware Human Video Generation ModelScope Text-to-Video Technical Report

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.748660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.748660Z digest=sha256:dd7fbc7cf84ecd284e25dfeadeb3f15e8254118a0744afeb60ec1b7ae231e674

Observation 74a42b92-080c-404c-a25d-ed8b1516d290 · outbound

This paper cites Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack.

Populate-A-Scene: Affordance-Aware Human Video Generation Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.682484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.682484Z digest=sha256:062092372efdd38e9f8611df4a28295a77a39b59263226b26e5c89cc5ccdece6

Observation 980abd31-064a-4b46-bc7d-3dfcd4079937 · outbound

This paper cites The Llama 3 Herd of Models.

Populate-A-Scene: Affordance-Aware Human Video Generation The Llama 3 Herd of Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.693914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.693914Z digest=sha256:c22cf282b2c876a9b2fc0d985e68abff83b426aa2abc4d6a323a497cbf8b38dd

Observation 8d11e3cc-a032-4197-bf32-4e8ae0c94b9a · outbound

This paper cites Latte: Latent Diffusion Transformer for Video Generation.

Populate-A-Scene: Affordance-Aware Human Video Generation Latte: Latent Diffusion Transformer for Video Generation

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.725035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.725035Z digest=sha256:27f48819738e8a8352e15e47158090ae4e775f0a03de229e3f39dbc250498a48

Observation 3dbc78da-5f3a-41b8-b3be-4d77d77789c1 · outbound

This paper cites Imagen Video: High Definition Video Generation with Diffusion Models.

Populate-A-Scene: Affordance-Aware Human Video Generation Imagen Video: High Definition Video Generation with Diffusion Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.705127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.705127Z digest=sha256:5126415d7e26c38b24b096ef0949a1815c8d0f26197b2f4f11f434f98362b715

Observation ffc8f9e7-cb9e-492a-af0e-0bccb47cd800 · outbound

This paper cites Flow map matching with stochastic interpolants: A mathematical framework for consistency models.

Populate-A-Scene: Affordance-Aware Human Video Generation Flow map matching with stochastic interpolants: A mathematical framework for consistency models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:23:12.674110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:23:12.674110Z digest=sha256:49b9910b6b6b2e3e6207c82dd289132c72c15118f8027e5648b41568798da67c

Observation 38463b39-f0da-4dd0-bc6e-d479b6d341fd · outbound

This paper cites an unresolved cited work.

Populate-A-Scene: Affordance-Aware Human Video Generation Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:23:13.072739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:23:12.669420Z digest=sha256:c62df8b9e3d707861512e5b8cb15c7938cab64b3d3a154630c8c9ed67629fd2d

Observation c9db72b1-ff8b-4e03-9fa0-216c54c3cc74 · outbound

This paper cites doi: 10.1145/3715140.https://doi.org/10.1145/3715140.

Populate-A-Scene: Affordance-Aware Human Video Generation doi: 10.1145/3715140.https://doi.org/10.1145/3715140

Reference 2025

Resolution
verified exact
doi, observed 2026-08-06T21:23:12.791955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:23:12.678507Z digest=sha256:b435f67312aac6a56cc5118e295d15eeaa13757dbf49ed8684e2f6532a6ac7c7

Pith citing papers

No inbound Pith citation observations are available.