Pith. sign in

Paper Citation Record · LEDGER

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos

As of 20 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2504.19165.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.19165 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T06:04:43.030542Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9e333019-eb5d-42ca-8890-1a0a29c583e3 · outbound

This paper cites Ren- derdiffusion: Image diffusion for 3d reconstruction, inpaint- ing and generation.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Ren- derdiffusion: Image diffusion for 3d reconstruction, inpaint- ing and generation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:44.122100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.689246Z digest=sha256:5dd4f7121e36694388a99d0c3b37b97bbd6e7fe7b5351f5f18e026d2e0ee1e6e

Observation b94f6888-e32b-4b65-93b3-450bc55bec79 · outbound

This paper cites Rignerf: Fully controllable neu- ral 3d portraits.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Rignerf: Fully controllable neu- ral 3d portraits

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:44.095579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.699281Z digest=sha256:7eee7dfa99f9aa172a0f7c2fdededd10d90c5f9f0d9f78f1e5993383c0f02306

Observation 9f4b36f2-5ab5-45f1-9690-a29b1ab5b300 · outbound

This paper cites Learning personal- ized high quality volumetric head avatars from monocular rgb videos.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Learning personal- ized high quality volumetric head avatars from monocular rgb videos

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:44.074919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.705663Z digest=sha256:7e2ebb79f11224ff310d403314732e032a8c6a87f00b18e6e90e9de2799665ee

Observation 378239ad-3201-4bcb-9938-53408d3a4e01 · outbound

This paper cites Effi- cient 3d implicit head avatar with mesh-anchored hash table blendshapes.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Effi- cient 3d implicit head avatar with mesh-anchored hash table blendshapes

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:44.053506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.712611Z digest=sha256:884697f95801209c993d653c869f3433cf574887af07714005662f68c3d90193

Observation f024a557-b17d-4947-8471-1ccb28dde4d8 · outbound

This paper cites Lumiere: A Space-Time Diffusion Model for Video Generation.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Lumiere: A Space-Time Diffusion Model for Video Generation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.719894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.719894Z digest=sha256:0170c1088fe812f4675e60267c682092f7ec7266f3d897df2912d20d6afa2c8c

Observation 01eb0f3e-d7a3-4cd6-ae6a-7d42ce818c39 · outbound

This paper cites A morphable model for the synthesis of 3d faces.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos A morphable model for the synthesis of 3d faces

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:44.027180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.729900Z digest=sha256:8e30716705e393e810b7833c289663d28efe6e35f9c91942e92412b243555e0e

Observation 5eba6ddc-2940-4e48-a583-cf6855fd68d2 · outbound

This paper cites Hyperreenact: One-shot reenactment via jointly learning to refine and re- target faces.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Hyperreenact: One-shot reenactment via jointly learning to refine and re- target faces

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.998122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.739204Z digest=sha256:41e2e243991c3d16ab64682aef6baa2233ff0c43c24dd48ca3ccd62afb930af6

Observation 97c1bdfd-af73-4ade-a76a-de5b23a2abf9 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos In- structpix2pix: Learning to follow image editing instructions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.977011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.745807Z digest=sha256:572eb417c2be613440335663057e85bc14e4e60c5a2f3ffea5257bd4d83fd073

Observation e72ed37a-45cf-4eb8-995a-fc5287ec76b9 · outbound

This paper cites Magicpose: Realistic human poses and facial expressions retargeting with identity-aware diffusion.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Magicpose: Realistic human poses and facial expressions retargeting with identity-aware diffusion

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.948332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.753875Z digest=sha256:ed75a607f1443c74d0901d307a6c1d1e4310a78542ba4ff118c26edb32a4a05e

Observation 48ed9474-459c-4938-ae06-907bc907b23f · outbound

This paper cites Personalized face modeling for improved face reconstruction and motion retargeting.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Personalized face modeling for improved face reconstruction and motion retargeting

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.922932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.758306Z digest=sha256:3663685418dd7287bb4c1143fb46f2c425f858b75401e8d4bd40f5d2e9559da8

Observation 895e8b83-1c67-4f1b-9941-8863cd1cdb9e · outbound

This paper cites Monogaus- sianavatar: Monocular gaussian point-based head avatar.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Monogaus- sianavatar: Monocular gaussian point-based head avatar

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.763564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.763564Z digest=sha256:6e25b8e811cb2bcfa149b99ddff3a896a99bf398b4911eb7007d914c30520b12

Observation ba3a9968-905c-4f72-a998-8d863298458c · outbound

This paper cites Portrait4D-v2: Pseudo Multi-View Data Creates Better 4D Head Synthesizer.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Portrait4D-v2: Pseudo Multi-View Data Creates Better 4D Head Synthesizer

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.770897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.770897Z digest=sha256:f9ca4dec86698940961d137b2bb8eed605bac270f648550d9ed26ad091befad1

Observation 640397b4-7ffa-4da1-ab8f-2b8de54be33f · outbound

This paper cites Portrait4d-v2: Pseudo multi-view data creates better 4d head synthesizer.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Portrait4d-v2: Pseudo multi-view data creates better 4d head synthesizer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.887468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.782178Z digest=sha256:9b00e58f959156d6f0895a7b5ed1416bff78f34d24f4de9a69032622212e738d

Observation 01bd7109-b9fd-4153-97ed-c81b9234a12c · outbound

This paper cites Diffusionrig: Learning personalized priors for facial appearance editing.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Diffusionrig: Learning personalized priors for facial appearance editing

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.864549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.787249Z digest=sha256:78666a2048835982b5dd322d22d9647299d6e19abae7f8bff2180e32002c8927

Observation fda516f8-b0b8-456b-9825-c49319601a2b · outbound

This paper cites Headgan: One-shot neural head synthesis and editing.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Headgan: One-shot neural head synthesis and editing

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.792301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.792301Z digest=sha256:276a1de79b03fd20754ba971e5dfbc72b2986f7a4c2da85c9162ff41551887dd

Observation 11f32ccb-80c6-4d9e-926d-b8a42fe7f759 · outbound

This paper cites Megaportraits: One-shot megapixel neural head avatars.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Megaportraits: One-shot megapixel neural head avatars

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.797059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.797059Z digest=sha256:6e8eaf1f3c352abf3709e695a26ee4d63b5c95722ceec6c6389fee14076bf12f

Observation abaf13a3-df2a-4c40-be28-1d1866c73ef4 · outbound

This paper cites Emoportraits: Emotion-enhanced multimodal one-shot head avatars.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Emoportraits: Emotion-enhanced multimodal one-shot head avatars

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.801463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.804358Z digest=sha256:ff8c71697721021d54805b7a33d08905d1a8380e426776c50dd586487919bebb

Observation d372b808-2574-46b1-987d-c68d507c1891 · outbound

This paper cites From data to functa: Your data point is a function and you can treat it like one.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos From data to functa: Your data point is a function and you can treat it like one

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.808853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.808853Z digest=sha256:ba17e203765f51c1c3b20f857f18804b283de77af7586b3cd2aebb43d430fb02

Observation f23719a9-bfef-4050-8b9b-c4ce35a2948f · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.815231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.815231Z digest=sha256:a4b58e2d3436f95a66508eb14e2b032e576b4ea113fecb39b17a9d5e7663b09e

Observation f0ef7c4c-82b8-4796-9e41-ac83ceee92a1 · outbound

This paper cites Dynamic neural radiance fields for monocular 4d facial avatar reconstruction.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Dynamic neural radiance fields for monocular 4d facial avatar reconstruction

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.763482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.823293Z digest=sha256:87f417e558dc9963a691d97c64de2b1248596ecb51188de1869b3881b6e3ff67

Observation d34828a1-671b-4dd2-90ab-8d8fd1ea7b41 · outbound

This paper cites Reconstructing personalized se- mantic facial nerf models from monocular video.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Reconstructing personalized se- mantic facial nerf models from monocular video

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.744379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.830093Z digest=sha256:a8fc607a9b29bdb96ac4c717ead090ae6ad46723c2930cce675bde3c6487cd64

Observation 61f4a6b8-c4da-4d6a-b175-4696513ac966 · outbound

This paper cites Neural head avatars from monocular rgb videos.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Neural head avatars from monocular rgb videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.834353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.834353Z digest=sha256:91dc3ebbdcda59856b8b174d36e7802cf9b05d19c22e305afa1c6db6ac5544cf

Observation 0c7d28ad-919f-4f73-b1b6-7067a7b63ce6 · outbound

This paper cites LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.839847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.839847Z digest=sha256:2df8d2c19b8234d541d1dbf021afda5bc23fd08a2abe43725514be5bd8445697

Observation 8f732032-42c1-4b34-826f-0295548be098 · outbound

This paper cites Gans trained by a 9 two time-scale update rule converge to a local nash equilib- rium.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Gans trained by a 9 two time-scale update rule converge to a local nash equilib- rium

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.702421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.848945Z digest=sha256:2eadfeef3f9f469ebcc992406d9c761ab0f1532ef10b18c736d07061ce9e8cff

Observation dd33650e-159a-4490-abff-a6bdc108ffa2 · outbound

This paper cites Denoising dif- fusion probabilistic models.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Denoising dif- fusion probabilistic models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.853830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.853830Z digest=sha256:403ebe6eed654fd6fb7a119bd32791235d791ab5d313f806f730397d80d665ae

Observation ac1e49a8-87d8-44d2-87c0-2096b4e9dace · outbound

This paper cites Shap-E: Generating Conditional 3D Implicit Functions.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Shap-E: Generating Conditional 3D Implicit Functions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.860950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.860950Z digest=sha256:01298fcd91c8269b20bf8f8755bd3da563a93164f7fbbf15915af980aff120fb

Observation 3c4c3c10-414e-47eb-91fe-c613a3fd1dd9 · outbound

This paper cites Ln3diff: Scalable latent neural fields diffusion for speedy 3d generation.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Ln3diff: Scalable latent neural fields diffusion for speedy 3d generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.662889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.867234Z digest=sha256:e60f80e17817da55254c91b371c5dbdbe9add5ec791175030d554061c7291e7c

Observation 432661a3-44c1-4231-b5f7-d3c5ddf3eba4 · outbound

This paper cites Learning a model of facial shape and expression from 4d scans.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Learning a model of facial shape and expression from 4d scans

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.643270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.873514Z digest=sha256:7399fd438044800993e9cebccfa9b93f78a3d1e4147e575d2bbc7859701dd8f1

Observation 77826726-8948-4f95-8553-75f8e1546fd8 · outbound

This paper cites Raft-stereo: Multilevel recurrent field transforms for stereo matching.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Raft-stereo: Multilevel recurrent field transforms for stereo matching

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.618994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.880282Z digest=sha256:57625c0e6fd5baf44998876ab3352fb36afb4b26e4d4dc091d878e9d09093a4e

Observation 4ee9bd4f-a068-4061-890e-5c481df67c80 · outbound

This paper cites Diffdub: Person-generic visual dubbing using inpaint- ing renderer with diffusion auto-encoder.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Diffdub: Person-generic visual dubbing using inpaint- ing renderer with diffusion auto-encoder

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.888933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.888933Z digest=sha256:f544a6bbf27be94763f01ae92eed8ae0c715a6d38abe7b9695ab6abf1d0e96e6

Observation 71f1f70a-8ad2-4c60-b363-0cc1a1b8f390 · outbound

This paper cites Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.894821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.894821Z digest=sha256:df76c891918f3c27374f038e4c9ce5344bf2768fddd4d551152f36b3850f4e9f

Observation 00131728-8648-4758-a9c8-32939d87bcbc · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view syn- thesis.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Nerf: Representing scenes as neural radiance fields for view syn- thesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.899570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.899570Z digest=sha256:37fdd167610293bda9be67ec428a384e23d7f8fd94a34f1a708defca27a17845

Observation 48d35746-f573-4d2e-9b30-d498affff066 · outbound

This paper cites Diffrf: Rendering-guided 3d radiance field diffusion.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Diffrf: Rendering-guided 3d radiance field diffusion

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.904536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.904536Z digest=sha256:1f88c65a281bc2553c801ab9842214cf39bec92d4ef8066ff93a1a2053f7ee46

Observation fb07274f-9a60-4674-83b7-3438574cec79 · outbound

This paper cites Gaus- sianavatars: Photorealistic head avatars with rigged 3d gaus- sians.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Gaus- sianavatars: Photorealistic head avatars with rigged 3d gaus- sians

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.910487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.910487Z digest=sha256:ea693ad8d9758778ceb8e1b01fc31c9d2bfe2ee3843a1eafdb1c92bf519564aa

Observation d27f5b89-d8e8-47d0-bec8-6edaa42a9550 · outbound

This paper cites Relightable gaussian codec avatars.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Relightable gaussian codec avatars

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.922075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.922075Z digest=sha256:9ac7e36448d1377e8f7f17d805fc2b0f4afba230738f080cd7a09732689070b0

Observation 5ab4cd2f-7a74-44ff-bcc8-8a2199f171c3 · outbound

This paper cites First order motion model for image animation.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos First order motion model for image animation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.929358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.929358Z digest=sha256:e601e82af984031a4f438bd94f9b4a973d01a8c773857f2a1cf27f79176c2c89

Observation b06ce8dc-4c39-4c49-afaa-76d685eec1b4 · outbound

This paper cites Diffusion with forward models: Solv- ing stochastic inverse problems without direct supervision.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Diffusion with forward models: Solv- ing stochastic inverse problems without direct supervision

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.494852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.934970Z digest=sha256:778e1c9dc9c86073714dbffd3f0ab2ab74d823931cc4d48b3aeded5f8848babb

Observation 7a03ed31-f658-4d2b-a3e0-b7e790661664 · outbound

This paper cites Chan, Chao Liu, Zhiding Yu, Sameh Khamis, Manmohan Chandraker, Ravi Ramamoorthi, and Koki Nagano.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Chan, Chao Liu, Zhiding Yu, Sameh Khamis, Manmohan Chandraker, Ravi Ramamoorthi, and Koki Nagano

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.474005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.941945Z digest=sha256:879b2722540397c626f2add7fc2b16e239a637e709aba4aae288876b990b5324

Observation 4c7d81ea-6620-4153-8df6-cbe30b6b1007 · outbound

This paper cites Single-view view synthe- sis with multiplane images.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Single-view view synthe- sis with multiplane images

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.449010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.947702Z digest=sha256:12d531748ec266519e883c38486aacbe2211ae0336cb1b8e7168f5541288f305

Observation 900e13e8-31ec-43c3-a5ab-d05d2ae1a11a · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.952677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.952677Z digest=sha256:e94af55dd11f993ea6daa9b20fed66ed4ca6fad222a5a6476ab7936f018477a3

Observation 36235e18-5975-4621-9fc0-56d40abafb2a · outbound

This paper cites Lion: Latent point dif- fusion models for 3d shape generation.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Lion: Latent point dif- fusion models for 3d shape generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.958332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.958332Z digest=sha256:ab1c0765ea9012bb1588eb346c4f6d4791d53542e78589c5d1345c6a78f73ee0

Observation aeaaa91b-c604-488f-9f48-7e07fa705751 · outbound

This paper cites Rodin: A generative model for sculpting 3d digital avatars using diffusion.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Rodin: A generative model for sculpting 3d digital avatars using diffusion

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.964174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.964174Z digest=sha256:5bb29f1e108b6d33622e8f9eb9e26e2988bb4982cb086e5ba40e6aeb0860e1fe

Observation bd29e840-3d99-4e15-ad7e-70f1a3bdddd1 · outbound

This paper cites One-shot free-view neural talking-head synthesis for video conferenc- ing.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos One-shot free-view neural talking-head synthesis for video conferenc- ing

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.969788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.969788Z digest=sha256:cafbdbfab52caa94fb8f54cac6a757b44e26079f4ace6ab2cfca87799fb77577

Observation 0b5ea385-f09f-4d9c-89c4-35069fa852b7 · outbound

This paper cites Mvdd: Multi-view depth diffusion mod- els.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Mvdd: Multi-view depth diffusion mod- els

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.380066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.975073Z digest=sha256:a9a28cb20c8e1edae4c6e717ec251141ff41627ed35212f081fc31e8b9e578f7

Observation 5dd5e7d6-bf68-4966-91c9-a57862a1bc9b · outbound

This paper cites Vfhq: A high-quality dataset and bench- mark for video face super-resolution.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Vfhq: A high-quality dataset and bench- mark for video face super-resolution

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.361399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.979146Z digest=sha256:90c6a6a55099ec2f3d28f5968a9b3ab8808a20da2ecadb940b45df9f5173b6c7

Observation e8946256-01df-410d-8ac7-8e3d2f6af16a · outbound

This paper cites X-portrait: Expressive portrait anima- tion with hierarchical motion attention.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos X-portrait: Expressive portrait anima- tion with hierarchical motion attention

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.343837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.986581Z digest=sha256:4c7af3c8109818ac3c3d3f75643d5f4b2a15c204db6a3b155f68d88de6d3cd42

Observation 62f99e85-72b4-44c7-ab1d-439fb9459931 · outbound

This paper cites Depth Anything V2.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Depth Anything V2

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.990961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.990961Z digest=sha256:29f97aa121e550582e95389e2fedd5ef7602ee86766127ccae45aee8ab7ba438

Observation 126d6c94-fb95-4c97-8da8-a1eb2d1e686b · outbound

This paper cites Styleheat: One-shot high-resolution ed- itable talking face generation via pre-trained stylegan.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Styleheat: One-shot high-resolution ed- itable talking face generation via pre-trained stylegan

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.327775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:42.995717Z digest=sha256:46ea7629a7debbc43a0aa44995561d772498e5dabb1f19fb92c80415912e77cc

Observation fb2d03ea-0291-44fb-8ffb-d30ea944ce45 · outbound

This paper cites Fast bi-layer neural synthesis of one- shot realistic head avatars.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Fast bi-layer neural synthesis of one- shot realistic head avatars

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:42.999690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:42.999690Z digest=sha256:9613bf04bad30a8a76aa0aff190db58ef218474a093e0cca7bce92fa6cd429c2

Observation a1e65c85-ea80-49c9-a6d0-634f03333952 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Adding conditional control to text-to-image diffusion models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:43.006003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:43.006003Z digest=sha256:31fc85f379587785ec2c81072fb4dbac1e22a2dc8f1000a7b59aaf0bddbd7817

Observation 95ad8f95-f661-49b9-81f5-b902c32c4d98 · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos The unreasonable effectiveness of deep features as a perceptual metric

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:43.010461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:43.010461Z digest=sha256:a2cfca762334741c1e7aedcb97c1d09bdf278f39eea52cb24eeb30f04289541a

Observation f706462d-5f2b-4fa3-9fe2-2bea7e568185 · outbound

This paper cites Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T06:04:43.014480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T06:04:43.014480Z digest=sha256:c35da40c8fe663c07d4279940c87595c17aab4d576a03c81f91646cfa6205025

Observation 8693f762-f413-4ecb-9694-1ecb3ac09957 · outbound

This paper cites Generative multi- plane images: Making a 2d gan 3d-aware.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos Generative multi- plane images: Making a 2d gan 3d-aware

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.266884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:43.019943Z digest=sha256:cf6718059fcf4ae50d5213f6d9b24fc2b2d6226448b8d13a8ed0359d4d029838

Observation 82d0e6ce-f84a-4d63-a882-673dd3f98203 · outbound

This paper cites 2D diffu- sion + depth.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos 2D diffu- sion + depth

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.248955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:43.024683Z digest=sha256:550135f3d1f0f6ff55a19376c537ed57884ae3856a237f7cae0b2d5992ec9d43

Observation a42c6fbb-94f4-4a55-8a74-a8da4b964f9d · outbound

This paper cites In our experiments, all the side view render- ings, stereo renderings and rendering speed measurements are conducted through the first method.

IM-Portrait: Learning 3D-aware Video Diffusion for Photorealistic Talking Heads from Monocular Videos In our experiments, all the side view render- ings, stereo renderings and rendering speed measurements are conducted through the first method

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T06:04:43.231620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T06:04:43.030542Z digest=sha256:be45217fbe7209f83d68d22054a3f1f4a3f34804159a64351f4dd6105376a4bf

Pith citing papers

No inbound Pith citation observations are available.