Pith. sign in

Paper Citation Record · LEDGER

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

As of 9 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2512.13247.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.13247 v2

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T16:30:41.029982Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved87
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 242fbb76-35db-4825-a0dc-9859150589a7 · outbound

This paper cites Deep audio-visual speech recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(12):8717–8727, 2018.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Deep audio-visual speech recognition.IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(12):8717–8727, 2018

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:31.570218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:31.570218Z digest=sha256:cc3e35cdf4717124f68da0b50516eb7d377850ec41e83d1cace0308a63c7c0ac

Observation 6928ce21-def6-4d5e-8639-962e91a697bd · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits LRS3-TED: a large-scale dataset for visual speech recognition

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:31.687428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:31.687428Z digest=sha256:7f88340e5dbaa6626de0d1646111345aac8a402ba8af0ad3c499a678b7ca5424

Observation 191b26e3-a26e-4af3-ba90-f0aa1d6265b1 · outbound

This paper cites Rignerf: Fully controllable neu- ral 3d portraits.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Rignerf: Fully controllable neu- ral 3d portraits

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:31.788229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:31.788229Z digest=sha256:053b3c7aab8ca01d97bbbba33d01697d8a9a31444cd1c5b82ec12c981e4e9d06

Observation 46ffd0bc-69db-4bf5-b3c8-5fa83c3a83cc · outbound

This paper cites Id-to-3d: Expressive id-guided 3d heads via score distillation sampling.Advances in Neural Informa- tion Processing Systems, 2024.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Id-to-3d: Expressive id-guided 3d heads via score distillation sampling.Advances in Neural Informa- tion Processing Systems, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:31.865125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:31.865125Z digest=sha256:9f840a1cb46b2b11d1d77bd9ae360df3e6f616b4e26c6995823accb66a5bbeff

Observation 39edeeb2-bef9-4b81-8b09-7c4309d37a5b · outbound

This paper cites Wav2vec 2.0: A framework for self- supervised learning of speech representations.Advances in neural information processing systems, 2020.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Wav2vec 2.0: A framework for self- supervised learning of speech representations.Advances in neural information processing systems, 2020

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:31.955776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:31.955776Z digest=sha256:279ad2599908c4e3994d6078b9044ce2b2d8321c50bed8d789f2e0023cca3b29

Observation 752870c4-ffaa-47fb-a774-6fa0fa686ce2 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:32.166173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:32.166173Z digest=sha256:df977731c8681fe28b064ed2848a15395563e7596952adedc7c5a467e11e139c

Observation 1041f84d-e043-4222-978d-906f7abfede2 · outbound

This paper cites Align your latents: High-resolution video synthesis with la- tent diffusion models.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Align your latents: High-resolution video synthesis with la- tent diffusion models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:32.343965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:32.343965Z digest=sha256:551bf14be40be3751539765f2d3f1747a3f35cc09cd95265bb84236fc0215a8a

Observation 11e0462b-0318-4663-b54e-ef55b3a2de65 · outbound

This paper cites Masactrl: Tuning-free mu- tual self-attention control for consistent image synthesis and editing.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Masactrl: Tuning-free mu- tual self-attention control for consistent image synthesis and editing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:32.434438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:32.434438Z digest=sha256:e3e87201cfb82751d48667d6836a4e0d8e66930f0c46fdd2068a0ac2158ab935

Observation 65dafdc2-20ea-4a98-a534-ca4d4625b2a1 · outbound

This paper cites X-dyna: Ex- pressive dynamic human image animation.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits X-dyna: Ex- pressive dynamic human image animation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:32.553306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:32.553306Z digest=sha256:ebcc6e34985c2aae2b68a15ffff44bd9e6e45dc1d09108911b26c209e22e882b

Observation 3e328c9b-3b15-426b-bcdd-98d5185051b3 · outbound

This paper cites Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:32.624883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:32.624883Z digest=sha256:a0ca77ee2cad518c875a58c63b9a627eace31f4d084e1e18a2f7e2ba09b2497d

Observation 0b127765-b937-43a6-9ef4-c84c1a3a5510 · outbound

This paper cites Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:32.718987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:32.718987Z digest=sha256:4002253fd0fe5d4d7931c10a84e8970376dc5aef36d977ddb2e4caeddb92e2dd

Observation 4235c6d0-9187-438b-837a-31dde6e7e43d · outbound

This paper cites Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:32.803336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:32.803336Z digest=sha256:d1f54b300c5ca2d24fe1740782c184657a05d1a59e9be5e44ee4ec2b3b10d661

Observation 4636df29-4c2d-403f-8874-67529b654656 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Arcface: Additive angular margin loss for deep face recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:32.886547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:32.886547Z digest=sha256:ac907b510854406e2654f6c202dc15fdfc77b2cd2df5c252a1bd1a7621749be1

Observation ddeb31f7-6397-419e-8b5c-6bbbea3d046f · outbound

This paper cites Portrait4d-v2: Pseudo multi-view data creates better 4d head synthesizer.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Portrait4d-v2: Pseudo multi-view data creates better 4d head synthesizer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:33.061276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:33.061276Z digest=sha256:9aa6a6b355b8eb338a67ec225afcd8c65a1a55b6623876f93e27836eab87de13

Observation 8f4a3a6d-0f4e-4226-9a44-4141583dce20 · outbound

This paper cites Black, and Timo Bolkart.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Black, and Timo Bolkart

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:33.162001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:33.162001Z digest=sha256:a425acb38e980402ab201b4999d610408063f0db0c43d994f155231333c95a51

Observation fbb8817e-46d6-43cb-b651-a0698d69fe13 · outbound

This paper cites Visual Speech-Aware Perceptual 3D Facial Expression Reconstruction from Videos.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Visual Speech-Aware Perceptual 3D Facial Expression Reconstruction from Videos

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:33.250107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:33.250107Z digest=sha256:b730fb05ee10077cd74705312a537ead57930accc57b0ae7a2b820f4a393e218

Observation 826a90c1-2e46-49be-99d4-72bbfc6f7d5f · outbound

This paper cites Dynamic neural radiance fields for monocular 4d facial avatar reconstruction.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Dynamic neural radiance fields for monocular 4d facial avatar reconstruction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:33.301402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:33.301402Z digest=sha256:6284c30aefcd9326534e04f39f99d8a308ac56053a9a1149b11d051f2681ba0b

Observation 02398a05-1abe-420e-91c3-3d2b09f2683c · outbound

This paper cites Srinivasan, Jonathan T.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Srinivasan, Jonathan T

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:33.348874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:33.348874Z digest=sha256:6bffae3a845e09f2c1ec3d6afd2cddd7f37ea15708878519cfae11355dc9b6da

Observation e70596fc-8f93-476d-b22e-cf3df04675fc · outbound

This paper cites Animateme: 4d facial expressions via diffusion 9 models.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Animateme: 4d facial expressions via diffusion 9 models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:33.487172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:33.487172Z digest=sha256:999a2e0eb6dc4ac284e013eef62d847b91164d970bf5e4dcfece9079013cc293

Observation c759ee61-b77b-4c89-8983-4766e67a9909 · outbound

This paper cites Arc2avatar: Generating expressive 3d avatars from a single image via id guidance.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Arc2avatar: Generating expressive 3d avatars from a single image via id guidance

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:33.571623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:33.571623Z digest=sha256:06b97403d4203102a8a690cb187a3697015607a1b9ea3412e02412865b3c8bad

Observation cd2a97ec-3f29-4ef2-8fb6-7ed8878a06f0 · outbound

This paper cites Generative adversarial nets.Advances in neural information processing systems, 2014.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Generative adversarial nets.Advances in neural information processing systems, 2014

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:33.638606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:33.638606Z digest=sha256:07ab51bc73320fb27770c270755edfd3d4d7790f3fb1cf01e1ca88c0b4913384

Observation 94908cdd-1d12-4155-8dd2-9d204ff97007 · outbound

This paper cites Neural head avatars from monocular rgb videos.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Neural head avatars from monocular rgb videos

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:33.727435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:33.727435Z digest=sha256:ec8ab740143964b29410b3b42f04325db81b79e1c2d3426dd21c617f0d358308

Observation eca14e2b-53eb-46c0-9574-a87f43f2cd10 · outbound

This paper cites SparseCtrl: Adding Sparse Controls to Text-to-Video Diffusion Models.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits SparseCtrl: Adding Sparse Controls to Text-to-Video Diffusion Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:33.805660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:33.805660Z digest=sha256:4d89d2623a32a435100fd17465a111b35407061723ea81725a2ce7065fa13b0c

Observation b76ae828-863c-45a7-affd-335edc4d1fb8 · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning.Interna- tional Conference on Learning Representations, 2024.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Animatediff: Animate your personalized text-to- image diffusion models without specific tuning.Interna- tional Conference on Learning Representations, 2024

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:33.921329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:33.921329Z digest=sha256:c871b6ce63bdf63f8c9ea97e0b3a1d9efdc100cf08a25acfefd83bb1df2827d7

Observation ca5da7bf-262e-4e4e-8475-68cc9a4d705c · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems,.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Gans trained by a two time-scale update rule converge to a local nash equilib- rium.Advances in neural information processing systems,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:33.999403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:33.999403Z digest=sha256:b25d3fe82eb7298d59504f2dde27f4a9f44417ca84f532105ba04a86ef316462

Observation 81bd2dfa-07bb-465d-bcab-de3bd1cf0323 · outbound

This paper cites Denoising dif- fusion probabilistic models.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Denoising dif- fusion probabilistic models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:34.122254Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:34.122254Z digest=sha256:7cebef5e2a4fe31afe858a0d9b652e1884e62c7760a619521dfe32beef3ff6c8

Observation 394311bd-d768-4349-899e-cbff3bb9d025 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits LoRA: Low-Rank Adaptation of Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:34.235121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:34.235121Z digest=sha256:a61c7f8025a99d1c94a594744a8a99fe8eef8a332eaa5eb28669cbf4244c7d86

Observation 3a9d8aa3-a42e-4106-8eb1-3ad214e61c56 · outbound

This paper cites Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:34.344003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:34.344003Z digest=sha256:922aa7c17719dc34a32dc6f3663c99acedea011760a071ba57f66f6f40ec8f92

Observation 3defc845-6b14-46a8-af97-b4b7b2ac0b2a · outbound

This paper cites Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:34.497342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:34.497342Z digest=sha256:ba6857c51228f09112d53ff9a92d0115ceb84d7f3725695ebd136ecd06b5a45c

Observation f7011014-8836-48d4-92f4-1778890ff7a3 · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:34.578349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:34.578349Z digest=sha256:e42753d351135068a5787a603bd0ac1576363f4399cd05c8b5c0ea28cdabfbab

Observation fbe1d7fb-4050-4695-bb60-0affa5994581 · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits A style-based generator architecture for generative adversarial networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:34.698328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:34.698328Z digest=sha256:77174883222ea07243b4834882ccf599d1771b708abbaa8ec599f409297d2634

Observation f1eebd37-7836-4255-8be3-3a7ba2365f89 · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.ACM Trans.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits 3d gaussian splatting for real-time radiance field rendering.ACM Trans

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:34.809760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:34.809760Z digest=sha256:203f67af270e905263b8072ec0c4000bcf691a589da59438e19ce0b2f4010c3f

Observation e63a812e-cf2c-4259-a52a-25acf86467c8 · outbound

This paper cites Float: Generative motion latent flow matching for audio-driven talking portrait.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Float: Generative motion latent flow matching for audio-driven talking portrait

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:34.901469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:34.901469Z digest=sha256:2e9f7f0524e032c3df08687b1fcb882c4a01a8aac635ddb3c1fae7fdca94d854

Observation fbddb0f7-2a93-4c00-bee6-6f560aa15975 · outbound

This paper cites Nersemble: Multi-view ra- diance field reconstruction of human heads.ACM Trans.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Nersemble: Multi-view ra- diance field reconstruction of human heads.ACM Trans

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:35.001074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:35.001074Z digest=sha256:83e2f30df2f01178febac1bd5e11bcef30113a09d1fed05e12fb0af95583f7dc

Observation 808651a2-bde0-4c82-9029-654d3ead2f1b · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:35.074851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:35.074851Z digest=sha256:7a791dde00744a04aeebbae1850d59d8d498d73955becb42e3be6a828a1dcef1

Observation acbb6633-f732-42e2-87ef-e714ecfe718f · outbound

This paper cites Spherehead: stable 3d full-head synthesis with spherical tri-plane representation.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Spherehead: stable 3d full-head synthesis with spherical tri-plane representation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:35.215180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:35.215180Z digest=sha256:aa999dfb3188c5279170cab3b56ffc6486b383ba5c76c6f534845599b025b99d

Observation 0c3c7130-21ce-4f89-a332-050bb1e35d38 · outbound

This paper cites an unresolved cited work.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:35.272753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:35.272753Z digest=sha256:de4b6869ff3b6f01eeaa70e9afc1a01cacdffebb056a46825cda3698a84a1353

Observation 219c7d78-1777-488f-806a-e336e39069bc · outbound

This paper cites One-shot high-fidelity talking- head synthesis with deformable neural radiance field.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits One-shot high-fidelity talking- head synthesis with deformable neural radiance field

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:35.397161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:35.397161Z digest=sha256:b4db1d776f28181bcc32b5ac4a68650ae825ba56157c34994f5b22c1a19d9dd6

Observation b1df8b29-848c-4997-886e-0825dd37a5e7 · outbound

This paper cites Generalizable one-shot 3d neural head avatar.Advances in Neural Information Processing Systems,.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Generalizable one-shot 3d neural head avatar.Advances in Neural Information Processing Systems,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:35.485676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:35.485676Z digest=sha256:ba8b258ff8a49b4742d7d8cd9090bf3d18fe44c5a500293944a547659097cb79

Observation de78be16-963a-4df0-a1bb-134c7e7f8e1f · outbound

This paper cites Im-portrait: Learning 3d-aware video diffusion for photorealistic talking heads from monoc- ular videos.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Im-portrait: Learning 3d-aware video diffusion for photorealistic talking heads from monoc- ular videos

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:35.713904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:35.713904Z digest=sha256:76ec416ed5dabcf6a571074a7aa5cb8b60df1e24fc7737cee12cc4135c737878

Observation 95bb7677-80ca-4422-8d42-5443a6169694 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:35.858555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:35.858555Z digest=sha256:7f17ad1087f6dd375ead8b65cff68abf5eb60131ee4902fc0165714662abe6d5

Observation 106a0397-d1e4-489f-99c1-90de931acb3f · outbound

This paper cites Anitalker: animate vivid and di- verse talking faces through identity-decoupled facial motion encoding.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Anitalker: animate vivid and di- verse talking faces through identity-decoupled facial motion encoding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:35.996816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:35.996816Z digest=sha256:30d489c75e4e0d6cf9d02951c8cc5b87cabb59629322fcd554c0a4ccc7004f71

Observation 7928c3f5-233c-4514-9e8a-d159f6038bfe · outbound

This paper cites Decoupled Weight Decay Regularization.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Decoupled Weight Decay Regularization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:36.083220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:36.083220Z digest=sha256:bd11a54eec82a7f2b976c698fdd5769b1365680a2ccaf15d48bfd839593024b9

Observation 2551374a-9f34-420f-bd87-8dfba11de37c · outbound

This paper cites Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.Advances in Neural Information Processing Systems, 35:5775–5787,.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Dpm-solver: A fast ode solver for diffusion probabilistic model sampling in around 10 steps.Advances in Neural Information Processing Systems, 35:5775–5787,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:36.291689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:36.291689Z digest=sha256:92718ae5b4eaa9edca76baf1d89bedf8dcb17e7797166d81c9da61cd935a7bc1

Observation 4409eef2-a474-42c9-ae30-2d1397b1cfbb · outbound

This paper cites DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:36.393186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:36.393186Z digest=sha256:8734de5592ce186a7183470845bb46a672ce13f7e13d1842af506c5e679c8165

Observation d06f37eb-e0b2-44da-9ec9-c547503dee50 · outbound

This paper cites Visual Speech Recognition for Multiple Languages in the Wild.Na- ture Machine Intelligence, 4:930–939, 2022.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Visual Speech Recognition for Multiple Languages in the Wild.Na- ture Machine Intelligence, 4:930–939, 2022

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:36.532426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:36.532426Z digest=sha256:4a41403fd4426a57c75d3c6f21f079bad847f052de2bc1040befb78c8abdb27f

Observation 6742e00e-5f10-43fe-b9e8-9e6f3c674845 · outbound

This paper cites Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Follow Your Pose: Pose-Guided Text-to-Video Generation using Pose-Free Videos

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:36.601149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:36.601149Z digest=sha256:84e7b5f452e1770f2811e8aeedd21fd5aad888a4e052431f3ba976df67d840e0

Observation 84c2ebfe-5a4c-4cd6-888a-ec9a7696e51e · outbound

This paper cites Otavatar: One-shot talking face avatar with control- lable tri-plane rendering.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Otavatar: One-shot talking face avatar with control- lable tri-plane rendering

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:36.700349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:36.700349Z digest=sha256:9a14698ecc5999c3589803cbebce5e8c65e4bdbb91789b772a1e9986e48ae702

Observation 875c9368-921a-4d6b-ba46-5987c70a55d1 · outbound

This paper cites Isambard-ai: a leadership-class supercomputer opti- mised specifically for artificial intelligence.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Isambard-ai: a leadership-class supercomputer opti- mised specifically for artificial intelligence

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:36.819914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:36.819914Z digest=sha256:5277fc224a0d7ff35f955548b46b1c47e271ddd5a6843f0e83469204dee4ddfe

Observation 6b360ebc-d83e-4f0f-93b9-4653f93eb4a9 · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view syn- thesis.Communications of the ACM, 2021.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Nerf: Representing scenes as neural radiance fields for view syn- thesis.Communications of the ACM, 2021

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:36.924756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:36.924756Z digest=sha256:4f0feadd220de88871028e140b266d13355e41484d8481b5d9b219d732eeb344

Observation 37fb9548-2a67-43a1-9c86-92146d94137b · outbound

This paper cites Elucidating the Exposure Bias in Diffusion Models.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Elucidating the Exposure Bias in Diffusion Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:37.031365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:37.031365Z digest=sha256:30521536eb9533ffc5c6ba745122afbcc9a8c458e9eefc51f2dd4d73aa252bc9

Observation e2a1faac-0c16-4d4c-ac14-03e3f58ea969 · outbound

This paper cites Arc2face: A foundation model for id-consistent human faces.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Arc2face: A foundation model for id-consistent human faces

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:37.182264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:37.182264Z digest=sha256:00c1aeb94ef1a568bb2975bb8faa10e158782eb4ae685c4e179eb5d64a698b50

Observation 99561777-a0c5-4a4c-9079-480bb99f893c · outbound

This paper cites Scalable diffusion models with transformers.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Scalable diffusion models with transformers

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:37.271905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:37.271905Z digest=sha256:f902d0bdfb0091364c00367dbdca2bfd34a11e61f812085eba58d1b0b3df2e2b

Observation 25ac8a62-8686-42e7-943f-88d069e31bff · outbound

This paper cites Movie Gen: A Cast of Media Foundation Models.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Movie Gen: A Cast of Media Foundation Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:37.377778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:37.377778Z digest=sha256:028910ce9f3c27f2f9b0f22471eb43895f4b6d8b6d95aae2a2a4d50b3ade8c9c

Observation 3ad7472c-0c50-4815-b330-02b91167216f · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits A lip sync expert is all you need for speech to lip generation in the wild

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:37.557723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:37.557723Z digest=sha256:a935caad4e3d808dfcbeb9e2097ebbbc9270663ab9b22d6b705f183506681900

Observation f1139cd7-bb04-4461-98f0-8dc917732474 · outbound

This paper cites Gaus- sianavatars: Photorealistic head avatars with rigged 3d gaus- sians.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Gaus- sianavatars: Photorealistic head avatars with rigged 3d gaus- sians

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:37.628661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:37.628661Z digest=sha256:ae44737aa6fdad26f008a6fbe57afc5e615c5abf837d9371deff34e22ec2f88a

Observation 9ea2e7f3-6988-4242-861f-416d7ffb215b · outbound

This paper cites Generalization in Generation: A closer look at Exposure Bias.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Generalization in Generation: A closer look at Exposure Bias

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:37.723334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:37.723334Z digest=sha256:dbbe76f2964b3ced072da2633989f6f57f1bfc2fcb8dc8ea70a2135f721bb9bb

Observation 8696e58f-7177-4cd2-8f24-63c0c5ab6eed · outbound

This paper cites wav2vec: Unsupervised Pre-training for Speech Recognition.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:37.850388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:37.850388Z digest=sha256:8a42bae6194c77dc878e1325774de42e68a44c2c379a533bf7612d0ea44bc2ae

Observation 94b10062-5343-4d70-af25-a483816c7257 · outbound

This paper cites Difftalk: Crafting diffusion models for generalized audio-driven portraits animation.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Difftalk: Crafting diffusion models for generalized audio-driven portraits animation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:37.996463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:37.996463Z digest=sha256:f7f513344b07802ce3dd76ceb644a45330fd5c4b2f0a277f441a6aa4f9eb7354

Observation afb30cb3-8e71-4d4d-a5c9-fd3ee329c9aa · outbound

This paper cites Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Zero123++: a Single Image to Consistent Multi-view Diffusion Base Model

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:38.106522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:38.106522Z digest=sha256:b65263bda2a977f322880f29ab9fd4e6802cab54806ac53e221eda7b34d74bcf

Observation 7d76e169-e41f-4b66-8a70-5183dbb912f4 · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:38.220611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:38.220611Z digest=sha256:81549cbb3630b8d6c6ac9bf5836df41d941d6a31dbdbff0bb87f8268d943a140

Observation 403fa4b0-489d-48a1-9b2f-ddc09129078c · outbound

This paper cites an unresolved cited work.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Unresolved cited work

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:38.374267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:38.374267Z digest=sha256:0c422c8a737fc3dd6776436811b33f9fa3157247758df00766127f1949dee51f

Observation 2182c974-2c76-440d-b922-171e293a5fce · outbound

This paper cites Emo: Emote portrait alive - generating expressive portrait videos with audio2video diffusion model under weak conditions.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Emo: Emote portrait alive - generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:38.549141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:38.549141Z digest=sha256:f2abf45040f0e21982de45b89e03a9a308ce8665d1985af172268a66db8d45df

Observation 86fa87d0-f67a-42b4-974a-3c28957a479a · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:38.620629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:38.620629Z digest=sha256:16dde52c3f4175635eb333270917cf47acc0420f7be68612faba0d295a82aceb

Observation 25212b15-ad3b-4487-90a2-96751809c19c · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Wan: Open and Advanced Large-Scale Video Generative Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:38.778119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:38.778119Z digest=sha256:a5378eaf82a7b2a1cb0995c993057ea152804024f886bb91c59f75baf8b57e3e

Observation 61b7f1b7-40a3-48d5-9cff-3319c4fa43aa · outbound

This paper cites V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:38.891584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:38.891584Z digest=sha256:62ed6adc5419bcc13a00cf62080801d415f5944b4aa2f786caed9c5d774491fd

Observation 0a529e18-e194-4a28-a382-592abd4bf2f1 · outbound

This paper cites DisCo: Disentangled Control for Realistic Human Dance Generation.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits DisCo: Disentangled Control for Realistic Human Dance Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.076074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:39.076074Z digest=sha256:be5607d3a6c03d2262a1ba946eb557b772072e3d7736a3e2753acc52f61439b8

Observation 46b4810d-e767-432a-b36c-1fd296247647 · outbound

This paper cites One-shot free-view neural talking-head synthesis for video conferenc- ing.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits One-shot free-view neural talking-head synthesis for video conferenc- ing

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.171775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:39.171775Z digest=sha256:1117911cb3c789caaa2a0923b5d07ab8d4eebbc01d7b171d151ddbc751d8337e

Observation 8c32dc19-7b63-4871-be44-8fe6b464f2f7 · outbound

This paper cites Videocomposer: Compositional video synthesis 11 with motion controllability.Advances in Neural Information Processing Systems, 36:7594–7611, 2023.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Videocomposer: Compositional video synthesis 11 with motion controllability.Advances in Neural Information Processing Systems, 36:7594–7611, 2023

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.261453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:39.261453Z digest=sha256:aee4a0f7247214abd228cd326fc05f90b24e50c8735f8a44d7f3cc93bad08547

Observation ff743ba2-f1ba-4f2e-a92c-58fd02cb41be · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.321612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:39.321612Z digest=sha256:3b9ccd3e1cfa11b440497b57ccaf09ac6690cd31c9cf78cc7ac2059b361c0964

Observation b2b28a9d-f196-468a-b568-d28665937126 · outbound

This paper cites Vfhq: A high-quality dataset and bench- mark for video face super-resolution.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Vfhq: A high-quality dataset and bench- mark for video face super-resolution

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.444135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:39.444135Z digest=sha256:9a9603f4ad8e2569bf784f5ba1ec16166f956205f8848631b37a5b1f81433ff6

Observation 697de81c-4365-4064-9772-1108bc316543 · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.500409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:39.500409Z digest=sha256:d4ea62a8423cb957b0b9c1818ace0d4573d04540760197f302d375f1ec14d1c2

Observation c7ec8840-3b21-4209-98b5-fa134dc439b1 · outbound

This paper cites Magicanimate: Temporally consistent human im- age animation using diffusion model.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Magicanimate: Temporally consistent human im- age animation using diffusion model

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.609205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:39.609205Z digest=sha256:2d8c4bf1c54597581a66ba9d09926d8b72491694dd820093784cc680f0265c60

Observation bc47b44f-f2a5-4b9d-8206-2a83171a46a9 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.690354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:39.690354Z digest=sha256:87978f6ab14ae366348a32aa538d6cde8b1226eaed69da7265443012db64a077

Observation 7ef293b0-cfe9-45be-98c4-564a80875a1b · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.753939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:39.753939Z digest=sha256:664c9bd9db4d82c16229eb4a8057151cd5264a774dc4708525f1ee3fb38fbc2e

Observation fc693490-f38d-4fd0-87b7-d482deef00c1 · outbound

This paper cites Real3d-portrait: One-shot realistic 3d talking portrait synthesis.ICLR, 2024.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Real3d-portrait: One-shot realistic 3d talking portrait synthesis.ICLR, 2024

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.826431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:39.826431Z digest=sha256:a68b43c453aec91e1c5ab05b2261f066950d185a64591de1f5dfa9eb33e44dab

Observation ac30a21f-9050-49d2-9389-d6189c4f4a24 · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.882819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:39.882819Z digest=sha256:728f5843ee2371fc49912159414bffb5e05bf86d20c08c84bd99ae40b0d9350a

Observation 76a5c066-29ab-4f57-b98b-acf33db80f00 · outbound

This paper cites Mimicmo- tion: High-quality human motion video generation with confidence-aware pose guidance.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Mimicmo- tion: High-quality human motion video generation with confidence-aware pose guidance

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.994204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:39.994204Z digest=sha256:5aa7409ed12c7267b5c476febfee354203f059918433186b8eded91868458648

Observation c35f110c-30da-43c8-8594-4d4e93781f0f · outbound

This paper cites Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:40.064138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:40.064138Z digest=sha256:30923bdbc292cb9f0d3c0d9027914b975979ab843e17bd7c6518527b98f1b97a

Observation 21ebe2c0-91ce-485d-9fc4-08c0052c3459 · outbound

This paper cites Ilsh: The imperial light- stage head dataset for human head view synthesis.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Ilsh: The imperial light- stage head dataset for human head view synthesis

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:40.120294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:40.120294Z digest=sha256:d01ea1ae76c3f2c6059f7a96ba002a024195785894c43f054330954e227f92b6

Observation 5a29f13e-d058-458c-9746-b63839fe5b4b · outbound

This paper cites B¨uhler, Xu Chen, Michael J.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits B¨uhler, Xu Chen, Michael J

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:40.190583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:40.190583Z digest=sha256:13dcd7fed9e5ddf6a47710bf7f9bbd91e45eb3aaefc8f68bd7e1d78c795ce7ea

Observation 8ebe6b7d-b312-4357-8c2b-290b972c8fe8 · outbound

This paper cites MagicVideo: Efficient Video Generation With Latent Diffusion Models.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits MagicVideo: Efficient Video Generation With Latent Diffusion Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:40.316004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:40.316004Z digest=sha256:820f1fa062b368792c0fc0b79404df700e8c1dd20cd28d9070304b767edbf44a

Observation 9860946a-f8c4-4167-ad30-9ee003bb6dbd · outbound

This paper cites CelebV- HQ: A large-scale video facial attributes dataset.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits CelebV- HQ: A large-scale video facial attributes dataset

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:40.445550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:40.445550Z digest=sha256:70c86fca82766922c87950d67414416c6285d816d3a5ad6a7b398686a8486d3e

Observation ef0c333e-bd7a-4404-9e9e-d197366969ae · outbound

This paper cites Champ: Controllable and consistent human image an- imation with 3d parametric guidance.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Champ: Controllable and consistent human image an- imation with 3d parametric guidance

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:40.568751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:40.568751Z digest=sha256:60a5030873c14b06e28f556371250e5edf447939da7919b39f6e607d6d01c9c9

Observation f5ecf44e-a19b-49ab-b12e-ae5dd013550a · outbound

This paper cites Webface260m: A benchmark unveiling the power of million-scale deep face recognition.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Webface260m: A benchmark unveiling the power of million-scale deep face recognition

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:40.730781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:40.730781Z digest=sha256:4628c0a9e5e60cf7f6b8dedd1c80677bccb11a774712e665d98bf5becad74ac9

Observation 0601cb3c-0ac3-41c5-99e2-462d3e251d58 · outbound

This paper cites in-the- wild.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits in-the- wild

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:40.887733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:40.887733Z digest=sha256:99c0b226964762e26e7a0f2f2e27142b06c33c41e93a1cfdcb4a603f36dc95d9

Observation 75d4d27b-7a0e-4cb8-af1d-913e06ed112e · outbound

This paper cites in-the-wild.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits in-the-wild

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:41.029982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:41.029982Z digest=sha256:050f6f0fa7793627aec31af91347fb918675ba24cc5250806fc3982c05ffc57e

Pith citing papers

No inbound Pith citation observations are available.