Pith. sign in

Paper Citation Record · LEDGER

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation

As of 8 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 2 inbound Pith citation observations for arXiv:2508.06511.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.06511 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:43:27.553611Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:01:54.160626Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:13:27.684326Z

Reference resolution

71 of 71 outbound references displayed

  • verified exact0
  • verified fuzzy57
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 23e7f1ae-2355-4938-bc93-9c8cc521c53c · outbound

This paper cites Spatio- temporal energy-guided diffusion model for zero-shot video synthesis and editing,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Spatio- temporal energy-guided diffusion model for zero-shot video synthesis and editing,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:38.779195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:20.438409Z digest=sha256:8536f42694fda75d6090d66de30d4264306517d98a6d94f7f6c1a39c30153ef6

Observation 61d0c4ad-4374-47cc-be30-e068769bfe9b · outbound

This paper cites Tvg: A training-free transition video generation method with diffusion models,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Tvg: A training-free transition video generation method with diffusion models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:38.565940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:20.607686Z digest=sha256:fbc9997571a1ab1e383acd953377288b561740dbf907c244ad2ad546cf9a54a1

Observation 0d1cd090-bf90-4703-a118-f3a4e125a3dc · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:20.763128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:20.763128Z digest=sha256:b5dfd072cd8fba561e659d3182beebfa595f5c424b812c06c9139985aa1fa717

Observation 76b08039-2173-4305-a945-c8532ab21bea · outbound

This paper cites Corrtalk: Correlation between hierarchical speech and facial activity variances for 3d animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Corrtalk: Correlation between hierarchical speech and facial activity variances for 3d animation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:38.365671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:20.872571Z digest=sha256:67908c2f1884fed020e380815b11d0ec03070400337121bf32982930a049f02b

Observation 15db9d35-dc96-49b4-b9bb-608b5e91ebef · outbound

This paper cites Wonderjourney: Going from anywhere to everywhere,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Wonderjourney: Going from anywhere to everywhere,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:38.228910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:20.959506Z digest=sha256:49e7f8976e8e63a7c6c3900b73552faf7d1d1cd01e04149164d1aa15bb0adaef

Observation ac4c71df-2336-42b2-87ca-4dfcb924d932 · outbound

This paper cites Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Omnihuman-1: Rethinking the scaling-up of one-stage conditioned human animation models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.998274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:21.076719Z digest=sha256:700b0cb23a716d66236eac8e0ae016db9aabc7e1767e3774f265b61ffe45181a

Observation a84b4c09-8d90-4c46-b4de-15bb8d934705 · outbound

This paper cites Audio-semantic enhanced pose-driven talking head generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Audio-semantic enhanced pose-driven talking head generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.769317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:21.154548Z digest=sha256:70fb9d1572a0275887b8dcfeb5ef08d7588ff699fb84481135557d47c286c873

Observation b29b8937-4508-40fc-9517-030d68a8c97b · outbound

This paper cites Alleviating one- to-many mapping in talking head synthesis with dynamic adaptation context and style adapter,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Alleviating one- to-many mapping in talking head synthesis with dynamic adaptation context and style adapter,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.587619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:21.325531Z digest=sha256:64efc44c84c9ce64524671e9c83c47807083605b503701385e9ff943d5b7bd9f

Observation 76152d6b-50ad-45db-a1ca-01b56ee10fb1 · outbound

This paper cites Stochastic latent talking face generation toward emotional expressions and head poses,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Stochastic latent talking face generation toward emotional expressions and head poses,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.403599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:21.435507Z digest=sha256:b6e63b10e24cea4c00755de902297e247f0ee97608ef5e2f2593e51a9d07f04d

Observation d19d0b5e-9631-4218-a31d-96a41128f5e8 · outbound

This paper cites Hallo2: Long-duration and high-resolution audio-driven portrait image animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Hallo2: Long-duration and high-resolution audio-driven portrait image animation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.249313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:21.543133Z digest=sha256:842f216034edf6d599b6f0590fc75725e7ab149bd4a5622d0ce4daa404095868

Observation 452e0136-51ca-4206-a242-b16880eb494e · outbound

This paper cites Out of time: automated lip sync in the wild,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Out of time: automated lip sync in the wild,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:37.073999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:21.646913Z digest=sha256:a41d1e85d28cd64d6a6e18ca39d72159e358baa9d0bd854a606a5d3e9d83b0d9

Observation d97cfdfd-9130-4f64-aa87-18fec9500e84 · outbound

This paper cites Styletalk++: A unified framework for controlling the speaking styles of talking heads,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Styletalk++: A unified framework for controlling the speaking styles of talking heads,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:36.881020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:21.773061Z digest=sha256:06b96970ab8fab87b5016f6336c869d8b7d838c74598a40c06b7dfbd4a477acc

Observation 1788aa76-1a4c-4014-80e0-d6f57b38682f · outbound

This paper cites Multimodal inputs driven talking face generation with spatial–temporal dependency,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Multimodal inputs driven talking face generation with spatial–temporal dependency,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:36.666264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:21.885332Z digest=sha256:0cea0a26cd9d176de956c52ee145bf37045e5384589c14cfede87e0d8772d99f

Observation 6d903a5b-71d8-4679-b2f9-6a477f61c1e7 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation A lip sync expert is all you need for speech to lip generation in the wild,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:36.466007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:22.054746Z digest=sha256:cb11dd8edcaa08805122d4ae6cdcb345d6c02cead0db580fa26ba01a3fbc75fc

Observation 6d23a5b6-e5de-4f06-b5a4-c1ed4db6f079 · outbound

This paper cites Difftalk: Crafting diffusion models for generalized audio-driven portraits anima- tion,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Difftalk: Crafting diffusion models for generalized audio-driven portraits anima- tion,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:36.226831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:22.113727Z digest=sha256:8d03aeb04e2d496078444b0c55a2f2e1d5196361b22c0944d254a213c809ecc9

Observation 717e74f8-781b-4d37-8072-472b6c6aad71 · outbound

This paper cites InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation InstructAvatar: Text-Guided Emotion and Motion Control for Avatar Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:22.208132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:22.208132Z digest=sha256:92b6ffefc55a7d80947fc18541afb2a98bf2bee19d9ac846e90a6555fe56871b

Observation 88dafea1-90d0-4a4f-91e9-b5bf13a1187e · outbound

This paper cites Moee: Mixture of emotion experts for audio-driven portrait animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Moee: Mixture of emotion experts for audio-driven portrait animation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:36.088689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:22.291883Z digest=sha256:2993a477a7fb2247949768e0bbb49eed982c7b3ce3a5a4ca949da195bc2a9b3a

Observation d38c91a5-5f69-4b5b-a381-49458102688f · outbound

This paper cites Face recognition based on fitting a 3d mor- phable model,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Face recognition based on fitting a 3d mor- phable model,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:35.845916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:22.352937Z digest=sha256:8b54a5d42f55fd969650a117e05576c45485c13f95366c579658ccc78eaaa485

Observation 35c75f7c-e5a3-489b-8dce-e3561e78b74a · outbound

This paper cites Echomimic: Lifelike audio-driven portrait animations through editable landmark conditioning,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Echomimic: Lifelike audio-driven portrait animations through editable landmark conditioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:35.645001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:22.461146Z digest=sha256:840fc2aba86136fb8e57f487f328c4232314c6d9c0353d01412bf6c23d8cee58

Observation b95cff9c-01c3-4fe1-92c9-21f702ab2b5e · outbound

This paper cites Hallo3: Highly dynamic and realistic portrait image animation with diffusion transformer networks,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Hallo3: Highly dynamic and realistic portrait image animation with diffusion transformer networks,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:35.470591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:22.572249Z digest=sha256:d1482d0737fc5d7ee44705a089e11be8a03812370414e36943e751912f6541f0

Observation 791123d2-719c-4b7b-9375-4d5aab3d8325 · outbound

This paper cites Styletalk: One-shot talking head generation with controllable speaking styles,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Styletalk: One-shot talking head generation with controllable speaking styles,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:35.296235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:22.625988Z digest=sha256:d6243b4cc682e917e46c4c827b360a26e2ff301b1b7935579a5916dff6a6b834

Observation d8ebe8a7-1bc3-4514-86e7-30e255eb9b84 · outbound

This paper cites Style2talker: High-resolution talking head generation with emotion style and art style,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Style2talker: High-resolution talking head generation with emotion style and art style,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:35.101665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:22.709334Z digest=sha256:a1c42f0e70e4484e59457add3406ad805e6be6f957de7805c80c56fba0a498d3

Observation 0dbb7566-bc70-4c97-a2e8-b83e276edda7 · outbound

This paper cites Edtalk: Efficient disentanglement for emotional talking head synthesis,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Edtalk: Efficient disentanglement for emotional talking head synthesis,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:34.858085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:22.797427Z digest=sha256:12e0898c3ccdf35fd07f7a6e8819ff6b742cc455236eee96ba61c72507fdd544

Observation c14acd06-90e2-43af-a9dc-f3c400550026 · outbound

This paper cites Say anything with any style,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Say anything with any style,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:34.701760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:22.909984Z digest=sha256:c4d8af9a0c0cfbc15b66333c4aa66677fb4f8f82c914ae29c7bba4b1cf5b2849

Observation 0d175107-6644-4bba-96a0-2bc4906775bd · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:34.476957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:22.942861Z digest=sha256:044db0f4eb8b0fd872cba2fc870afc10a74ee1ab9a086d73af1a9351a8e0d091

Observation b24e3996-28ca-44c3-81fa-d4db7c263a02 · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:34.307021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:23.033694Z digest=sha256:b2234724273404258167f04e59b69806b45b85396ac442cfec6d8e40225995c8

Observation 46cb5150-1c18-4cd5-a097-6fe6d9085ed0 · outbound

This paper cites Real3d-portrait: One-shot realistic 3d talking portrait synthesis,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Real3d-portrait: One-shot realistic 3d talking portrait synthesis,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:34.138099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:23.090722Z digest=sha256:230d4c18ef9fa5a43b3f0f96a8a5c6f6890f488c1da6ebbab237ab8e5f501872

Observation 7c348509-6106-48d2-b11a-2e557bc8570f · outbound

This paper cites Scalable diffusion models with transformers,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Scalable diffusion models with transformers,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.907031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:23.147505Z digest=sha256:ac69e449a416be5a4051c282073a7f8bb05f1b6b221f044db80328fde330f42a

Observation 6c8df61a-d7fe-40cd-9420-e1bc79c41d92 · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Cogvideox: Text-to-video diffusion models with an expert transformer,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.727718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:23.241228Z digest=sha256:e882f9a80d21b304db6b7d5e2f52b75ca9ae2e740b3627a5dd25ac02a01acd90

Observation fd53d714-bd38-4778-b867-a7e1fb121158 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:23.332696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:23.332696Z digest=sha256:49dc01166659d5df0e926f000ea2a179fe5db1e421568237e9f982c91d30d40d

Observation 64874710-1360-4236-96b1-bf5fe39e4031 · outbound

This paper cites Efficient emotional adaptation for audio-driven talking-head generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Efficient emotional adaptation for audio-driven talking-head generation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.566500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:23.427386Z digest=sha256:de34c85e1503233dc6d75e5b17d0bb690e72e21084d56ea3d748fe8d7bd458d5

Observation 86f3dfd8-79ed-4ef6-b804-61215e02fe25 · outbound

This paper cites Talkclip: Talking head generation with text-guided expressive speaking styles,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Talkclip: Talking head generation with text-guided expressive speaking styles,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.404635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:23.500353Z digest=sha256:828a53951db9533959b77c0ec38e42a5ce3f37e789e39e8c879fb95db4f0e5de

Observation 9faeff50-dbcb-4bb2-8b88-134aec1e951d · outbound

This paper cites Whisperx: Time-accurate speech transcription of long-form audio,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Whisperx: Time-accurate speech transcription of long-form audio,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:23.589831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:23.589831Z digest=sha256:4e80a367297355ccab164a54618b6891f146e1a6ec63e0dbf4aa439ec16e0c6d

Observation 499289a0-6f66-421d-af3e-61575593d220 · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:23.684599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:23.684599Z digest=sha256:9beda31f46becc79c01c37a0f4ccd7df53b552b6fe5e809166ee7dbd1b5191fd

Observation 9da71d78-ef20-4e73-bb0d-8d84202ec70f · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:23.782760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:23.782760Z digest=sha256:dbfa4ff351dc9bb5b23ab55be89e977277b06be9d57ea024db14f64e5ac745cc

Observation 85e08829-a9c3-4242-8017-f0cb3d426677 · outbound

This paper cites Rep- resentation alignment for generation: Training diffusion transformers is easier than you think,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Rep- resentation alignment for generation: Training diffusion transformers is easier than you think,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.244426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:23.845045Z digest=sha256:afa145dcc98748110b9eb42a1eb268daa0c851726637a6b02ad2ccdb2bafa5cd

Observation 73f9ec82-391c-47bf-95dd-a89513ed5f6a · outbound

This paper cites Dinov2: Learning robust visual features without supervision,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Dinov2: Learning robust visual features without supervision,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:33.038179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:23.915048Z digest=sha256:431a79a5672417b3f2bcdb4f40f49d714347b67dc50e71eabb334b8e2a0be8da

Observation 0c8471b9-20c6-470d-bb11-dddd135e1d93 · outbound

This paper cites Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:32.847464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:23.985124Z digest=sha256:c977b6a66440b1011fa2eafeae92e70cabcd192ab3370feca58a56cf1e8d6c3b

Observation 8be51ba6-c19e-4c1e-a37d-b2156889883f · outbound

This paper cites CelebV-HQ: A large-scale video facial attributes dataset,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation CelebV-HQ: A large-scale video facial attributes dataset,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:32.636580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:24.080905Z digest=sha256:d8acf41907cb95bb1229db82d5ade90bc68c62463a3c8b832807fe9d0f8334c6

Observation fe30ef19-9d24-4883-be37-144cd3bed17a · outbound

This paper cites Hierarchical feature warping and blending for talking head animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Hierarchical feature warping and blending for talking head animation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:32.433129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:24.238790Z digest=sha256:91ea72f80ac2b7f14fb7dcefd22768075f47b8bbb097edcae66c7c31941b61a5

Observation 75d24d69-e53d-4113-a87b-be33cd9f0c26 · outbound

This paper cites Denoising diffusion probabilistic models,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Denoising diffusion probabilistic models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:24.356376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:24.356376Z digest=sha256:6ccbb62bd58239696cd2779af148973139a689850c565c9482d0f5879fa2b712

Observation 7a07d779-5799-4c25-bbcc-a72fdf3677c9 · outbound

This paper cites Diffused heads: Diffusion models beat gans on talking-face generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Diffused heads: Diffusion models beat gans on talking-face generation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:32.261330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:24.523648Z digest=sha256:8acdeffe762a66a0069c3ad46dea10d64791b64a9aec462be7abaa30b1ac82d3

Observation e1da9b9c-3384-4687-9fd7-cb1f31ca2b0f · outbound

This paper cites Loopy: Taming audio-driven portrait avatar with long-term motion dependency,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Loopy: Taming audio-driven portrait avatar with long-term motion dependency,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:32.085060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:24.615081Z digest=sha256:bdb5c2a8cc493ebeb869bbcb2954bd2aac7c34068236dbfebaa3e3428403a596

Observation 977c85de-1616-4d4b-978d-883bfa9f98d4 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:24.682758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:24.682758Z digest=sha256:82dc4937ceeec6635a3437902be37732306fff9c0f2afbac5f1e9e24525d1048

Observation ca2ad9aa-9256-48cc-b7c2-f94119409e40 · outbound

This paper cites Efficient emotional adaptation for audio-driven talking-head generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Efficient emotional adaptation for audio-driven talking-head generation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:31.863138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:24.780806Z digest=sha256:3a286b77d4ab6117055d609b372ab39d20c68ac5dfb3d808cf7bb8a861e7d4a6

Observation bbde5224-6056-47f0-aa05-f09de5ef64a0 · outbound

This paper cites Emmn: Emotional motion memory network for audio-driven emotional talking face generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Emmn: Emotional motion memory network for audio-driven emotional talking face generation,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:31.669046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:24.877887Z digest=sha256:8ba0d6b420a604958461195be0811e142f1627e88e83b4457600e4e5451270d0

Observation 001f16cd-4748-4652-a42f-0e9d09616507 · outbound

This paper cites Talking face gener- ation with audio-deduced emotional landmarks,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Talking face gener- ation with audio-deduced emotional landmarks,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:31.452342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:24.940996Z digest=sha256:ddf9319efd57601cbf9a95ab0f4c1dc507ba3cdf323d680d3b42da1a4bf3e986

Observation 82de0386-a1b7-4c06-8e82-c8e3a89ccabe · outbound

This paper cites Eamm: One-shot emotional talking face via audio-based emotion-aware motion model,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Eamm: One-shot emotional talking face via audio-based emotion-aware motion model,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:31.294588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:25.046119Z digest=sha256:8bb3b0acccc090a4acdf0b610e982ac1b201863adf060626ddd3142a3a00f645

Observation e06e0b86-af4b-4930-8ac4-68a5f866dc0d · outbound

This paper cites Neural discrete representation learning,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Neural discrete representation learning,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:31.103654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:25.180962Z digest=sha256:b452ba5503b9561183631ecdec0c30cbb2a8f14f47918cecffdb13269df0aa1f

Observation 1df25f53-5755-424b-b07b-c463d8e2d844 · outbound

This paper cites Progressive disentangled representation learning for fine-grained controllable talking head synthesis,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Progressive disentangled representation learning for fine-grained controllable talking head synthesis,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:30.839573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:25.288675Z digest=sha256:1f1e184100124e7083e33355a6f07480a9aa31dadfbb5546426d3be826827fec

Observation 4841976a-bf37-4262-a744-c86f37d48914 · outbound

This paper cites PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:25.371332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:25.371332Z digest=sha256:9193cf2f70c15173fc11fedd31b246a5396185388256b5a9cee4567e687dd1e1

Observation 12e53d0b-d73e-49f8-9cd5-eaf8efb3e117 · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Animate anyone: Consistent and controllable image-to-video synthesis for character animation,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:30.685967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:25.473087Z digest=sha256:fd4d15906e2ecbfdf7a7acd64039ea1960a9e9ae61f9e7d8621cadae1757bca5

Observation d8204951-6415-48ad-8c77-e803742d4b90 · outbound

This paper cites Megactor- σ: Unlocking flexible mixed-modal control in portrait animation with diffusion transformer,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Megactor- σ: Unlocking flexible mixed-modal control in portrait animation with diffusion transformer,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:30.479263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:25.616433Z digest=sha256:7ef05f11da6fb45398dc1014bc918f93d18b6b43e4716eeecb0910dd46948508

Observation fedecec2-efc7-46d6-9efa-406140d0c4be · outbound

This paper cites Vasa-1: Lifelike audio-driven talking faces generated in real time,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Vasa-1: Lifelike audio-driven talking faces generated in real time,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:30.267860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:25.797404Z digest=sha256:f90228c9d8d1600d3b8b7733047cee3bc7392ee08d56194d0c74ad6e9183dabd

Observation d810176f-352f-4331-9db9-d1a337d234a0 · outbound

This paper cites High- resolution image synthesis with latent diffusion models,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation High- resolution image synthesis with latent diffusion models,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:25.877250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:25.877250Z digest=sha256:e3be7f82ed2c740be3a21ca95e2359b10cd2c06d584702718a8d9e9564fac493

Observation 1395334b-d60f-4ab7-82ed-0b771d9c44b9 · outbound

This paper cites Easyanimate: A high-performance long video gen- eration method based on transformer architecture,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Easyanimate: A high-performance long video gen- eration method based on transformer architecture,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:25.953112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:25.953112Z digest=sha256:1debee24398dac6bf8649a6c60e43d6dea76d5fc99d3f280ad45557754860c65

Observation c4f017e6-91b8-4be8-aa3d-ed68e2a76e37 · outbound

This paper cites Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:26.010676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:26.010676Z digest=sha256:d8042f299f81061869577a0e34d51c9fd3b02ba5719437c47eb78a2c67ed2d60

Observation 99c38c09-4ff1-48b9-b933-742b6633e492 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:30.081629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:26.094954Z digest=sha256:98b740be969f2ef2d063242dba86e4ab5a776f231e8483d111d53865fa9744d4

Observation 07ce0a9e-4aa9-42c5-9f5c-51fbb3c3124a · outbound

This paper cites Effective whole-body pose estimation with two-stages distillation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Effective whole-body pose estimation with two-stages distillation,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:29.879551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:26.179445Z digest=sha256:3106b1128b809aa53ac127156a597992f78eb5f1102cc084ec4dc805961d456d

Observation 97dc7f49-0224-44b6-967f-5e19fca39027 · outbound

This paper cites Stylecrafter: Enhancing stylized text-to-video generation with style adapter,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Stylecrafter: Enhancing stylized text-to-video generation with style adapter,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:29.655657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:26.283689Z digest=sha256:1d07c4b955432d2d6c802b6642dd5b24786b1bb3c617af87cd6dd063fa5d2903

Observation b0e833b1-9f78-4b4d-97ba-b86315713c97 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Learning transferable visual models from natural language supervision,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:29.491485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:26.373131Z digest=sha256:340cbe852b9f0ab57f1c214f74d619020673b7a7477eb60824ed58421e89cc1e

Observation 0bc9ccad-1c30-45bf-bf48-2b8db2da1e75 · outbound

This paper cites Vision transformer with quad- rangle attention,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Vision transformer with quad- rangle attention,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:29.346309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:26.457543Z digest=sha256:57597b24f032bf7ab02f0ecf8bdc1049d961cca5c51497c0748342178b4f204b

Observation 10ac8c89-d956-4722-8868-8d3baa0e54ed · outbound

This paper cites Celebv-text: A large-scale facial text-video dataset,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Celebv-text: A large-scale facial text-video dataset,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:29.150137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:26.612613Z digest=sha256:784189a4706d64d2817e3d72744ad7fa940b3adaa5580a745ae7f27ba034c5f2

Observation eb870c7b-82ee-490b-9c7b-310cc136010f · outbound

This paper cites DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation DH-FaceVid-1K: A Large-Scale High-Quality Dataset for Face Video Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:26.732825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:26.732825Z digest=sha256:0cd71d24840e8c15ea6488066f5de880464c7d8625df04eb0867cd65fa43e681

Observation cfad3b58-5566-457a-8896-b88f521e39f0 · outbound

This paper cites Mead: A large-scale audio-visual dataset for emotional talking-face generation,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Mead: A large-scale audio-visual dataset for emotional talking-face generation,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:28.940438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:26.853076Z digest=sha256:3839d6152e83a0f76aa1312ef7b6c6d95d5bcd092f5ca1f76af366cc0c133064

Observation 74c75899-b98c-4529-acdf-2cdec7496d4e · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:27.004117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:27.004117Z digest=sha256:6c561fd0180335ac26270c3df597706030e425be4bc048bab80d32f1598ca21b

Observation 6d984e16-3249-45e2-a821-dcc5d65678b2 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Gans trained by a two time-scale update rule converge to a local nash equilibrium,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:28.758591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:27.132662Z digest=sha256:de00c48b606e24634a2395b4a7bcab55b56b08720c682d8705b174d667cc5c80

Observation d756d998-225c-4e56-8b4b-280b72fec00a · outbound

This paper cites Video-to-video synthesis,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Video-to-video synthesis,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:28.559570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:27.290711Z digest=sha256:ef16b42baa7c52fb2e070f807e1719272705c3d57d142b2726c11052ff6f937c

Observation 7db45693-961c-4dc3-afe1-336b58182499 · outbound

This paper cites Ani- mating arbitrary objects via deep motion transfer,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Ani- mating arbitrary objects via deep motion transfer,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:28.352098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:27.374229Z digest=sha256:3066737434da0f76d6889fecba6d475eaa0f475a809e167c4c184882c7fe1cd2

Observation fb8e6625-bf93-43b8-9a23-5b65c68f696c · outbound

This paper cites Seeing what you said: Talking face generation guided by a lip reading expert,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Seeing what you said: Talking face generation guided by a lip reading expert,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:28.145129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:27.454581Z digest=sha256:b16f89131dfda55e93518ea504ca5dea8d686ae9ffc474616c3d8442b76c42e2

Observation 64953955-4a54-4cc1-a894-2bdbaeebdd65 · outbound

This paper cites Towards robust blind face restoration with codebook lookup transformer,.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Towards robust blind face restoration with codebook lookup transformer,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:43:27.914596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T12:43:27.553611Z digest=sha256:4bc53e3681647b1f861f07ab7896898df2b88553bf99212cb4415efce0d55c0d

Pith citing papers

Observation d773345f-2744-4f8c-a9d3-c70dc31562d1 · inbound

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning cites this paper.

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:13:27.685726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:03:27.312548Z digest=sha256:95a45d093617c970f0996561004a76cb2de268dad073bb2a376ad88226b7021b

Observation be47e19f-0ab2-41bd-8864-6cb4b794748e · inbound

LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models cites this paper.

LaCache: Exact Caching and Precision-Adaptive Inference for Diffusion Large Language Models DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:01:54.160626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T01:01:54.160626Z digest=sha256:c1f46ffdfff56a1677288f68e49502a905702265f084a5624a567e57331e3123