Pith. sign in

Paper Citation Record · LEDGER

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis

As of 13 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2411.13209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13209 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:47:09.073686Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy27
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f79ae570-cab8-4f83-ab06-75077cbb4ac3 · outbound

This paper cites Simulation-based learning in higher education: A meta-analysis.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Simulation-based learning in higher education: A meta-analysis

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:10.137518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.809431Z digest=sha256:76f917cc5e652d83bd18a54e816c3ae29440db2980224f32b888668a6862ec47

Observation 9dd10fdb-49bc-4d6e-8a30-eba6a04a0831 · outbound

This paper cites Psychological foundations of emerging technologies for teaching and learning in higher education.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Psychological foundations of emerging technologies for teaching and learning in higher education

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:10.116225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.815259Z digest=sha256:6f92caebf4ed37d0f9745dce146b40f414437c63c5e2839dcc63b85c9474235f

Observation f15c7172-b18b-41a4-8d1f-96d407ea625f · outbound

This paper cites an unresolved cited work.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:47:10.097669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.821437Z digest=sha256:0f5b97ca0dff60b3286866e0c10988c9b4858570315142ff97f5d3cb756d8fb4

Observation b8c58379-42cb-4f6a-b13a-ba795041be66 · outbound

This paper cites Designing effective training programs for investigative interviewers of children.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Designing effective training programs for investigative interviewers of children

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:10.079749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.826723Z digest=sha256:6fe296e778de1d42e7e5f5c114110b0fe441fd71c8c8ace79846ebff151a40e3

Observation 7fb1a43d-1ef9-482a-9fa9-76a5447ead54 · outbound

This paper cites an unresolved cited work.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:47:10.058801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.831853Z digest=sha256:b61b3b6e0a31ae62202cb0372c482c95dfc4499db5ef64991a6783b9c1174a2f

Observation 06bbf19c-0514-4af6-aea9-aa1f47d53b63 · outbound

This paper cites Interviewing children.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Interviewing children

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:10.037742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.837457Z digest=sha256:aaa3cab4eb241639e23d8801a573712d7deebcba4666399ec2c321f50096186c

Observation d63462fa-9032-4607-a731-0f2b2fd92dff · outbound

This paper cites Tell me what happened: Questioning children about abuse.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Tell me what happened: Questioning children about abuse

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:10.013244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.845623Z digest=sha256:6cd32c309c0eb25753f105f53694e7c379a400c20461610da88658807c3ffd58

Observation 8de70706-2588-468e-9620-a250123246a0 · outbound

This paper cites An overview of mock interviews as a training tool for interviewers of children.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis An overview of mock interviews as a training tool for interviewers of children

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.996698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.854431Z digest=sha256:4a9cf328ea9e843946378c8deea91ea6ed734178903149761a9a5204cf02ad6f

Observation ef11a872-8f79-49ef-a3da-61ca9fdfec2f · outbound

This paper cites Towards an ai-driven talking avatar in virtual reality for investigative interviews of children.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Towards an ai-driven talking avatar in virtual reality for investigative interviews of children

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.979634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.860126Z digest=sha256:078b98d23397332c2a08e0dfe827db3f7dce3560dda449b7a28ecfc71737c5df

Observation e6547aa4-dfc5-4c06-ad48-0e81540c28da · outbound

This paper cites an unresolved cited work.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:47:09.962601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.866481Z digest=sha256:a04aacb495c3b5478cb3f1807f6b29897bf770ca8fe0e921374289890ec0815b

Observation 459553b4-ab63-4bf2-8d48-a903b32abb04 · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view synthesis.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Nerf: Representing scenes as neural radiance fields for view synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.871543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.871543Z digest=sha256:f41def6b175cb93f85b2fa99a9b55284094a9a71b8ddc0a93f3dc742a198741c

Observation 49f82dc4-9230-4af4-b7c9-93105dbfb566 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Robust speech recognition via large-scale weak supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.876364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.876364Z digest=sha256:1f8fcfd30f628261b3b11e7fca058cbea9bcc193374343830b4ff2568169c106

Observation 9afcc583-79bf-4bae-8a57-dd3d1b84e27f · outbound

This paper cites Whisper afe for talking heads generation.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Whisper afe for talking heads generation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.923125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.881311Z digest=sha256:19362efd626a100bc763e0106588d35b9a4b13506287fff27427d248d11b68a3

Observation 1e4731ba-445f-4ebe-85c7-2beadc0f5829 · outbound

This paper cites Technological acceptance of an avatar based interview training application: The development and technological acceptance study of the avbit application., 2021.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Technological acceptance of an avatar based interview training application: The development and technological acceptance study of the avbit application., 2021

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.906522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.886036Z digest=sha256:5f43df9b243550c0d33d73bc838850c14f0e641bb40e29dd54df812371ebab39

Observation 254f0f89-aefd-464a-8222-e8ef696b67b7 · outbound

This paper cites A field assessment of child abuse investigators’ engagement with a child-avatar to develop interviewing skills.Child Abuse & Neglect, 143:106324, 2023.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis A field assessment of child abuse investigators’ engagement with a child-avatar to develop interviewing skills.Child Abuse & Neglect, 143:106324, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.889013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.890840Z digest=sha256:dff296ef5b93535dcc739d30d912255d8365c9b558c4156540389cf7bea8e6a8

Observation 685af020-1d76-4329-ba5f-1a0148af9b01 · outbound

This paper cites Evaluation of a comprehensive interactive training system for investigative interviewers of children.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Evaluation of a comprehensive interactive training system for investigative interviewers of children

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.871318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.895764Z digest=sha256:128dfe238b36882f2b9ea5877c83751eebe68edf471d4622f77d5d4c47e0c296

Observation 702d9421-bf07-49a6-9884-87ad94a959c6 · outbound

This paper cites Training in investigative interviews of children: Serious gaming paired with feedback improves interview quality.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Training in investigative interviews of children: Serious gaming paired with feedback improves interview quality

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.842122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.900547Z digest=sha256:5b52063f04b7f8c9ddec8340aa679e5c182d792eed16fb75c35c12887d93d7bc

Observation 61322ef2-4dae-47ec-8f1a-1dda5755b770 · outbound

This paper cites How to prepare for conversations with children about suspicions of sexual abuse? evaluation of an interactive virtual reality training for student teachers.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis How to prepare for conversations with children about suspicions of sexual abuse? evaluation of an interactive virtual reality training for student teachers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.823380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.905396Z digest=sha256:02b5dadfde124c33c5e4acb5eb9d2ddee731078459035a5562f70cbcfce48f78

Observation 6d6295ca-6422-4315-b270-9bbbae55ebbd · outbound

This paper cites A theoretical and empirical analysis of 2d and 3d virtual environments in training for child interview skills.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis A theoretical and empirical analysis of 2d and 3d virtual environments in training for child interview skills

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.805367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.910586Z digest=sha256:7864a8087d7b7b5ad5fa594a56afd140abbcda80e01dc8f0e3effeb5296680b6

Observation c2fae4ba-10c1-41fd-8864-375aade7d8dc · outbound

This paper cites Live speech portraits: real-time photorealistic talking-head animation.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Live speech portraits: real-time photorealistic talking-head animation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.784817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.915545Z digest=sha256:e6423eabbc94eaae1afeb49ab475c9193d3fd4af52f631d578466c82fa6b046b

Observation 20e87e22-84b4-49f7-9d7e-c043b3586c9c · outbound

This paper cites Generative pre-training for speech with autoregressive predictive coding.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Generative pre-training for speech with autoregressive predictive coding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.764253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.921810Z digest=sha256:507187da76e35e704144c18544e42fb2520cdad8f49045f1875ad72a451af50d

Observation 968615b6-4508-44d6-aeb5-d66cba41690e · outbound

This paper cites RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis RealTalk: Real-time and Realistic Audio-driven Face Generation with 3D Facial Prior-guided Identity Alignment Network

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.927249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.927249Z digest=sha256:c266b21d15ce6feb48aa1a005178f1df00688fd94fbc80a4fba0a60ab06bd22e

Observation eccb74c8-ffda-4f2a-b55a-d225e16f4298 · outbound

This paper cites 3d gaussian splatting for real-time radiance field rendering.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis 3d gaussian splatting for real-time radiance field rendering

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.934320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.934320Z digest=sha256:da2c6f19e00ca10074de9b50d980ea6782b223fa5748dbf9bdc24b960f87dca2

Observation 73e3db43-c7ed-4143-8bb4-e062bdbeafc8 · outbound

This paper cites GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian Splatting.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis GSTalker: Real-time Audio-Driven Talking Face Generation via Deformable Gaussian Splatting

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.940014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.940014Z digest=sha256:1bdd82081a2eda15f9fe053407379c83694736b07a3c0c42fc62e9f63ad2f087

Observation b3479857-ac82-4187-914f-25a0c087d6b8 · outbound

This paper cites Gaussiantalker: Real-time talking head synthesis with 3d gaussian splatting.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Gaussiantalker: Real-time talking head synthesis with 3d gaussian splatting

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.733110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.945932Z digest=sha256:82b484a2d80f9057db38b11906a1e047afe72fa23855eaba2b0d90942874e9cf

Observation 5d2d5425-b974-4633-b021-4a6d0777df82 · outbound

This paper cites Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.951700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.951700Z digest=sha256:ba172661dff55242122ce4d6490e8aec85b432fd0fc4b268dc834a1faf58df47

Observation 340d058e-1a90-42b8-a0eb-a1ba860a06db · outbound

This paper cites Ad-nerf: Audio driven neural radiance fields for talking head synthesis.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Ad-nerf: Audio driven neural radiance fields for talking head synthesis

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.713679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.957073Z digest=sha256:315972c6b0fdf48e005041119c5992c426641cee2205bbd5cbe1150069005015

Observation 6a0bc73a-2840-4c52-a266-a925773d2d0e · outbound

This paper cites Efficient region-aware neural radiance fields for high-fidelity talking portrait synthesis.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Efficient region-aware neural radiance fields for high-fidelity talking portrait synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.693229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.962409Z digest=sha256:1252ed0f0ac834f7473680fbe0cafefa373dbdd844275f00159cf0984cd57437

Observation 64c5e362-c91c-447e-b9b9-abd19f84ba16 · outbound

This paper cites GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis GeneFace++: Generalized and Stable Real-Time Audio-Driven 3D Talking Face Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.967948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.967948Z digest=sha256:7959562d6c60b6782326fb43f3292957cc7c73f7a210811f18edfae8cd8e33b3

Observation d4842ba4-5216-4a80-80a5-f2b6cf9ca685 · outbound

This paper cites R2-Talker: Realistic Real-Time Talking Head Synthesis with Hash Grid Landmarks Encoding and Progressive Multilayer Conditioning.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis R2-Talker: Realistic Real-Time Talking Head Synthesis with Hash Grid Landmarks Encoding and Progressive Multilayer Conditioning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.974918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.974918Z digest=sha256:9cc9153f750389754ade3ba21f64e34cccf02da425b3e5890ca190794d92143a

Observation 332784ed-7baa-4373-bca8-fab16268a54e · outbound

This paper cites Deep speech 2: End-to-end speech recognition in english and mandarin.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Deep speech 2: End-to-end speech recognition in english and mandarin

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.980477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.980477Z digest=sha256:d496e7c4ab716cc9413fd9d17d1ed251fdbbfeebb6c969333b28f74a5198f205

Observation fb0395bf-b997-4d87-82f8-3aae258faa57 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.663290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.985782Z digest=sha256:a406b79b70f5f95ea928cf09c5aa0b86240144cdf58e586b3ec363f143940a49

Observation cf61b4c4-1e46-434f-b3a0-86b91a4766d9 · outbound

This paper cites Hubert: How much can a bad teacher benefit asr pre-training? In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages 6533–6537.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Hubert: How much can a bad teacher benefit asr pre-training? In ICASSP 2021-2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages 6533–6537

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.645909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:08.991193Z digest=sha256:e88a0757d8a06590d723a6451eec85fd3cfe90089973d3c0ee096b45822f7408

Observation 8f0216d4-df7f-4a21-a22a-ae3c0b888dcc · outbound

This paper cites Bidirectional recurrent neural networks.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Bidirectional recurrent neural networks

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:08.997783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:08.997783Z digest=sha256:c21fed6d6c96436ba3095f85d9f4855c1c83518982d92163a91b6a9914a02800

Observation 89851c0b-c735-49f0-8656-174646668c90 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Gaussian Error Linear Units (GELUs)

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.003205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.003205Z digest=sha256:721e7c357223eb48ca75acf6a50db9dd2de312a01923db1f3aaf145c0d73771f

Observation 058cc7c1-7830-4a7c-9327-49db50e7c999 · outbound

This paper cites vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.012586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.012586Z digest=sha256:a92c89c81846c2b1638c8a2da03c59b8c204b8a8313385d04ffabe7a59c20e0b

Observation 0dcb4ce5-457d-45d7-b455-1c25e382afc7 · outbound

This paper cites Product quantization for nearest neighbor search.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Product quantization for nearest neighbor search

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.018545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.018545Z digest=sha256:c5d02068bddc21a357493a654f78c3fa54ee6a836ed34f14e55ca057b7e0a19c

Observation bb09211e-51af-44f9-bdc5-38409a41b11b · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Librispeech: an asr corpus based on public domain audio books

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.026473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.026473Z digest=sha256:f1c7abd7d40dcaee945d6bdf5c7aa8b3cf0eb241a9fee50e38a6e5e63c739440

Observation 07afb9dd-47a8-4890-a6f3-30820aa4e43a · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Libri-light: A benchmark for asr with limited or no supervision

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.588803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:09.033426Z digest=sha256:570541da2b879c89b193cba689e9803c49b2cd5dbe6e5a5fa72630533b47c1e9

Observation f6ef059b-e217-4d56-a14b-a04a74f80190 · outbound

This paper cites Spanbert: Improving pre-training by representing and predicting spans.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Spanbert: Improving pre-training by representing and predicting spans

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.560601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:09.038620Z digest=sha256:f705ec960588a9883030da9473c783983e096261eb2903f889ecd1ac0194a4bb

Observation 3a3b8606-b8d1-438e-92b2-cead2e4d942e · outbound

This paper cites Aws polly.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Aws polly

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.527753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:09.043469Z digest=sha256:2e55bbe56558aa583c326af4fa3798f0ddf99a6c02b716dbbdd9e61e0be065ec

Observation 25de3783-3ddb-49b7-823b-b5f12ffaaf4a · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis The unreasonable effectiveness of deep features as a perceptual metric

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.048662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.048662Z digest=sha256:98d4fcc18b0c543b05554ee6c39b72993e7e54fc555a5a2fcd44fd071a29b8c4

Observation 35fb33ed-2fd1-432c-a89f-493632187426 · outbound

This paper cites Lip movements generation at a glance.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Lip movements generation at a glance

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.347158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:09.053891Z digest=sha256:5e59ba243b2240a3d6e5f2f3618d325cb653b6e9c4c72ddbc18646ac0740b5ca

Observation fc2861b9-2d0f-4d83-aeea-a5060dc8bc2d · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.058524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.058524Z digest=sha256:6849e144df47ba09f4e766d0b44510c7dafaf36fbd4da3c831b69b6c04c86c77

Observation feaa8353-5d57-414f-a279-d85d128132f7 · outbound

This paper cites Openface: an open source facial behavior analysis toolkit.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Openface: an open source facial behavior analysis toolkit

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.313668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:09.063648Z digest=sha256:8ce568f4b355ecd08ac113a7a46c87ec41457bf25f4d9671a88b727de5e2b3ee

Observation 41c8fe3c-3a32-4f46-a06b-c9de70833c37 · outbound

This paper cites Out of time: automated lip sync in the wild.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis Out of time: automated lip sync in the wild

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.068356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.068356Z digest=sha256:34a59f1637c26eea87eb4a8157612b73cdf5f647455674e15c4e1ad69812c7ad

Observation 7fb85829-6862-4cb8-9218-679ab8d50885 · outbound

This paper cites The uncanny valley [from the field].

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis The uncanny valley [from the field]

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:47:09.272448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T16:47:09.073686Z digest=sha256:48a1e3e5e5f19a915d0bd72857e86cdf9143d062c747b2b72b0c885d3e69f78c

Pith citing papers

No inbound Pith citation observations are available.