Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:50:49.884180Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2506.01270.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:50:49.884180Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:50:45.946937Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T11:50:50.003069Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c0bb2916-2367-4544-8184-3e624188d338 · outbound
Online Audio-Visual Autoregressive Speaker Extraction Online Audio-Visual Autoregressive Speaker Extraction
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ce03a151-9cda-4578-93e7-be345624fcf8 · outbound
Online Audio-Visual Autoregressive Speaker Extraction pseudo past extracted speech
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ec701100-3023-4995-92f5-82ac578580bf · outbound
Online Audio-Visual Autoregressive Speaker Extraction Dataset We mainly use the Lip Reading Sentences 3 (LRS3) dataset to validate our proposed method in this work [29], which is widely used in many A VSE studies [34–36]
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c4fd1eef-0518-4b67-bae9-21b283447642 · outbound
Online Audio-Visual Autoregressive Speaker Extraction All improve- ments are calculated relative to the unprocessed multi-talker speech signals
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e69b16ad-948b-4e99-82a6-270fe3813ccc · outbound
Online Audio-Visual Autoregressive Speaker Extraction The proposed visual en- coder, with its lightweight design and efficient processing, pro- vides a competitive alternative to the more complex visual en- coder
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4c98a724-67a7-4f3f-bb5c-f388384a8ce9 · outbound
Online Audio-Visual Autoregressive Speaker Extraction Some experiments on the recognition of speech, with one and with two ears,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation acacf362-dc1b-416a-8cd0-0c191d56ce1e · outbound
Online Audio-Visual Autoregressive Speaker Extraction Restoring speaking lips from occlusion for audio-visual speech recognition,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8a23a9c5-99ca-4986-88e3-c1be43aaf497 · outbound
Online Audio-Visual Autoregressive Speaker Extraction Deep clus- tering: Discriminative embeddings for segmentation and separa- tion,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 36529006-3bd3-44f2-8c80-b56bd9a14608 · outbound
Online Audio-Visual Autoregressive Speaker Extraction Conv-TasNet: Surpassing ideal time–frequency magnitude masking for speech separation,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f9d9b155-52f7-42cc-a7ad-9b5ba50224aa · outbound
Online Audio-Visual Autoregressive Speaker Extraction Dual-Path RNN: Efficient long sequence modeling for time-domain single-channel speech sepa- ration,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d5aaa220-4d04-42d1-9e13-1cf09631eb0a · outbound
Online Audio-Visual Autoregressive Speaker Extraction TF-GridNet: Making time-frequency domain models great again for monaural speaker separation,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation da0bb5f4-376a-45d2-ae30-93b40f7e729a · outbound
Online Audio-Visual Autoregressive Speaker Extraction Mossformer2: Combining transformer and rnn-free recurrent network for enhanced time- domain monaural speech separation,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 06a1a5a5-0872-4348-84f2-cde2b38d109f · outbound
Online Audio-Visual Autoregressive Speaker Extraction V oice- Filter: Targeted voice separation by speaker-conditioned spectro- gram masking,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0bcb3c14-7c28-4a3e-8859-7ed55aea51c1 · outbound
Online Audio-Visual Autoregressive Speaker Extraction SpeakerBeam: Speaker aware neural network for target speaker extraction in speech mixtures,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d0b41310-99cd-49b6-99a6-1a1969c83399 · outbound
Online Audio-Visual Autoregressive Speaker Extraction SpEx: Multi-scale time domain speaker extraction network,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 862ee3c5-15eb-4d2d-9b62-f452c8757511 · outbound
Online Audio-Visual Autoregressive Speaker Extraction Looking to listen at the cock- tail party: a speaker-independent audio-visual model for speech separation,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a9e535be-dc5c-4eec-8e24-6bd511b70bc3 · outbound
Online Audio-Visual Autoregressive Speaker Extraction Scenario-aware audio-visual TF- Gridnet for target speech extraction,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4e460beb-d76d-48e7-b698-c1f08372a26f · outbound
Online Audio-Visual Autoregressive Speaker Extraction Time domain audio visual speech separation,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 959f2c3a-c0eb-4694-89e5-c85f72243a22 · outbound
Online Audio-Visual Autoregressive Speaker Extraction Selective listening by synchro- nizing speech with lips,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f38ef15-19ca-418a-8d91-2e61a275589f · outbound
Online Audio-Visual Autoregressive Speaker Extraction FaceFilter: Audio-visual speech separation using still images,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3c11cfa0-3fa1-482c-ae8a-a436d0d73c27 · outbound
Online Audio-Visual Autoregressive Speaker Extraction Brain- informed speech separation (BISS) for enhancement of target speaker in multitalker speech perception,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5f239228-c3cf-4fd8-a929-251aae6a68f3 · outbound
Online Audio-Visual Autoregressive Speaker Extraction NeuroHeed: Neuro-Steered Speaker Extraction using EEG Signals
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e02594b-dcc2-4169-b27a-699abd8ab9cb · outbound
Online Audio-Visual Autoregressive Speaker Extraction NeuroHeed+: Improving neuro-steered speaker extraction with joint auditory attention detection,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b1beb070-f189-46f9-b306-c2d40b487bb3 · outbound
Online Audio-Visual Autoregressive Speaker Extraction V oiceFilter-Lite: Streaming targeted voice separation for on-device speech recog- nition,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 866ccd12-e1fe-49ff-b9cc-47965431eab0 · outbound
Online Audio-Visual Autoregressive Speaker Extraction Papez: Resource-efficient speech sepa- ration with auditory working memory,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1fa3fe15-e872-41fb-9b66-c67efa9a29b5 · outbound
Online Audio-Visual Autoregressive Speaker Extraction SkiM: Skipping memory LSTM for low-latency real-time continuous speech separation,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation afbcb701-099d-4d62-b252-086dbd33e94a · outbound
Online Audio-Visual Autoregressive Speaker Extraction Resource-efficient separation transformer,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a2d34585-d1dd-4d1f-ace2-1381f1283e16 · outbound
Online Audio-Visual Autoregressive Speaker Extraction RT-LA-V ocE: Real- time low-SNR audio-visual speech enhancement,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2fc7b33a-982c-4b07-871f-386bbaf5cf5e · outbound
Online Audio-Visual Autoregressive Speaker Extraction USEV: Universal speaker extraction with visual cue,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0fe65e12-aa9c-416e-a257-3586680a3dcf · outbound
Online Audio-Visual Autoregressive Speaker Extraction Real-time audio-visual end-to-end speech enhance- ment,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8db1e0e9-c7b2-4c06-9935-569a61406b40 · outbound
Online Audio-Visual Autoregressive Speaker Extraction BlazeFace: Sub-millisecond neural face detec- tion on mobile GPUs,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1a6f1f38-727e-4965-9ebd-232eee34fb2c · outbound
Online Audio-Visual Autoregressive Speaker Extraction Xception: Deep learning with depthwise separable convolutions,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8da9928e-d137-4ee7-9786-714afc877e7f · outbound
Online Audio-Visual Autoregressive Speaker Extraction PARIS: Pseudo-autoregressive siamese training for online speech separation,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e3be1cbd-ca63-4c40-8f67-cf50e492bb74 · outbound
Online Audio-Visual Autoregressive Speaker Extraction LRS3-TED: a large-scale dataset for visual speech recognition
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57341b6c-1d97-4a12-9078-2ee72405c140 · outbound
Online Audio-Visual Autoregressive Speaker Extraction Deep lip reading: A comparison of models and an online application,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ded28af7-8f67-41cd-adad-d2195784ee71 · outbound
Online Audio-Visual Autoregressive Speaker Extraction MuSE: Multi-modal target speaker extraction with visual cues,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 993a4242-7a2b-4b70-8726-b71ff01b3839 · outbound
Online Audio-Visual Autoregressive Speaker Extraction SDR– half-baked or well done?
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 000c2997-b4e3-464e-95b6-6ce93a75e03c · outbound
Online Audio-Visual Autoregressive Speaker Extraction A hybrid continuity loss to reduce over-suppression for time-domain target speaker extraction,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 506f3ef7-487d-4a88-9952-9d0921026f86 · outbound
Online Audio-Visual Autoregressive Speaker Extraction ReVISE: Self-supervised speech resynthesis with visual input for universal and generalized speech regeneration,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bfb8f8cd-2e02-4c74-ac72-95ab2da7280b · outbound
Online Audio-Visual Autoregressive Speaker Extraction PIA VE: A pose-invariant audio- visual speaker extraction network,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6722909d-3a70-46f3-9a83-0ef8ada3c5d3 · outbound
Online Audio-Visual Autoregressive Speaker Extraction Audio- visual speech separation in noisy environments with a lightweight iterative model,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f4a7fef-d652-4b69-9a11-f4f5e11d8f50 · outbound
Online Audio-Visual Autoregressive Speaker Extraction Adam, a method for stochastic optimiza- tion,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ef72e77f-0bb0-4e40-8879-c998dffb32e6 · outbound
Online Audio-Visual Autoregressive Speaker Extraction Performance mea- surement in blind audio source separation,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9370e7b9-421e-42cd-9b3a-03339db643a6 · outbound
Online Audio-Visual Autoregressive Speaker Extraction Per- ceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eed5317f-9f84-4542-be5a-f322163b464a · outbound
Online Audio-Visual Autoregressive Speaker Extraction A short- time objective intelligibility measure for time-frequency weighted noisy speech,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83bd3914-87fb-43d7-a216-331d466c124d · outbound
Online Audio-Visual Autoregressive Speaker Extraction V oxCeleb2: Deep speaker recognition,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2d9ecaa1-10c4-408e-ab9c-fa32ea251442 · outbound
Online Audio-Visual Autoregressive Speaker Extraction TCD-TIMIT: An audio-visual corpus of continuous speech,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c0bb2916-2367-4544-8184-3e624188d338 · inbound
Online Audio-Visual Autoregressive Speaker Extraction Online Audio-Visual Autoregressive Speaker Extraction
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.