Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:46:04.812385Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 1 inbound Pith citation observation for arXiv:2506.09792.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:46:04.812385Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:46:04.622698Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T04:46:04.891505Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation dc2de33d-730c-49d9-8dc4-cc305f4e20d3 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Most existing studies focus on improving the audio-visual fusion mechanisms [1, 2, 3, 4, 5] or addressing visual cue-impaired scenarios [6, 7]
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bd51abf4-a74b-46fe-be2e-d77959f7cc4a · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bdec9dcb-8ec5-42ea-9e00-963626b2f9bd · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Dataset In this study, several experimental settings are considered: •Training Set:A two-speaker mixture training set is simu- lated following previous work [1, 6, 2, 3]
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6dfff320-741e-4807-8bea-7aaa4718b772 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 715afddc-d228-491b-bd38-b22495bee4f9 · outbound
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 554b79fa-87ed-4956-9e48-178c2cbd62e7 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f77eb3b3-5e62-42a7-a8a1-15d631fae36a · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction 62401377, Shenzhen Sci- ence and Technology Program (Shenzhen Key Laboratory, Grant No
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 277ee2e9-90dc-421c-8b38-bbdb74a5141e · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Muse: Multi-modal target speaker extraction with visual cues,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5ecfee67-d124-4cc9-b4a4-b30340e60df6 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Av-sepformer: Cross-attention sepformer for audio-visual target speaker extraction,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dc1f0523-48b4-4580-b3cb-863bb54eea93 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Avhumar: Audio- visual target speech extraction with pre-trained av-hubert and mask-and-recover strategy,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e9d1b61d-adb8-4864-84af-404c35a7c337 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Target speech extraction with pre-trained av-hubert and mask-and-recover strat- egy,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f2fce59-8df6-46fc-a264-defe611a2a0e · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction c 2av-tse: Context and confidence-aware audio visual target speaker extraction,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation fb5b5b21-c426-4760-ba87-16c0ddf158b4 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Imaginenet: Target speaker extraction with intermittent visual cue through embedding inpainting,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e0ad3710-a4cd-4595-8c2f-a4944678ed28 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Restoring speaking lips from occlusion for audio-visual speech recognition,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1d9e7c30-6caf-405d-8f3f-df8ed062ab1f · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Semantic en- coding during language comprehension at single-cell resolution,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 747ba67d-bf1f-43dd-bba5-3db1cc01ab66 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Hubert: Self-supervised speech representa- tion learning by masked prediction of hidden units,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ce3d5104-6aa5-4de9-8092-a6289ee7cd5f · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Large language model can transcribe speech in multi-talker scenarios with versatile instructions,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e133f23a-cfdd-4793-945f-bc8822d23e76 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Target speech extraction with pre-trained self-supervised learning models,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e15dd3bf-0e42-49e5-aeb5-e76d5bcad2a6 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Probing self-supervised learning models with target speech extraction,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 78c595a9-fa23-41e9-b7e4-e1ed7ad327c5 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction A large-scale evaluation of speech foundation models,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a571c4ce-40a4-4f71-b03d-16c5d536b767 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Transferring knowledge from large foundation models to small downstream models,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 62790262-403f-4144-b32a-f3b86abd677e · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Knowledge transfer from pre-trained language models to cif-based speech recognizers via hierarchical distillation,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b915031b-a18a-458b-90d7-4e2dd441ec03 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Speechtok- enizer: Unified speech tokenizer for speech large language mod- els,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 97115ffb-4df0-4dc0-9975-36526792fae7 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ddd5d4f-bc6b-4752-958a-3c196629d5fd · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction ALMTokenizer: A Low- bitrate and Semantic-rich Audio Codec Tokenizer for Audio Lan- guage Modeling,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a37c920-5726-402b-9940-466c6bb4effb · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Wavlm: Large-scale self- supervised pre-training for full stack speech processing,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c2378fd2-e3ac-4f81-96a9-249ff3ea9aac · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Roberta: A robustly optimized bert pretraining approach,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cccba63c-671b-4ce4-a72c-1fa36577a663 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Separate in the speech chain: cross-modal conditional audio-visual target speech extraction,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9e63d60d-5eac-43ff-b8f5-113ff6fe942a · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction V oxceleb2: Deep speaker recognition,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf384b2e-344f-4da7-a016-6670a52231d2 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Watch or listen: Ro- bust audio-visual speech recognition with visual corruption mod- eling and reliability scoring,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f3733b2-e3a8-4f34-9193-24faa3e99dcd · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Lrs3-ted: a large- scale dataset for visual speech recognition,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation df70e485-3872-4cd2-b00f-2584726a6d38 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Sdr – half-baked or well done?
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d35c370d-25fb-4aa2-92c0-6f0d5cbb1174 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Single-sided Real-time PESQ Score Estimation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d0040916-1427-4b42-8027-a37154c8e2ae · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction An al- gorithm for intelligibility prediction of time–frequency weighted noisy speech,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2feebede-1b54-43ca-a246-0f0eb293d75e · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction SpeechBERTScore: Reference-aware automatic evaluation of speech generation leveraging nlp evaluation metrics,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e82f49b8-00ac-4725-ad93-a1c3d3503b1e · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction How should we extract discrete audio tokens from self-supervised models?
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 673af5ad-5b81-4606-a91d-8a5f8194060c · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5426972f-cbe5-4462-b88d-d586d4b961b1 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction wav2vec 2.0: A framework for self-supervised learning of speech representa- tions,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c5af5730-58c1-48f1-870e-a77dab4ef6d3 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Learning audio-visual speech representation by masked multimodal cluster prediction,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4ac4cd68-4834-42c4-8b31-c4753ab3a989 · outbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Intuitive multilingual audio- visual speech recognition with a single-trained model,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bd51abf4-a74b-46fe-be2e-d77959f7cc4a · inbound
Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.