Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T18:07:26.693764Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2412.08247.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T18:07:26.693764Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T19:39:50.132135Z
A source-named dated measurement, never combined with another source.
Source: cited_works
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b60bfc47-188e-41e2-8494-d905c436ce73 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Visualvoice: Audio-visual speech separation with cross-modal consistency,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 493573c8-9508-4054-8131-58dfa4ef8dbe · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Audio-visual active speaker extraction for sparsely overlapped multi- talker speech,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8fab99be-6f86-4073-bdfe-d8918a0ed70e · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Audio-visual target speaker extraction with selective auditory attention,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 98d403b8-9a39-47f5-a643-c6e87d83afcf · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues A robust audio-visual speech enhancement model,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation eceddd33-faaa-449e-9799-8664a5f13015 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Audiovisual speaker separation with full- and sub-band modeling in the time-frequency domain,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 681def39-4c94-4964-b58a-cb98c2752a09 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues My Lips Are Concealed: Audio-Visual Speech Enhancement Through Ob- structions,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5fca883b-ea00-4f68-8d8b-29a062e35db9 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Switching variational auto-encoders for noise-agnostic audio-visual speech enhancement,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 28d5a172-97c6-4f7d-8108-ac760709df12 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Robust unsupervised audio-visual speech enhancement using a mixture of variational au- toencoders,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1a23a504-5856-4281-b145-0165c2030650 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Time-domain audio-visual speech separation on low quality videos,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 745e7521-9262-4e13-9e69-315ad84c62bf · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Imag- inenet: Target speaker extraction with intermittent visual cue through embedding inpainting,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d46e866a-3d73-4726-9d37-a8e8cd34563a · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Multimodal attention fusion for target speaker extraction,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8ab7c55d-3f6e-4602-a11a-6f43d5e24c07 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Multimodal speakerbeam: Single channel target speech extraction with audio-visual speaker clues.,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bb45d666-1c42-4dcc-ab39-5f07327e5f39 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues A two-stage audio-visual speech separation method without visual signals for testing and tuples loss with dynamic margin,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 314dd23d-c148-4bbe-a1ee-0f31e42e6c0f · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Iterative autoregression: a novel trick to improve your low-latency speech enhancement model,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 47145855-1901-4412-b07b-b3bfce7f5f03 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues An online speaker-aware speech separation approach based on time- domain representation,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0f2a94b7-8ea7-4b0f-b9d7-7af202be27ea · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Source- aware context network for single-channel multi-speaker speech separa- tion,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ec5f85cf-e046-4f85-8ef8-66629bee6d61 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Listening and grouping: an online autoregressive approach for monaural speech separation,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b1bccb2e-54ce-4264-9de5-74809139d0d7 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Paris: Pseudo-autoregressive siamese training for online speech separation,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d17dac2e-4a72-4569-b03e-cdeb744017df · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Robust Speaker Extraction Network Based on Iterative Refined Adaptation,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a95549d0-c92d-4918-bb80-334bfeaf513a · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Neuroheed: Neuro-steered speaker extraction using eeg signals,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e4874120-3ef5-46d4-b66a-2159e8564f0b · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Attentive training: A new training framework for speech enhancement,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cd17a2ad-6615-4985-9a61-2be2ec3dfdbf · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Coarse-to-fine target speaker extraction based on contextual information exploitation,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d81af2e2-c5db-4139-add9-95ec3078936b · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Wesep: A scalable and flexible toolkit towards generalizable target speaker extraction,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation bf9af09e-4c66-4939-9b8d-aab8a78dce54 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Multi-level speaker representation for target speaker extraction,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4630f28c-b9f7-4ea0-bff0-09a8de741d06 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Muse: Multi- modal target speaker extraction with visual cues,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ec1ef8fc-614b-4c50-9bb1-616fba8d1305 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Neural machine translation by jointly learning to align and translate,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 05ce3e2c-a246-4126-8ab4-dc1c88bfbacb · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Conv-tasnet: Surpassing ideal time– frequency magnitude masking for speech separation,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8c41f28-b5c4-4d8a-a58a-5dee04d51837 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues V oxCeleb2: Deep Speaker Recognition,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 93976f9a-6887-44f5-8f1f-6fe93fcf26f6 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Performance measurement in blind audio source separation,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fa09f51a-b794-4a5c-a919-5d6ff0eaee98 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b512a36e-86bf-4eda-8d64-2a98e86a5d1a · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues An algorithm for intelligibility prediction of time–frequency weighted noisy speech,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation af79fb60-7def-4cf6-9a0b-720b3fe33706 · outbound
MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues FaceFilter: Audio-Visual Speech Separation Using Still Images,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7e7e147e-96ba-4759-a544-6ae1f7d213e7 · inbound
CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.