Pith. sign in

Paper Citation Record · LEDGER

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions

As of 14 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 1 inbound Pith citation observation for arXiv:2507.15294.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15294 v2

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:40:09.564250Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T19:39:50.277622Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact0
  • verified fuzzy72
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 83c6922e-33a6-4d1b-9c21-65841d11fc41 · outbound

This paper cites Some experiments on the recognition of speech, with one and with two ears,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Some experiments on the recognition of speech, with one and with two ears,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.175918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.175918Z digest=sha256:07e3c67046202996acefae6c46e2c3041bb2adc74f537876efac532b96110fc4

Observation 744f2f48-ebd7-4a45-8006-d2f53eb17f8b · outbound

This paper cites The cocktail party phenomenon: A review of research on speech intelligibility in multiple-talker conditions,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions The cocktail party phenomenon: A review of research on speech intelligibility in multiple-talker conditions,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.181286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.181286Z digest=sha256:d15bdae25b062826541b0ab3431270656cc4525dd8107ea228bf80510014ce4c

Observation 8b3c2740-4d0c-4fe9-9692-f331dc2aa051 · outbound

This paper cites Seeing to hear better: evidence for early audio-visual interactions in speech identification,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Seeing to hear better: evidence for early audio-visual interactions in speech identification,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.949415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.186675Z digest=sha256:ccd84c8a92bad281da95b6459b1031cb0df4eeb9f0245f6a18221d3a015e7e2e

Observation 95ec1176-f5ff-4daf-aed3-5b0f40666070 · outbound

This paper cites Visual contribution to speech intelligibility in noise,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Visual contribution to speech intelligibility in noise,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.932674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.191605Z digest=sha256:68f969cc3918ffde1f6bd59e279a3a093e9621f880fc8ead2d33991ed5768b28

Observation 4ff5073d-3358-4e53-8caa-ed50b21744a5 · outbound

This paper cites Muse: Multi-modal target speaker extraction with visual cues,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Muse: Multi-modal target speaker extraction with visual cues,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.915156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.196835Z digest=sha256:fa6e0ad9d154518606068d481516726bd6c2048ba9086634b7ace97e10598171

Observation 2763e41b-075b-4632-925b-f5e4921a5a8c · outbound

This paper cites Usev: Universal speaker extraction with visual cue,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Usev: Universal speaker extraction with visual cue,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.885618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.201540Z digest=sha256:00729dae922ddd65b15ac02b2963a4ac9070c0e5be17e41350cedb4f4587a274

Observation d1c517a4-72d4-4dbd-b3a0-a2c100695cab · outbound

This paper cites Time domain audio visual speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Time domain audio visual speech separation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.866995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.207358Z digest=sha256:14ac2753dc9cc1e80983845a60807e75c52643187b22f877b4ef6235542a2fc8

Observation 094b854b-5973-4c6b-a35e-cf83b49c2596 · outbound

This paper cites Neural target speech extraction: An overview,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Neural target speech extraction: An overview,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.212164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.212164Z digest=sha256:f7f27fc1c2683c28f5dd170c595020140cf1402c4aa65db126d2d9fca3c7baca

Observation a76edd7b-75fe-4c81-b09d-a2a30fc5197c · outbound

This paper cites An overview of deep-learning-based audio-visual speech en- hancement and separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions An overview of deep-learning-based audio-visual speech en- hancement and separation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.841856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.217781Z digest=sha256:5b9ee29d8061a0c77f8df9749ba6fbfd16037e0492a92c84c15ff92a1ecbd41a

Observation 1a4b2d8f-8080-4e8e-a875-1c91d982f1bd · outbound

This paper cites PIA VE: A Pose-Invariant Audio- Visual Speaker Extraction Network,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions PIA VE: A Pose-Invariant Audio- Visual Speaker Extraction Network,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.826667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.222557Z digest=sha256:0900d2926992b0866966cd526a3adf91d84166bddbf91200985635094a7c6499

Observation 7d13eadd-9117-4305-83c8-634b94bccd40 · outbound

This paper cites Speaker extraction with co-speech gestures cue,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Speaker extraction with co-speech gestures cue,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.809836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.227254Z digest=sha256:d00913eb91365a9194cb9edda4b9c70da3e5ffcbbdbc49bf7a17515ec5143c0a

Observation 0f02fc49-23af-4400-a7a3-c539854bd177 · outbound

This paper cites Rethinking the Visual Cues in Audio-Visual Speaker Extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Rethinking the Visual Cues in Audio-Visual Speaker Extraction,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.792840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.231505Z digest=sha256:528aada3f882d0a35ddfeb51cef1ed5274c5b6923a14833364e53bea329a05d2

Observation 302839d1-bece-43a4-8150-2868d3c52695 · outbound

This paper cites FaceFilter: Audio- Visual Speech Separation Using Still Images,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions FaceFilter: Audio- Visual Speech Separation Using Still Images,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.767074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.235980Z digest=sha256:42427c67ab5c4edbd3518b9a7f04db61ad10c7271e02c32d18c7a350887ee4ca

Observation 91365b9e-2036-43e3-b47e-3086683cfa8c · outbound

This paper cites c 2av-tse: Context and confidence-aware audio visual target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions c 2av-tse: Context and confidence-aware audio visual target speaker extraction,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.751607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.239987Z digest=sha256:4dc3ae6890264096169d92295dde9064963e2763c57f6d906ebd4acb66488d45

Observation 0a5ad71e-687a-4064-b397-262808e4d4ae · outbound

This paper cites Incorporating linguistic constraints from external knowledge source for audio-visual target speech extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Incorporating linguistic constraints from external knowledge source for audio-visual target speech extraction,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.736790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.244151Z digest=sha256:19e0bcf749a263fa7ce5b951abf4ce14f9c2d927136aa27b339f00ce7f9e3b54

Observation 982b878b-e97d-4872-a5db-3910f7170acb · outbound

This paper cites Iianet: An intra- and inter-modality attention network for audio-visual speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Iianet: An intra- and inter-modality attention network for audio-visual speech separation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.721532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.248745Z digest=sha256:b959244cb4dbffbab80f8586328619fbcb7bd54922fde8c0472145e9fe074fe9

Observation 68eefdee-0a10-4910-844e-7f8f666b21d1 · outbound

This paper cites Hearing lips in noise: Universal viseme-phoneme mapping and transfer for robust audio-visual speech recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Hearing lips in noise: Universal viseme-phoneme mapping and transfer for robust audio-visual speech recognition,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.704696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.253262Z digest=sha256:c7f93414427c30e062c9df5595db0961299eec5a9be0acebd62b601263e7967e

Observation 5516288b-28fe-4e07-9780-6485d57fdb4e · outbound

This paper cites Diarization is hard: Some experiences and lessons learned for the jhu team in the inaugural dihard challenge.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Diarization is hard: Some experiences and lessons learned for the jhu team in the inaugural dihard challenge

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.688983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.257996Z digest=sha256:2d8c6a2fdd8abfcc06236f870384eb79fcee658550856c76058e5c7ab6dc7cea

Observation 251799e1-7c01-4669-abea-64bcb8f411b9 · outbound

This paper cites Noise- disentanglement metric learning for robust speaker verification,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Noise- disentanglement metric learning for robust speaker verification,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.672219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.263291Z digest=sha256:d13f45a4b819d695cd8903719d5f9a2bec1d777ad9b1d383a20f3e97e38d1786

Observation 9b109bc0-8f1d-4e45-9b0c-b6a4cc57e48f · outbound

This paper cites Momuse: Momentum multi-modal target speaker extraction for real-time scenarios with impaired visual cues,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Momuse: Momentum multi-modal target speaker extraction for real-time scenarios with impaired visual cues,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.656814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.267901Z digest=sha256:672333b9b5484659a06b1b073ad95b35a94ed9fab6e08b906f4b07636cc24fe7

Observation 0cb45b8b-12b0-4d2d-bc15-9b7e1d9cb579 · outbound

This paper cites Ravss: Robust audio- visual speech separation in multi-speaker scenarios with missing visual cues,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Ravss: Robust audio- visual speech separation in multi-speaker scenarios with missing visual cues,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.641107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.273119Z digest=sha256:b53db84b6f95ef9de017714da25111abe96ab078f658f4be54536485f617d749

Observation ca417b6c-9d8f-497e-817a-1d9c3321d122 · outbound

This paper cites Switching variational auto- encoders for noise-agnostic audio-visual speech enhancement,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Switching variational auto- encoders for noise-agnostic audio-visual speech enhancement,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.621867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.277199Z digest=sha256:f1c3a4b62cbce8dbabae3a9beaeccc9522f9e5e50996ebcb925359121cfb1c18

Observation d4283792-3a6a-4479-af6c-1321dd6c1cc4 · outbound

This paper cites Robust unsupervised audio-visual speech enhance- ment using a mixture of variational autoencoders,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Robust unsupervised audio-visual speech enhance- ment using a mixture of variational autoencoders,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.606055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.281465Z digest=sha256:0b74d5e52ba4fc1ebd904300d88808f5c0015aef3ddb47f9154062a83406a38c

Observation 5f16851f-04db-428a-988a-61d4a7da1a78 · outbound

This paper cites Time-domain audio-visual speech separation on low quality videos,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Time-domain audio-visual speech separation on low quality videos,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.590626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.286027Z digest=sha256:2c6e4a2951c02e6e428d8aced8f2531c5e2718965ff886cf41ea84b2b09d2eaa

Observation 952af7c7-1dfe-4740-b23e-d0bd0ab8c3a9 · outbound

This paper cites Imaginenet: Target speaker extraction with intermittent visual cue through embedding inpainting,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Imaginenet: Target speaker extraction with intermittent visual cue through embedding inpainting,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.574988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.291033Z digest=sha256:b3fd425ee9bc96d701c807a41e16ce363effb7cbca1b3ec01b18b0fd44fe9721

Observation e7dcd13f-e53e-4227-a762-1dd98aa07395 · outbound

This paper cites My lips are concealed: Audio-visual speech enhancement through obstructions,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions My lips are concealed: Audio-visual speech enhancement through obstructions,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.559734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.296351Z digest=sha256:b9b1197885f184a297bc838752254c3a9cdb1154375e3c7d1ca1b8ab63140649

Observation 6619f601-ea29-4e01-9e11-6d48e8324806 · outbound

This paper cites Multi-cue guided semi-supervised learning toward target speaker separation in real environments,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multi-cue guided semi-supervised learning toward target speaker separation in real environments,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.545050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.300822Z digest=sha256:01947020ccc930b8d151130b14d77130d8f32f25f9eedb6d979caac57b293413

Observation f4f9bccf-176a-4260-a95b-6ca2b83a4649 · outbound

This paper cites A two-stage audio-visual speech separation method without visual signals for testing and tuples loss with dynamic margin,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions A two-stage audio-visual speech separation method without visual signals for testing and tuples loss with dynamic margin,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.528996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.305414Z digest=sha256:75844aa7a1163d9a66b841e6e991d28817161dc75882c5e05e62792e0d8352f7

Observation 84bce1f1-23d6-4331-9ed0-991756082d7f · outbound

This paper cites Cross-modal speech separation without visual information during testing,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Cross-modal speech separation without visual information during testing,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.513641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.311071Z digest=sha256:36ded96951aaa386cc8b99be236abe343d034416b3ff0d8442e736b8c5d9043a

Observation d62436f8-6c39-4acd-acbb-d67d25f39ccc · outbound

This paper cites Robust audio-visual speech enhancement: Correcting misassignments in complex environments with advanced post-processing,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Robust audio-visual speech enhancement: Correcting misassignments in complex environments with advanced post-processing,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.498357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.317501Z digest=sha256:1852b7ce4d4ed69659c5c97cef336f5068b463b95241edcfae84e58b9f008ff5

Observation c459807a-0a79-4ed6-8058-0cea66615d59 · outbound

This paper cites Cocktail party listening in a dynamic multitalker environment,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Cocktail party listening in a dynamic multitalker environment,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.482734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.323438Z digest=sha256:03a0869df505851a3d30ecc595b1998cc1eb34968284f784e16dfffd98220f97

Observation e408aa7c-a002-478f-a24e-080c71069d95 · outbound

This paper cites The advantage of knowing where to listen,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions The advantage of knowing where to listen,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.467314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.329382Z digest=sha256:fde9a4078d2440ef3bf0e645b15972963f37e595bd74189a69be610b18f3b419

Observation c7f82ef1-b792-452b-a6db-e60f3923d14c · outbound

This paper cites Object continuity enhances selective auditory attention,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Object continuity enhances selective auditory attention,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.450245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.334812Z digest=sha256:f6f34abe21005f6bc15a3d47e7ba57d70dee3787725a89889d5d23c06aad589b

Observation 67e2d8c8-66db-4084-b3cc-1a7574113c4e · outbound

This paper cites Enhanced learning through multimodal training: evidence from a comprehensive cognitive, physical fitness, and neuroscience intervention,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Enhanced learning through multimodal training: evidence from a comprehensive cognitive, physical fitness, and neuroscience intervention,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.432061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.339775Z digest=sha256:4f2ad1782e93d58520bbc25e931f5dbf16cafab40265af7a2c998b472767699d

Observation c202bef3-25b1-4d8c-a92a-f2e3d8b3cbbe · outbound

This paper cites The important role of contex- tual information in speech perception in cochlear implant users and its consequences in speech tests,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions The important role of contex- tual information in speech perception in cochlear implant users and its consequences in speech tests,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.414053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.344866Z digest=sha256:b3bafad40614331e662955ddde8cce286f838ae90f1c5e4f8d72a0fec5646f01

Observation 92d5150e-1fa1-4fa2-8fe1-bc72d4bc2acf · outbound

This paper cites Attention and working memory in human auditory cortex,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Attention and working memory in human auditory cortex,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.395278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.350053Z digest=sha256:34359c0d698fd38681f54200bc818113bfa07afdc2018a7b0f378458166d8126

Observation 877109a9-6ee5-418f-aefb-59121d7b42b2 · outbound

This paper cites Interactions between attention and working memory,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Interactions between attention and working memory,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.377778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.354448Z digest=sha256:3fa78a33935897f0eb02729c76c2f80664a79a99b485cf7f4e1387eda83fca5b

Observation f1508676-d493-4fde-b68c-adfcae7d7c55 · outbound

This paper cites Look once to hear: Target speech hearing with noisy examples,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Look once to hear: Target speech hearing with noisy examples,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.361411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.358543Z digest=sha256:cc53daee1998d9aa83869ec9c5d25348cfda73039f07bbfb2ffb0dcd4a31e7ff

Observation e077f5c0-d99c-48aa-9f92-fb7421d248c6 · outbound

This paper cites Multimodal attention fusion for target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multimodal attention fusion for target speaker extraction,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.362651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.362651Z digest=sha256:c1e80be7ce8e53de8050adcf062a0b231f66b6704213c93936b77572ab45969c

Observation d6d0b610-dd7f-440e-bebd-a6e46823b311 · outbound

This paper cites Multimodal speakerbeam: Single channel target speech extraction with audio-visual speaker clues.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multimodal speakerbeam: Single channel target speech extraction with audio-visual speaker clues

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.332485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.366891Z digest=sha256:62a4e89b0a6a2a135383b56fc2c2145b4800d203ad8c86ff87d82de1b86fa41d

Observation 5987cf6c-44af-40fa-a653-65ff48ffcdec · outbound

This paper cites Memory networks,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Memory networks,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.311379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.371554Z digest=sha256:316d4773163a8f446cbeccde9c50c4d6141cdf946b7526318477d010ba5dd644

Observation 3b1f9cce-ca92-414a-97b1-3694b4dd8106 · outbound

This paper cites End-to-end memory net- works,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions End-to-end memory net- works,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.289641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.377652Z digest=sha256:8c89c7742ee74e7ccc7d8529119ed4a7ff5f000239f6416ea57555fd6629b731

Observation c9af338d-836f-4736-ac78-904d87b15541 · outbound

This paper cites Unsupervised feature learning via non-parametric instance discrimination,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Unsupervised feature learning via non-parametric instance discrimination,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.268598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.382343Z digest=sha256:511b379a9584c2cfd9c821099b99864aef1c8eb4b064ef7eedf17518a016d15a

Observation 91992967-d5e3-4a92-b2e2-d45a6a180914 · outbound

This paper cites Cromm-vsr: Cross-modal memory augmented visual speech recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Cromm-vsr: Cross-modal memory augmented visual speech recognition,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.250036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.386674Z digest=sha256:6b1a9f36c035c3aeb80bd4680e95d3d72094bc5bcd122d1ab400bb551bf549d3

Observation f400ee2c-ab6b-4fc8-8a4a-bb3348d44d10 · outbound

This paper cites Multi-temporal lip-audio memory for visual speech recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multi-temporal lip-audio memory for visual speech recognition,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.235219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.391493Z digest=sha256:c5ccaa7ed88c6ef638f74a98216b93d182cdcc04b5bb712aae5e34dfe973b1a5

Observation 3d307f7b-3734-4902-b509-f0f6e5b525b0 · outbound

This paper cites Multi-modality associative bridging through memory: Speech sound recollected from face video,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multi-modality associative bridging through memory: Speech sound recollected from face video,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.220095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.395575Z digest=sha256:7171ed461487758f6ed0ffeb49c9718425f23aba1ac6c10de0032fefc7150d16

Observation b7540409-2519-41ea-b221-7034dc95db8b · outbound

This paper cites Akvsr: Audio knowledge empowered visual speech recognition by compressing audio knowledge of a pretrained model,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Akvsr: Audio knowledge empowered visual speech recognition by compressing audio knowledge of a pretrained model,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.202608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.400260Z digest=sha256:b2dfdf929708b81e44f9b95421345ab9545546ee1e434021d3ab107d806af3a8

Observation 62dbf9dc-1e99-402c-a311-d97e8e29a2b3 · outbound

This paper cites Distinguishing homophenes using multi-head visual-audio memory for lip reading,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Distinguishing homophenes using multi-head visual-audio memory for lip reading,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.184305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.405373Z digest=sha256:e57456278f0dccea2379ef7d2996c1dc850afe4a9bae4976a138da28a1c0fa61

Observation 8b8a75df-6b28-4901-82c0-914c3a569f20 · outbound

This paper cites Speech reconstruction with reminiscent sound via visual voice memory,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Speech reconstruction with reminiscent sound via visual voice memory,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.166533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.409884Z digest=sha256:cd9b7be905066515a366bb4bf133732f37dbeddaf0e8bd564f99ea139a74e8be

Observation aec87788-c5ef-4974-a09f-11d8b2aa3a8e · outbound

This paper cites Modeling attention and memory for auditory selection in a cocktail party environment,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Modeling attention and memory for auditory selection in a cocktail party environment,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.151632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.414371Z digest=sha256:e4bbe430b3e35b16e57ba8ce81fd3efbcf3d4411270e3d285edbdf4138f35b5a

Observation 4057b1ff-ab00-44c2-a304-2d26fb5eb04f · outbound

This paper cites Explicit-memory multiresolution adaptive framework for speech and music separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Explicit-memory multiresolution adaptive framework for speech and music separation,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.135206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.419236Z digest=sha256:ab9c8545dc639547c4c141f2ec86a67981846b7f052a3e9dec0d6f9bc2904ff8

Observation f8dedcdf-0717-4d31-ace6-d546dc222950 · outbound

This paper cites On the effectiveness of enrollment speech augmentation for target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions On the effectiveness of enrollment speech augmentation for target speaker extraction,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.118919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.423757Z digest=sha256:554f9d5be5e86f3df2fa1b2c2b090c6877cb032f8a3678d2ba1738bf4c3339d8

Observation 09f89054-bda5-48b8-8c9f-f5e021deb3c4 · outbound

This paper cites Selective listening by synchronizing speech with lips,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Selective listening by synchronizing speech with lips,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.102594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.428367Z digest=sha256:f8490d960dcb72f0ac92d7e00ddaba82fe557ae81132efcca4558db279ff968c

Observation 02d088bf-9e49-468c-b71e-80c73883c938 · outbound

This paper cites Robust Speaker Extraction Network Based on Iterative Refined Adap- tation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Robust Speaker Extraction Network Based on Iterative Refined Adap- tation,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.086812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.432582Z digest=sha256:28d81ea48026ad4415614f98fde882983246e037b81600537144af3e3d028077

Observation 7752d5ec-a8d0-4ad4-b71a-94b4668935d2 · outbound

This paper cites Listening and group- ing: an online autoregressive approach for monaural speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Listening and group- ing: an online autoregressive approach for monaural speech separation,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.070125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.437077Z digest=sha256:4c374179ef6fa671b0304466119edae85377022448b96b2e965ea5f4021f722f

Observation ec8da356-41c7-4aa4-ad74-9509c544109d · outbound

This paper cites Source-aware context network for single-channel multi-speaker speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Source-aware context network for single-channel multi-speaker speech separation,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.053802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.441118Z digest=sha256:4ac61d5d160a1e41df1be35b63c155211667ce1581592312b3290bc9949e8a35

Observation 7b63a849-45b7-4346-b50a-5dfc98bd0baa · outbound

This paper cites An online speaker-aware speech separation approach based on time-domain representation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions An online speaker-aware speech separation approach based on time-domain representation,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.037326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.445488Z digest=sha256:e5a1e268ea4386c66686d21f391e5ae139ae11be6f10cc1981a9363eb24f2945

Observation c75e5f9b-344c-48e1-a81c-cbb23b9a1946 · outbound

This paper cites Iterative autoregression: a novel trick to improve your low-latency speech enhancement model,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Iterative autoregression: a novel trick to improve your low-latency speech enhancement model,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.020906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.450282Z digest=sha256:54a1c5cb86b4a2f539f182d9e0430b24ecbe0c3333bcff8020cf6ab01895176c

Observation f3f3203c-8165-4df5-8d2f-c873e09cbeab · outbound

This paper cites Neuroheed: Neuro- steered speaker extraction using eeg signals,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Neuroheed: Neuro- steered speaker extraction using eeg signals,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.004328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.454651Z digest=sha256:3e32dfa815ab82a11807850e0727d264686de873153060d9d6a200c141ae4c3a

Observation e6b96c13-e4cd-4786-a305-f25608f9ce1e · outbound

This paper cites Neuroheed+: Improving neuro-steered speaker extraction with joint auditory attention detection,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Neuroheed+: Improving neuro-steered speaker extraction with joint auditory attention detection,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.985128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.458874Z digest=sha256:6d869642f8e3ecaabe880f8d2dff62def116068991f154c6557f5752386e4b17

Observation 0f2bcac3-be4f-490a-ab12-cecd3933a5e9 · outbound

This paper cites Paris: Pseudo-autoregressive siamese training for online speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Paris: Pseudo-autoregressive siamese training for online speech separation,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.968849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.463998Z digest=sha256:30ae9669674f29292a8ddf7bfe9cd3a98a99b1003e8f443eb1d637c150512618

Observation 900d7c57-838a-4484-a067-9c545a703533 · outbound

This paper cites On- line Audio-Visual Autoregressive Speaker Extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions On- line Audio-Visual Autoregressive Speaker Extraction,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.949809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.469257Z digest=sha256:3329b7bb7a7fb246e15d81f65d5defeac39ee08e9e47f3f62613ee3a9a1a6e69

Observation d6aa6f40-20ba-4b58-94f3-1292d36b9aab · outbound

This paper cites Attention is all you need,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Attention is all you need,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.473959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.473959Z digest=sha256:b796661fe57f6378cd2bc753e6468100ca3b583ee1d1d673306965f5db46aad0

Observation 81286334-2444-45bf-baeb-f40273a994ad · outbound

This paper cites Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.916620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.478541Z digest=sha256:b334c5a8ef384702ccbaa6e5d92e860a5e5131bb6882233a770d8512ce7634ba

Observation f360c19f-93a5-4926-8a80-b8ebe629387f · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Wespeaker: A research and production oriented speaker embedding learning toolkit,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.897397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.483332Z digest=sha256:b3916d2e764ff9e5df5de835e6b18000dc2ddfaaaa145f699f86be89dbfcf9c1

Observation a5e18d5e-3975-41f5-83d8-e7782d961065 · outbound

This paper cites Curriculum learning,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Curriculum learning,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.488014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.488014Z digest=sha256:778f5c68437cd1f0095046371d362b3f900603a933a3e4952eeb6d8737b906aa

Observation c59cc09d-4efc-4930-8960-00e534280cab · outbound

This paper cites A learning algorithm for continually running fully recurrent neural networks,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions A learning algorithm for continually running fully recurrent neural networks,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.492392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.492392Z digest=sha256:9e6f6a2a7496379ff949e9b10c9fafeaee7840338023815564ccb66e83482e29

Observation 3b0191e3-e40d-42e9-88a2-1557279844d5 · outbound

This paper cites Scheduled sampling for sequence prediction with recurrent neural networks,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Scheduled sampling for sequence prediction with recurrent neural networks,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.858309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.497009Z digest=sha256:909e7be41ab6108a008f16b05f70dfcc1b8d70490f35aaadc8e039ce54aa7958

Observation 17078865-8ccd-4862-b091-1f5803386d60 · outbound

This paper cites Delving into high-quality syn- thetic face occlusion segmentation datasets,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Delving into high-quality syn- thetic face occlusion segmentation datasets,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.842492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.502291Z digest=sha256:642bb624aa4a26c0b3b8994570d66fc49e5c8fd49237ddfdef9d0490eea48111

Observation ac2c226e-f775-4761-9c08-30a9a67f8a2a · outbound

This paper cites Watch or listen: Robust audio-visual speech recognition with visual corruption modeling and reliability scoring,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Watch or listen: Robust audio-visual speech recognition with visual corruption modeling and reliability scoring,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.825636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.506931Z digest=sha256:7ad6fbf76d09c5c99855e9d1fb9695b711e3c29952335182fc68a72731f521d1

Observation a38afe9a-a0f5-4d0f-b894-fb657b1d27c7 · outbound

This paper cites V oxceleb2: Deep speaker recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions V oxceleb2: Deep speaker recognition,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.806766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.511270Z digest=sha256:f2d459163c4250c88464e246423fa26ceb1fe72614f560d03dc0662e16d75c56

Observation 60b84ecc-6bd7-4797-a666-4cc5be932f54 · outbound

This paper cites Avhumar: Audio-visual target speech extraction with pre-trained av-hubert and mask-and-recover strategy,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Avhumar: Audio-visual target speech extraction with pre-trained av-hubert and mask-and-recover strategy,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.787987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.515654Z digest=sha256:0e9636487564e86e530af4af6d84d850d20e74b873a68af718fa2ccb6856fca5

Observation 7e5d0481-f3bb-42df-8a8c-4fe34c912811 · outbound

This paper cites Target speech extraction with pre-trained av-hubert and mask-and-recover strategy,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Target speech extraction with pre-trained av-hubert and mask-and-recover strategy,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.766513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.521222Z digest=sha256:55680a0a54dad79fcf9831b6cac94cabe4f5ba5e00efde007b5da6b1941cd06d

Observation 89aae3a2-1857-492b-bb66-7adf0a6cd7da · outbound

This paper cites Sdr–half-baked or well done?.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Sdr–half-baked or well done?

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.749611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.525852Z digest=sha256:13d1d1939a41bbd145fc6530235fb0717e1bfe40de7ed46e1601fffb7ee48ec7

Observation 03e57a72-d5f5-46cf-8ff2-e8a5e8c6791b · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.733522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.530574Z digest=sha256:42c7a617ff4351baaed5ded901ae7715c9b5d99c9febac5d8868e19f85b23d8f

Observation c9560885-5aa4-4169-8ffe-8241c468c3dd · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions A short- time objective intelligibility measure for time-frequency weighted noisy speech,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.713313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.535064Z digest=sha256:81ded0adb0275ac5120c2a99e6f87bd67711ac0f0cfe37543ee784a0ec6c3d78

Observation f1c2c8e9-7380-4dc0-aab9-0c7cabd69ed8 · outbound

This paper cites Time domain audio visual speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Time domain audio visual speech separation,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.693485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.540459Z digest=sha256:326dd7f0420c6ecb79d1a90b7daf25e7ff973386b76b55b6f37b2dc60a138b47

Observation 7cf1d4a5-6276-456f-9ad5-bdb419330ec7 · outbound

This paper cites Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.545087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.545087Z digest=sha256:7c403abb25e6dff24aab725e1f05ca071548487f616c4a0230224db221aa99d6

Observation 4453c919-dd08-4af9-a213-2da60ed02d89 · outbound

This paper cites Music source separation with band-split rnn,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Music source separation with band-split rnn,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.549514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.549514Z digest=sha256:ad90c26c203fe63270309556db689d3868d8ffa7e025833bd20a4f1c0fabee7c

Observation 88f831c1-baf9-4a3f-9125-e6d48eb7bc4d · outbound

This paper cites Wesep: A scalable and flexible toolkit towards generalizable target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Wesep: A scalable and flexible toolkit towards generalizable target speaker extraction,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.648914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.554592Z digest=sha256:7c34143d067c27b9a2f5db66007eb0726daa6f5c8cf2cdcd65022e3b62a629de

Observation 8e34c688-aeea-432b-b1fc-c5cc47cd47d5 · outbound

This paper cites Audio-visual target speaker extraction with selective auditory attention,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Audio-visual target speaker extraction with selective auditory attention,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.631585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T15:40:09.559373Z digest=sha256:67a733aedc545f6757ac6b5461372601935a0ef908827fa14588b73e4acc648d

Observation aa5cd0ee-cf22-4bee-8798-2d7dd572e36e · outbound

This paper cites Attention is all you need in speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Attention is all you need in speech separation,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.564250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.564250Z digest=sha256:822520674875a5e91ae9ecbaeca4b304644137dd9dca994bbafaa0cba2608c74

Pith citing papers

Observation a093622f-23ed-4eee-b9d4-df82da779cc2 · inbound

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction cites this paper.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.277622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.277622Z digest=sha256:50299e0eb0ccc8a39551508a62bc86e30c4d80a3e980ea7e9075292c7a8734f0