Pith. sign in

Paper Citation Record · LEDGER

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions

As of 8 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 1 inbound Pith citation observation for arXiv:2507.15294.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15294 v2

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:40:09.564250Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T19:39:50.277622Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact0
  • verified fuzzy72
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 83c6922e-33a6-4d1b-9c21-65841d11fc41 · outbound

This paper cites Some experiments on the recognition of speech, with one and with two ears,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Some experiments on the recognition of speech, with one and with two ears,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.175918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.175918Z digest=sha256:1d47d18ccd7fe95d469606fbec0c9983cbdee2dae4a9d205d33133fd62770ea2

Observation 744f2f48-ebd7-4a45-8006-d2f53eb17f8b · outbound

This paper cites The cocktail party phenomenon: A review of research on speech intelligibility in multiple-talker conditions,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions The cocktail party phenomenon: A review of research on speech intelligibility in multiple-talker conditions,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.181286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.181286Z digest=sha256:6425496fe6635776db08ca49f9ce34a698ae52c197d7696ee44c385b19444793

Observation 8b3c2740-4d0c-4fe9-9692-f331dc2aa051 · outbound

This paper cites Seeing to hear better: evidence for early audio-visual interactions in speech identification,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Seeing to hear better: evidence for early audio-visual interactions in speech identification,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.949415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.186675Z digest=sha256:c9cfd14715ca0450c2ccd639b06a59c115930e2a5e3e53ae894e5039bedeb7d9

Observation 95ec1176-f5ff-4daf-aed3-5b0f40666070 · outbound

This paper cites Visual contribution to speech intelligibility in noise,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Visual contribution to speech intelligibility in noise,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.932674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.191605Z digest=sha256:cd5fed987b3bd4babd6147ea525df1a48f0fae6e848b36242b35d672354d4c5c

Observation 4ff5073d-3358-4e53-8caa-ed50b21744a5 · outbound

This paper cites Muse: Multi-modal target speaker extraction with visual cues,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Muse: Multi-modal target speaker extraction with visual cues,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.915156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.196835Z digest=sha256:0b33ea1c4e54e9394bec43e83fdc203e6b761d808ded07f86af3adc47acf40aa

Observation 2763e41b-075b-4632-925b-f5e4921a5a8c · outbound

This paper cites Usev: Universal speaker extraction with visual cue,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Usev: Universal speaker extraction with visual cue,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.885618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.201540Z digest=sha256:9791dab814a0eacfbce4339c04e2bc9e8e476aa1e5608d8a1f01ce999965a6ee

Observation d1c517a4-72d4-4dbd-b3a0-a2c100695cab · outbound

This paper cites Time domain audio visual speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Time domain audio visual speech separation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.866995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.207358Z digest=sha256:738fcdd59d05694a9c3c2d272cbcae9664ee6db0544976896d4e4d36e0bc8117

Observation 094b854b-5973-4c6b-a35e-cf83b49c2596 · outbound

This paper cites Neural target speech extraction: An overview,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Neural target speech extraction: An overview,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.212164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.212164Z digest=sha256:0dce6467ef7b5a86c1ed52a6ba1ec27da2afc2aaa23e5007ee04a29cd125806b

Observation a76edd7b-75fe-4c81-b09d-a2a30fc5197c · outbound

This paper cites An overview of deep-learning-based audio-visual speech en- hancement and separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions An overview of deep-learning-based audio-visual speech en- hancement and separation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.841856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.217781Z digest=sha256:22cf95314d0c6b902392539515c9e7f4fe46b36d2673feaf156797131e22b515

Observation 1a4b2d8f-8080-4e8e-a875-1c91d982f1bd · outbound

This paper cites PIA VE: A Pose-Invariant Audio- Visual Speaker Extraction Network,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions PIA VE: A Pose-Invariant Audio- Visual Speaker Extraction Network,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.826667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.222557Z digest=sha256:07d371e8ff7cb0b3e43de1885313e7b941daf30b414b1c46b797982f1d1df944

Observation 7d13eadd-9117-4305-83c8-634b94bccd40 · outbound

This paper cites Speaker extraction with co-speech gestures cue,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Speaker extraction with co-speech gestures cue,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.809836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.227254Z digest=sha256:7496af49378ceb75249537d227f6d092be8691ea3803de24812f2eda45906054

Observation 0f02fc49-23af-4400-a7a3-c539854bd177 · outbound

This paper cites Rethinking the Visual Cues in Audio-Visual Speaker Extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Rethinking the Visual Cues in Audio-Visual Speaker Extraction,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.792840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.231505Z digest=sha256:ddce8ea18cb1a7dab84c232587aa549f53310f45726394132b6312a1845b2b7a

Observation 302839d1-bece-43a4-8150-2868d3c52695 · outbound

This paper cites FaceFilter: Audio- Visual Speech Separation Using Still Images,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions FaceFilter: Audio- Visual Speech Separation Using Still Images,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.767074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.235980Z digest=sha256:1377c2e5afd3c0f862818cab671cc89fbd316ce5c4f9798528c395c21c3206cf

Observation 91365b9e-2036-43e3-b47e-3086683cfa8c · outbound

This paper cites c 2av-tse: Context and confidence-aware audio visual target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions c 2av-tse: Context and confidence-aware audio visual target speaker extraction,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.751607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.239987Z digest=sha256:18fcb341366ab745be0bf05d9b072a27e0f205765fba29ae829b9045cf062ee1

Observation 0a5ad71e-687a-4064-b397-262808e4d4ae · outbound

This paper cites Incorporating linguistic constraints from external knowledge source for audio-visual target speech extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Incorporating linguistic constraints from external knowledge source for audio-visual target speech extraction,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.736790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.244151Z digest=sha256:318d2a20221683fc5c2d5c2f801800c39194505ddc105dbe1969b72ace5a9e78

Observation 982b878b-e97d-4872-a5db-3910f7170acb · outbound

This paper cites Iianet: An intra- and inter-modality attention network for audio-visual speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Iianet: An intra- and inter-modality attention network for audio-visual speech separation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.721532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.248745Z digest=sha256:e5cdbc2e8bda05a77819ebb4f40a4aff20d0df4269ea27a9eafad6f07f49389e

Observation 68eefdee-0a10-4910-844e-7f8f666b21d1 · outbound

This paper cites Hearing lips in noise: Universal viseme-phoneme mapping and transfer for robust audio-visual speech recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Hearing lips in noise: Universal viseme-phoneme mapping and transfer for robust audio-visual speech recognition,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.704696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.253262Z digest=sha256:d9338d14df7a3ac111605176bc707baa47d782956ddcd811bfade11f05c5cc37

Observation 5516288b-28fe-4e07-9780-6485d57fdb4e · outbound

This paper cites Diarization is hard: Some experiences and lessons learned for the jhu team in the inaugural dihard challenge.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Diarization is hard: Some experiences and lessons learned for the jhu team in the inaugural dihard challenge

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.688983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.257996Z digest=sha256:ab7586656b12eaa52b2fc5af325ba66e8fc66744cfb988c6e2c94c8f5c4b7bd9

Observation 251799e1-7c01-4669-abea-64bcb8f411b9 · outbound

This paper cites Noise- disentanglement metric learning for robust speaker verification,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Noise- disentanglement metric learning for robust speaker verification,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.672219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.263291Z digest=sha256:6e21c3e1c0c50399bdb70f2448290b5a5d9158fa09c230217c4c999e763645a7

Observation 9b109bc0-8f1d-4e45-9b0c-b6a4cc57e48f · outbound

This paper cites Momuse: Momentum multi-modal target speaker extraction for real-time scenarios with impaired visual cues,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Momuse: Momentum multi-modal target speaker extraction for real-time scenarios with impaired visual cues,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.656814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.267901Z digest=sha256:5993e410479c5bc920370e272572a53769db1ec7cc08776fc10f249e36952eee

Observation 0cb45b8b-12b0-4d2d-bc15-9b7e1d9cb579 · outbound

This paper cites Ravss: Robust audio- visual speech separation in multi-speaker scenarios with missing visual cues,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Ravss: Robust audio- visual speech separation in multi-speaker scenarios with missing visual cues,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.641107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.273119Z digest=sha256:069ff903d37badca4ec9d73cd14d6ef98edfa1b59b0d0446eb9bb449ee5a9ce8

Observation ca417b6c-9d8f-497e-817a-1d9c3321d122 · outbound

This paper cites Switching variational auto- encoders for noise-agnostic audio-visual speech enhancement,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Switching variational auto- encoders for noise-agnostic audio-visual speech enhancement,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.621867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.277199Z digest=sha256:30f6038c3f1c87c0dcbfd8136040fc7d2d95725c621f3ca80cc6c96ce9c240c5

Observation d4283792-3a6a-4479-af6c-1321dd6c1cc4 · outbound

This paper cites Robust unsupervised audio-visual speech enhance- ment using a mixture of variational autoencoders,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Robust unsupervised audio-visual speech enhance- ment using a mixture of variational autoencoders,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.606055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.281465Z digest=sha256:3657392ed433d099b7760be52716c32cf264d5d64d6a4dc99273edec0d382516

Observation 5f16851f-04db-428a-988a-61d4a7da1a78 · outbound

This paper cites Time-domain audio-visual speech separation on low quality videos,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Time-domain audio-visual speech separation on low quality videos,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.590626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.286027Z digest=sha256:7d1b05c56dea6f6d384ac79974487138c4040b70c80a419617b4460f8ec1648e

Observation 952af7c7-1dfe-4740-b23e-d0bd0ab8c3a9 · outbound

This paper cites Imaginenet: Target speaker extraction with intermittent visual cue through embedding inpainting,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Imaginenet: Target speaker extraction with intermittent visual cue through embedding inpainting,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.574988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.291033Z digest=sha256:63449b1c019527f4719d7d85608fa210ea42f67603ce5689bbc7d444329687b2

Observation e7dcd13f-e53e-4227-a762-1dd98aa07395 · outbound

This paper cites My lips are concealed: Audio-visual speech enhancement through obstructions,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions My lips are concealed: Audio-visual speech enhancement through obstructions,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.559734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.296351Z digest=sha256:ee4e5b88c8c74ebf8f79cb007771d555128dcffc6f1f4c6f63419695444ff3ed

Observation 6619f601-ea29-4e01-9e11-6d48e8324806 · outbound

This paper cites Multi-cue guided semi-supervised learning toward target speaker separation in real environments,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multi-cue guided semi-supervised learning toward target speaker separation in real environments,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.545050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.300822Z digest=sha256:c0f59922a1f3ed17d2f96da67b7f1d5a06a08210d59e9eb577def41b75a06f76

Observation f4f9bccf-176a-4260-a95b-6ca2b83a4649 · outbound

This paper cites A two-stage audio-visual speech separation method without visual signals for testing and tuples loss with dynamic margin,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions A two-stage audio-visual speech separation method without visual signals for testing and tuples loss with dynamic margin,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.528996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.305414Z digest=sha256:ae23e466814111be9f6bd40935f1c2d68b32eb690361e18b7cf82e205daa8bfb

Observation 84bce1f1-23d6-4331-9ed0-991756082d7f · outbound

This paper cites Cross-modal speech separation without visual information during testing,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Cross-modal speech separation without visual information during testing,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.513641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.311071Z digest=sha256:3f07ac2c44ac707a9f9f48ed531dd6759289f3f106170b63c49ac8affb8893d7

Observation d62436f8-6c39-4acd-acbb-d67d25f39ccc · outbound

This paper cites Robust audio-visual speech enhancement: Correcting misassignments in complex environments with advanced post-processing,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Robust audio-visual speech enhancement: Correcting misassignments in complex environments with advanced post-processing,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.498357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.317501Z digest=sha256:fd975d3753df1140858e56470b90c0c98c58f62041c821ede9659a0988797c80

Observation c459807a-0a79-4ed6-8058-0cea66615d59 · outbound

This paper cites Cocktail party listening in a dynamic multitalker environment,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Cocktail party listening in a dynamic multitalker environment,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.482734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.323438Z digest=sha256:525d9df6c41d83ece2e1e10a4539d77e09f87383aa0e61663bebcfc066d2a990

Observation e408aa7c-a002-478f-a24e-080c71069d95 · outbound

This paper cites The advantage of knowing where to listen,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions The advantage of knowing where to listen,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.467314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.329382Z digest=sha256:9d39814df9c88ef226566f207e382c348affee86bd43def40708868a2d9e5e60

Observation c7f82ef1-b792-452b-a6db-e60f3923d14c · outbound

This paper cites Object continuity enhances selective auditory attention,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Object continuity enhances selective auditory attention,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.450245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.334812Z digest=sha256:8e83151b0b71b6613915e26dd48dfcbcfdb3b4b9d70e72dd251f19ad555f4d8c

Observation 67e2d8c8-66db-4084-b3cc-1a7574113c4e · outbound

This paper cites Enhanced learning through multimodal training: evidence from a comprehensive cognitive, physical fitness, and neuroscience intervention,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Enhanced learning through multimodal training: evidence from a comprehensive cognitive, physical fitness, and neuroscience intervention,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.432061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.339775Z digest=sha256:9267c2ed47ab92fd1b52c79d70fdaea5451c486d14b1717270e20560b74040f6

Observation c202bef3-25b1-4d8c-a92a-f2e3d8b3cbbe · outbound

This paper cites The important role of contex- tual information in speech perception in cochlear implant users and its consequences in speech tests,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions The important role of contex- tual information in speech perception in cochlear implant users and its consequences in speech tests,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.414053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.344866Z digest=sha256:dcc31ea00daef817355845a322cb8d35adbc3314ad544e06aae007c7cfb94971

Observation 92d5150e-1fa1-4fa2-8fe1-bc72d4bc2acf · outbound

This paper cites Attention and working memory in human auditory cortex,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Attention and working memory in human auditory cortex,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.395278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.350053Z digest=sha256:435dc3cfffddd1ed6308a4465754edb9670aaee1ab3357f8954aaace5c95c762

Observation 877109a9-6ee5-418f-aefb-59121d7b42b2 · outbound

This paper cites Interactions between attention and working memory,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Interactions between attention and working memory,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.377778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.354448Z digest=sha256:5db4aa589c22ab115f67c26d08f8e1c6aa11dc20733fa7d068839a05a27b2711

Observation f1508676-d493-4fde-b68c-adfcae7d7c55 · outbound

This paper cites Look once to hear: Target speech hearing with noisy examples,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Look once to hear: Target speech hearing with noisy examples,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.361411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.358543Z digest=sha256:f52305e04c164d0c37692e846a8ced3caeb732d085320587fdf80e417f131bbc

Observation e077f5c0-d99c-48aa-9f92-fb7421d248c6 · outbound

This paper cites Multimodal attention fusion for target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multimodal attention fusion for target speaker extraction,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.362651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.362651Z digest=sha256:7d3d01b55049b6bfef152bbf35215828e251cc821cca531ed2d9bd64be527126

Observation d6d0b610-dd7f-440e-bebd-a6e46823b311 · outbound

This paper cites Multimodal speakerbeam: Single channel target speech extraction with audio-visual speaker clues.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multimodal speakerbeam: Single channel target speech extraction with audio-visual speaker clues

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.332485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.366891Z digest=sha256:307b135634007c856670f6db366b19ed2785656da48f97f717cc05e7f8b026f2

Observation 5987cf6c-44af-40fa-a653-65ff48ffcdec · outbound

This paper cites Memory networks,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Memory networks,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.311379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.371554Z digest=sha256:25f61854b3f5d877ed86d4544523139b937cbc5f51de0bff48f3b0e19329ac7d

Observation 3b1f9cce-ca92-414a-97b1-3694b4dd8106 · outbound

This paper cites End-to-end memory net- works,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions End-to-end memory net- works,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.289641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.377652Z digest=sha256:ac3a98dad590149a29e8112a91045ff72957e5897329bf1d5025e8c2bd1e0b03

Observation c9af338d-836f-4736-ac78-904d87b15541 · outbound

This paper cites Unsupervised feature learning via non-parametric instance discrimination,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Unsupervised feature learning via non-parametric instance discrimination,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.268598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.382343Z digest=sha256:650216f6619d30142bc6d8fbf2be189b884258c589e9b0da34a4003c4fc06eb6

Observation 91992967-d5e3-4a92-b2e2-d45a6a180914 · outbound

This paper cites Cromm-vsr: Cross-modal memory augmented visual speech recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Cromm-vsr: Cross-modal memory augmented visual speech recognition,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.250036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.386674Z digest=sha256:d3b1d67a57a7c2d08c15ca6cbc1bb4223abf96d6da206b2c0f1481b225571fed

Observation f400ee2c-ab6b-4fc8-8a4a-bb3348d44d10 · outbound

This paper cites Multi-temporal lip-audio memory for visual speech recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multi-temporal lip-audio memory for visual speech recognition,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.235219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.391493Z digest=sha256:fbfa5578644e0b860f132d94a88fb2cdc96b0f01d83a1ea407aa2a922205081e

Observation 3d307f7b-3734-4902-b509-f0f6e5b525b0 · outbound

This paper cites Multi-modality associative bridging through memory: Speech sound recollected from face video,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multi-modality associative bridging through memory: Speech sound recollected from face video,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.220095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.395575Z digest=sha256:2175ad5bd7f88028e364841e0dcda49516e84b26cadd01d154c6764f1d5ea513

Observation b7540409-2519-41ea-b221-7034dc95db8b · outbound

This paper cites Akvsr: Audio knowledge empowered visual speech recognition by compressing audio knowledge of a pretrained model,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Akvsr: Audio knowledge empowered visual speech recognition by compressing audio knowledge of a pretrained model,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.202608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.400260Z digest=sha256:60dfeb8c7212e6bdc0c8a91cd6469e8a39709d00da90c4b75fe072f56cf0d08f

Observation 62dbf9dc-1e99-402c-a311-d97e8e29a2b3 · outbound

This paper cites Distinguishing homophenes using multi-head visual-audio memory for lip reading,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Distinguishing homophenes using multi-head visual-audio memory for lip reading,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.184305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.405373Z digest=sha256:6edb261eca5c14c3297981add0f2064d1d32575047f3c658e0ec0058f1df16be

Observation 8b8a75df-6b28-4901-82c0-914c3a569f20 · outbound

This paper cites Speech reconstruction with reminiscent sound via visual voice memory,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Speech reconstruction with reminiscent sound via visual voice memory,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.166533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.409884Z digest=sha256:b6de0ed595a143b75e91bdd214ce896ef03a16c7fc587f63777cb049358ef800

Observation aec87788-c5ef-4974-a09f-11d8b2aa3a8e · outbound

This paper cites Modeling attention and memory for auditory selection in a cocktail party environment,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Modeling attention and memory for auditory selection in a cocktail party environment,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.151632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.414371Z digest=sha256:f96ca7d3646f1e2eb6003f55b63fbb09ac02e145364240c63ea09f430ff85312

Observation 4057b1ff-ab00-44c2-a304-2d26fb5eb04f · outbound

This paper cites Explicit-memory multiresolution adaptive framework for speech and music separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Explicit-memory multiresolution adaptive framework for speech and music separation,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.135206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.419236Z digest=sha256:dc8707d15d4d8bd4a1df483e70aa47b71b510fd39e64b88750f9f1771a125f93

Observation f8dedcdf-0717-4d31-ace6-d546dc222950 · outbound

This paper cites On the effectiveness of enrollment speech augmentation for target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions On the effectiveness of enrollment speech augmentation for target speaker extraction,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.118919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.423757Z digest=sha256:c835e6b927b0e64d2b420e04a56d9dfefbb405849101991f4586adb63daa2238

Observation 09f89054-bda5-48b8-8c9f-f5e021deb3c4 · outbound

This paper cites Selective listening by synchronizing speech with lips,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Selective listening by synchronizing speech with lips,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.102594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.428367Z digest=sha256:1b3bbe7f112af28a4d31394c452197c51b64621110b3b16099a2dd93564510be

Observation 02d088bf-9e49-468c-b71e-80c73883c938 · outbound

This paper cites Robust Speaker Extraction Network Based on Iterative Refined Adap- tation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Robust Speaker Extraction Network Based on Iterative Refined Adap- tation,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.086812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.432582Z digest=sha256:05befe0fef68b8a7ac9ebfa141139700a8bf1392bdbd34415651e520c2c0df78

Observation 7752d5ec-a8d0-4ad4-b71a-94b4668935d2 · outbound

This paper cites Listening and group- ing: an online autoregressive approach for monaural speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Listening and group- ing: an online autoregressive approach for monaural speech separation,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.070125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.437077Z digest=sha256:90f063b34a4c0c226adb075d9716eab155cf5a7120fcaa77c9d047930e36b837

Observation ec8da356-41c7-4aa4-ad74-9509c544109d · outbound

This paper cites Source-aware context network for single-channel multi-speaker speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Source-aware context network for single-channel multi-speaker speech separation,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.053802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.441118Z digest=sha256:42d6c5b3210ba17a247be72183cc0fe2d7db07aa7f7313231e0dd7d38aa5c604

Observation 7b63a849-45b7-4346-b50a-5dfc98bd0baa · outbound

This paper cites An online speaker-aware speech separation approach based on time-domain representation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions An online speaker-aware speech separation approach based on time-domain representation,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.037326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.445488Z digest=sha256:cf3928810f6df8cdbbf5338ee77f525334caf44498a6be9056170bd7132cad90

Observation c75e5f9b-344c-48e1-a81c-cbb23b9a1946 · outbound

This paper cites Iterative autoregression: a novel trick to improve your low-latency speech enhancement model,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Iterative autoregression: a novel trick to improve your low-latency speech enhancement model,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.020906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.450282Z digest=sha256:2df50b9deb66455620db49be7afa198a36cf7f79226062ca40b20cf857293a2b

Observation f3f3203c-8165-4df5-8d2f-c873e09cbeab · outbound

This paper cites Neuroheed: Neuro- steered speaker extraction using eeg signals,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Neuroheed: Neuro- steered speaker extraction using eeg signals,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.004328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.454651Z digest=sha256:c3a177e7675346f6ea4daf79afe519eaec1667b8c933b0bd6864cae479e96513

Observation e6b96c13-e4cd-4786-a305-f25608f9ce1e · outbound

This paper cites Neuroheed+: Improving neuro-steered speaker extraction with joint auditory attention detection,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Neuroheed+: Improving neuro-steered speaker extraction with joint auditory attention detection,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.985128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.458874Z digest=sha256:ee884054a53028bd6a0e1eee82287bae3da49d5d85a16344334658cbeb9bec6b

Observation 0f2bcac3-be4f-490a-ab12-cecd3933a5e9 · outbound

This paper cites Paris: Pseudo-autoregressive siamese training for online speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Paris: Pseudo-autoregressive siamese training for online speech separation,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.968849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.463998Z digest=sha256:dd5bb51ae66aa91428d1f138858b4bfb72c8c1ca2c59f3f4e5e5ad54e548687b

Observation 900d7c57-838a-4484-a067-9c545a703533 · outbound

This paper cites On- line Audio-Visual Autoregressive Speaker Extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions On- line Audio-Visual Autoregressive Speaker Extraction,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.949809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.469257Z digest=sha256:3a0486a74ab5d69a054d906490c56db1b575139d6be2d16af41f8dbdb9639d65

Observation d6aa6f40-20ba-4b58-94f3-1292d36b9aab · outbound

This paper cites Attention is all you need,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Attention is all you need,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.473959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.473959Z digest=sha256:88c66ea7049671b176e90764e942de4cf6db93b68521f3492ae8bbed0a50d685

Observation 81286334-2444-45bf-baeb-f40273a994ad · outbound

This paper cites Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.916620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.478541Z digest=sha256:9726ed49a2fa69735201d72b2b43a92737d2b4b918c2fe4329546b328ea6de06

Observation f360c19f-93a5-4926-8a80-b8ebe629387f · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Wespeaker: A research and production oriented speaker embedding learning toolkit,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.897397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.483332Z digest=sha256:f8c408b92c400e8fa1311b99533835e6d1b5dbbfe7f1dbacf6f9424f7ef91bb0

Observation a5e18d5e-3975-41f5-83d8-e7782d961065 · outbound

This paper cites Curriculum learning,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Curriculum learning,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.488014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.488014Z digest=sha256:b612a014456467703c5c228f74974311798862708678336c2449b71777e29c6c

Observation c59cc09d-4efc-4930-8960-00e534280cab · outbound

This paper cites A learning algorithm for continually running fully recurrent neural networks,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions A learning algorithm for continually running fully recurrent neural networks,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.492392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.492392Z digest=sha256:f3be002460d4d1a67891de8a6e62bddf1624e2a5d391bdf35f24dd71233e8e17

Observation 3b0191e3-e40d-42e9-88a2-1557279844d5 · outbound

This paper cites Scheduled sampling for sequence prediction with recurrent neural networks,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Scheduled sampling for sequence prediction with recurrent neural networks,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.858309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.497009Z digest=sha256:b61d0bc1e4f7438b02f9fc2870c8f33346f19f298686f3f65a4eb0cb0d7bc124

Observation 17078865-8ccd-4862-b091-1f5803386d60 · outbound

This paper cites Delving into high-quality syn- thetic face occlusion segmentation datasets,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Delving into high-quality syn- thetic face occlusion segmentation datasets,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.842492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.502291Z digest=sha256:7967c0d764387191e58696b052df7ee117b0bba360772ba79f822f24986d1fc5

Observation ac2c226e-f775-4761-9c08-30a9a67f8a2a · outbound

This paper cites Watch or listen: Robust audio-visual speech recognition with visual corruption modeling and reliability scoring,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Watch or listen: Robust audio-visual speech recognition with visual corruption modeling and reliability scoring,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.825636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.506931Z digest=sha256:a2c1888396bebcb4011ca0c0a164d8427c2e41226dc3c97c1127162c214b2ce7

Observation a38afe9a-a0f5-4d0f-b894-fb657b1d27c7 · outbound

This paper cites V oxceleb2: Deep speaker recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions V oxceleb2: Deep speaker recognition,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.806766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.511270Z digest=sha256:6f22219f22c6de86d2c6ab9bd9f165594bc9b4326f88d5130176535f913879fc

Observation 60b84ecc-6bd7-4797-a666-4cc5be932f54 · outbound

This paper cites Avhumar: Audio-visual target speech extraction with pre-trained av-hubert and mask-and-recover strategy,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Avhumar: Audio-visual target speech extraction with pre-trained av-hubert and mask-and-recover strategy,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.787987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.515654Z digest=sha256:68616250d71ea3ec7b2344c90721a0a52a55a3ff959b9b2c2e098ddc3e57d8c7

Observation 7e5d0481-f3bb-42df-8a8c-4fe34c912811 · outbound

This paper cites Target speech extraction with pre-trained av-hubert and mask-and-recover strategy,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Target speech extraction with pre-trained av-hubert and mask-and-recover strategy,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.766513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.521222Z digest=sha256:7d5415180dce9dfd4543a93dc4c7dd674f2e45e7a5e29522a0ca051f9f71243f

Observation 89aae3a2-1857-492b-bb66-7adf0a6cd7da · outbound

This paper cites Sdr–half-baked or well done?.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Sdr–half-baked or well done?

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.749611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.525852Z digest=sha256:99f151e205090a92682ac8a6e939a61a3241d40f4bb4025c2caa7146449353e9

Observation 03e57a72-d5f5-46cf-8ff2-e8a5e8c6791b · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.733522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.530574Z digest=sha256:e5e046b0590e78346a731925fb631d0e3f106713c196ec7429247af37073e22f

Observation c9560885-5aa4-4169-8ffe-8241c468c3dd · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions A short- time objective intelligibility measure for time-frequency weighted noisy speech,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.713313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.535064Z digest=sha256:1480bda12da4edcf3f36a7c093edc9d18826606aadbbf740137653f80740f742

Observation f1c2c8e9-7380-4dc0-aab9-0c7cabd69ed8 · outbound

This paper cites Time domain audio visual speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Time domain audio visual speech separation,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.693485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.540459Z digest=sha256:f113f5c4e4194f05dd7c75c87a373e99e08701efcd40f81556fd204d5925c077

Observation 7cf1d4a5-6276-456f-9ad5-bdb419330ec7 · outbound

This paper cites Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.545087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.545087Z digest=sha256:97d98e298cdf69011d4133e63228886ac80dc6f371e241647c2d807ed1c15028

Observation 4453c919-dd08-4af9-a213-2da60ed02d89 · outbound

This paper cites Music source separation with band-split rnn,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Music source separation with band-split rnn,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.549514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.549514Z digest=sha256:8679fdca82ef19bde48555762741bc0095db34b7d691729040d642f529eb5e00

Observation 88f831c1-baf9-4a3f-9125-e6d48eb7bc4d · outbound

This paper cites Wesep: A scalable and flexible toolkit towards generalizable target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Wesep: A scalable and flexible toolkit towards generalizable target speaker extraction,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.648914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.554592Z digest=sha256:16f66f0347591dc623844fca0434092f753bbe228fde8efb8ffd8ce537a40ddc

Observation 8e34c688-aeea-432b-b1fc-c5cc47cd47d5 · outbound

This paper cites Audio-visual target speaker extraction with selective auditory attention,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Audio-visual target speaker extraction with selective auditory attention,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.631585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T15:40:09.559373Z digest=sha256:afb703b036ff2510b6cdf9b7f9298841dfe239775887da728a6ce55881bb53d5

Observation aa5cd0ee-cf22-4bee-8798-2d7dd572e36e · outbound

This paper cites Attention is all you need in speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Attention is all you need in speech separation,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.564250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.564250Z digest=sha256:8c2455dfa714b5b00b22e18b8b70d32f2b98ed725e0dcf51daf780f418f78af4

Pith citing papers

Observation a093622f-23ed-4eee-b9d4-df82da779cc2 · inbound

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction cites this paper.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.277622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.277622Z digest=sha256:d9d1fe50f4dde4afc16d77ebfeb29a8c73b849a09083d54af8f747895e810b0a