Pith. sign in

Paper Citation Record · LEDGER

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions

As of 8 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 1 inbound Pith citation observation for arXiv:2507.15294.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15294 v2

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:40:09.564250Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T19:39:50.277622Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact0
  • verified fuzzy72
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 83c6922e-33a6-4d1b-9c21-65841d11fc41 · outbound

This paper cites Some experiments on the recognition of speech, with one and with two ears,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Some experiments on the recognition of speech, with one and with two ears,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.175918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.175918Z digest=sha256:1d47d18ccd7fe95d469606fbec0c9983cbdee2dae4a9d205d33133fd62770ea2

Observation 744f2f48-ebd7-4a45-8006-d2f53eb17f8b · outbound

This paper cites The cocktail party phenomenon: A review of research on speech intelligibility in multiple-talker conditions,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions The cocktail party phenomenon: A review of research on speech intelligibility in multiple-talker conditions,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.181286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.181286Z digest=sha256:6425496fe6635776db08ca49f9ce34a698ae52c197d7696ee44c385b19444793

Observation 8b3c2740-4d0c-4fe9-9692-f331dc2aa051 · outbound

This paper cites Seeing to hear better: evidence for early audio-visual interactions in speech identification,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Seeing to hear better: evidence for early audio-visual interactions in speech identification,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.949415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.186675Z digest=sha256:e7d6f01a87b5369d33ce9e5583d429e3a5f53e2280e0010303df7300702d9d09

Observation 95ec1176-f5ff-4daf-aed3-5b0f40666070 · outbound

This paper cites Visual contribution to speech intelligibility in noise,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Visual contribution to speech intelligibility in noise,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.932674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.191605Z digest=sha256:e8ba72290157b3f97cf1a6afad89d77931931430d33a4b0a97db783f6d47ebe8

Observation 4ff5073d-3358-4e53-8caa-ed50b21744a5 · outbound

This paper cites Muse: Multi-modal target speaker extraction with visual cues,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Muse: Multi-modal target speaker extraction with visual cues,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.915156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.196835Z digest=sha256:160016c9aa019eb674431ddf2974a818bb687f804c7b49bdaa8af4973bba084f

Observation 2763e41b-075b-4632-925b-f5e4921a5a8c · outbound

This paper cites Usev: Universal speaker extraction with visual cue,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Usev: Universal speaker extraction with visual cue,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.885618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.201540Z digest=sha256:a11667ed6a6906ce307d112bbd7ea802c1a617438c924dbc8eea8e6a5085a5a7

Observation d1c517a4-72d4-4dbd-b3a0-a2c100695cab · outbound

This paper cites Time domain audio visual speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Time domain audio visual speech separation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.866995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.207358Z digest=sha256:dbdb7321f50ae01f64934d79bcd04d2902fdfcf0dea41c41df23fac28eddacbb

Observation 094b854b-5973-4c6b-a35e-cf83b49c2596 · outbound

This paper cites Neural target speech extraction: An overview,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Neural target speech extraction: An overview,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.212164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.212164Z digest=sha256:0dce6467ef7b5a86c1ed52a6ba1ec27da2afc2aaa23e5007ee04a29cd125806b

Observation a76edd7b-75fe-4c81-b09d-a2a30fc5197c · outbound

This paper cites An overview of deep-learning-based audio-visual speech en- hancement and separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions An overview of deep-learning-based audio-visual speech en- hancement and separation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.841856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.217781Z digest=sha256:47ce261bbe630a11d60c0ed1545fa93732fa1ecb06efcc8749c831453aa9cdd6

Observation 1a4b2d8f-8080-4e8e-a875-1c91d982f1bd · outbound

This paper cites PIA VE: A Pose-Invariant Audio- Visual Speaker Extraction Network,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions PIA VE: A Pose-Invariant Audio- Visual Speaker Extraction Network,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.826667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.222557Z digest=sha256:5d3395ef9b5185064b61d5c81cfc9e5057c968b0bc92fbe65a69ae7960c9ab50

Observation 7d13eadd-9117-4305-83c8-634b94bccd40 · outbound

This paper cites Speaker extraction with co-speech gestures cue,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Speaker extraction with co-speech gestures cue,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.809836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.227254Z digest=sha256:2d96e0f2778b84dc7ef4475139b65772da4f77384a38e1d281b06e6ce0d83ebc

Observation 0f02fc49-23af-4400-a7a3-c539854bd177 · outbound

This paper cites Rethinking the Visual Cues in Audio-Visual Speaker Extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Rethinking the Visual Cues in Audio-Visual Speaker Extraction,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.792840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.231505Z digest=sha256:1f2b946cc39d1c758393e0f677112caac552157a85bcf886e9c78c08bc5a1762

Observation 302839d1-bece-43a4-8150-2868d3c52695 · outbound

This paper cites FaceFilter: Audio- Visual Speech Separation Using Still Images,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions FaceFilter: Audio- Visual Speech Separation Using Still Images,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.767074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.235980Z digest=sha256:679ea9b283decf7e0c169dafaf1c2d2ec9f59c23a481a4880187ed05eac47879

Observation 91365b9e-2036-43e3-b47e-3086683cfa8c · outbound

This paper cites c 2av-tse: Context and confidence-aware audio visual target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions c 2av-tse: Context and confidence-aware audio visual target speaker extraction,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.751607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.239987Z digest=sha256:c9dc4cdf91e083b75dcf94784889536db6f7a247a365b838ae997eafc25d55b9

Observation 0a5ad71e-687a-4064-b397-262808e4d4ae · outbound

This paper cites Incorporating linguistic constraints from external knowledge source for audio-visual target speech extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Incorporating linguistic constraints from external knowledge source for audio-visual target speech extraction,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.736790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.244151Z digest=sha256:0fd5a5c46167dc0d650aefa8886a5385c3dc3fbac7e967794a7023f2565079de

Observation 982b878b-e97d-4872-a5db-3910f7170acb · outbound

This paper cites Iianet: An intra- and inter-modality attention network for audio-visual speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Iianet: An intra- and inter-modality attention network for audio-visual speech separation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.721532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.248745Z digest=sha256:a83cc5dd1b9b78c572ac9a66ad8449c9e6da673c54e72b3906cef22422dccccc

Observation 68eefdee-0a10-4910-844e-7f8f666b21d1 · outbound

This paper cites Hearing lips in noise: Universal viseme-phoneme mapping and transfer for robust audio-visual speech recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Hearing lips in noise: Universal viseme-phoneme mapping and transfer for robust audio-visual speech recognition,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.704696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.253262Z digest=sha256:5f2d735fa18775ceae718f4251ff21455d13b15bdce4d973ed65fd3739011000

Observation 5516288b-28fe-4e07-9780-6485d57fdb4e · outbound

This paper cites Diarization is hard: Some experiences and lessons learned for the jhu team in the inaugural dihard challenge.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Diarization is hard: Some experiences and lessons learned for the jhu team in the inaugural dihard challenge

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.688983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.257996Z digest=sha256:3b0c345238a0a5f0285d53d9a7a87d6e507841557bce9dc88f023b1d0783443b

Observation 251799e1-7c01-4669-abea-64bcb8f411b9 · outbound

This paper cites Noise- disentanglement metric learning for robust speaker verification,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Noise- disentanglement metric learning for robust speaker verification,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.672219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.263291Z digest=sha256:6f71ec1337a09c7a60f0d2b1e2b3b7873ce270ee382cc4e4ea147cc1b3b0cd62

Observation 9b109bc0-8f1d-4e45-9b0c-b6a4cc57e48f · outbound

This paper cites Momuse: Momentum multi-modal target speaker extraction for real-time scenarios with impaired visual cues,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Momuse: Momentum multi-modal target speaker extraction for real-time scenarios with impaired visual cues,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.656814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.267901Z digest=sha256:a48e81f321fa4728865aa1d0a7043656c7c2de6f20a32111f892514411d2e73d

Observation 0cb45b8b-12b0-4d2d-bc15-9b7e1d9cb579 · outbound

This paper cites Ravss: Robust audio- visual speech separation in multi-speaker scenarios with missing visual cues,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Ravss: Robust audio- visual speech separation in multi-speaker scenarios with missing visual cues,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.641107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.273119Z digest=sha256:b9029eef746edad05d4d3c179d0f3491a97bb59f006e5e2bd64ebb0efd74b1cb

Observation ca417b6c-9d8f-497e-817a-1d9c3321d122 · outbound

This paper cites Switching variational auto- encoders for noise-agnostic audio-visual speech enhancement,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Switching variational auto- encoders for noise-agnostic audio-visual speech enhancement,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.621867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.277199Z digest=sha256:0a42d9e298c6687b75cf9f040dfaec9c0ecd42af320101c7aef6d58855a26a16

Observation d4283792-3a6a-4479-af6c-1321dd6c1cc4 · outbound

This paper cites Robust unsupervised audio-visual speech enhance- ment using a mixture of variational autoencoders,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Robust unsupervised audio-visual speech enhance- ment using a mixture of variational autoencoders,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.606055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.281465Z digest=sha256:b7a7c53811396ba919b6b9e4a446c3ddd7a53fe6c4d77947fc9613036bceb393

Observation 5f16851f-04db-428a-988a-61d4a7da1a78 · outbound

This paper cites Time-domain audio-visual speech separation on low quality videos,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Time-domain audio-visual speech separation on low quality videos,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.590626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.286027Z digest=sha256:acf1ad7306bb908334a4be4420e7e71a2715e4c5666a1051c22eaf0fa4612ae6

Observation 952af7c7-1dfe-4740-b23e-d0bd0ab8c3a9 · outbound

This paper cites Imaginenet: Target speaker extraction with intermittent visual cue through embedding inpainting,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Imaginenet: Target speaker extraction with intermittent visual cue through embedding inpainting,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.574988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.291033Z digest=sha256:6c892901924b216f06550368e6f368fd0a884aba2aad25dedac9b0a955e4f5e7

Observation e7dcd13f-e53e-4227-a762-1dd98aa07395 · outbound

This paper cites My lips are concealed: Audio-visual speech enhancement through obstructions,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions My lips are concealed: Audio-visual speech enhancement through obstructions,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.559734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.296351Z digest=sha256:a7593f005ffdf37df5517d7136eae4bfb22da9f280dca7ec398fcd7f180f8cfa

Observation 6619f601-ea29-4e01-9e11-6d48e8324806 · outbound

This paper cites Multi-cue guided semi-supervised learning toward target speaker separation in real environments,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multi-cue guided semi-supervised learning toward target speaker separation in real environments,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.545050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.300822Z digest=sha256:4295c9239bd2a183f27bb35fc873668fbfa229f81b5384dcbdf0c2daddc24a6a

Observation f4f9bccf-176a-4260-a95b-6ca2b83a4649 · outbound

This paper cites A two-stage audio-visual speech separation method without visual signals for testing and tuples loss with dynamic margin,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions A two-stage audio-visual speech separation method without visual signals for testing and tuples loss with dynamic margin,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.528996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.305414Z digest=sha256:29c6c8130e16a2d2504fa7fd723a6f3a527fda53a9225e1b731851726f7e001c

Observation 84bce1f1-23d6-4331-9ed0-991756082d7f · outbound

This paper cites Cross-modal speech separation without visual information during testing,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Cross-modal speech separation without visual information during testing,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.513641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.311071Z digest=sha256:ec0d8d928fa838bd0d2dcc1f808c66752596516ffa5fe836e03c60e66bc9337c

Observation d62436f8-6c39-4acd-acbb-d67d25f39ccc · outbound

This paper cites Robust audio-visual speech enhancement: Correcting misassignments in complex environments with advanced post-processing,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Robust audio-visual speech enhancement: Correcting misassignments in complex environments with advanced post-processing,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.498357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.317501Z digest=sha256:67f4eda0ec376c1c2798f759e5c0edc12897cc5e1747b26d64488acf04fd6408

Observation c459807a-0a79-4ed6-8058-0cea66615d59 · outbound

This paper cites Cocktail party listening in a dynamic multitalker environment,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Cocktail party listening in a dynamic multitalker environment,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.482734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.323438Z digest=sha256:4ca9fbbe7d5f1df84c025040f20706672e39c239dad7e9107d6063af040d538e

Observation e408aa7c-a002-478f-a24e-080c71069d95 · outbound

This paper cites The advantage of knowing where to listen,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions The advantage of knowing where to listen,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.467314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.329382Z digest=sha256:a19e71667aa8304e8ed97c50f04339e9c589529b635fa94d803733c9affb551e

Observation c7f82ef1-b792-452b-a6db-e60f3923d14c · outbound

This paper cites Object continuity enhances selective auditory attention,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Object continuity enhances selective auditory attention,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.450245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.334812Z digest=sha256:36e399f964bb22167bd81c9ad9551d093ea4de19deb9d705f8c9f44a68a67d46

Observation 67e2d8c8-66db-4084-b3cc-1a7574113c4e · outbound

This paper cites Enhanced learning through multimodal training: evidence from a comprehensive cognitive, physical fitness, and neuroscience intervention,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Enhanced learning through multimodal training: evidence from a comprehensive cognitive, physical fitness, and neuroscience intervention,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.432061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.339775Z digest=sha256:68b674fa458e5f657cb29270f95e53f0bb4137964835eaed721c8187adbf8c4b

Observation c202bef3-25b1-4d8c-a92a-f2e3d8b3cbbe · outbound

This paper cites The important role of contex- tual information in speech perception in cochlear implant users and its consequences in speech tests,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions The important role of contex- tual information in speech perception in cochlear implant users and its consequences in speech tests,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.414053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.344866Z digest=sha256:ee7966b68930b9e9faf7ef71133cb88ea4d994efcdb45a82050d7747c9dab0bb

Observation 92d5150e-1fa1-4fa2-8fe1-bc72d4bc2acf · outbound

This paper cites Attention and working memory in human auditory cortex,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Attention and working memory in human auditory cortex,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.395278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.350053Z digest=sha256:d1fd3e97314ef99846e19c8de6dd971ecf3b418c66868c7756b080da8d2fa1fd

Observation 877109a9-6ee5-418f-aefb-59121d7b42b2 · outbound

This paper cites Interactions between attention and working memory,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Interactions between attention and working memory,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.377778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.354448Z digest=sha256:fc36a778f18743b58426b43c3737dc7339ed0d74cd181447110b2da82a574ec8

Observation f1508676-d493-4fde-b68c-adfcae7d7c55 · outbound

This paper cites Look once to hear: Target speech hearing with noisy examples,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Look once to hear: Target speech hearing with noisy examples,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.361411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.358543Z digest=sha256:ff6fcb6547935e0951eb35bc57eb9d7e8f710009dd04142ed2825d4d4d8e3886

Observation e077f5c0-d99c-48aa-9f92-fb7421d248c6 · outbound

This paper cites Multimodal attention fusion for target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multimodal attention fusion for target speaker extraction,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.362651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.362651Z digest=sha256:7d3d01b55049b6bfef152bbf35215828e251cc821cca531ed2d9bd64be527126

Observation d6d0b610-dd7f-440e-bebd-a6e46823b311 · outbound

This paper cites Multimodal speakerbeam: Single channel target speech extraction with audio-visual speaker clues.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multimodal speakerbeam: Single channel target speech extraction with audio-visual speaker clues

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.332485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.366891Z digest=sha256:09fedddbf189bab870e5613e3b08c0569564facef654347c3f875dafa3cb262f

Observation 5987cf6c-44af-40fa-a653-65ff48ffcdec · outbound

This paper cites Memory networks,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Memory networks,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.311379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.371554Z digest=sha256:d4c3ad1361c1f26571d80b36f284159efea116431e78d557a6e5792ec8bfa672

Observation 3b1f9cce-ca92-414a-97b1-3694b4dd8106 · outbound

This paper cites End-to-end memory net- works,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions End-to-end memory net- works,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.289641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.377652Z digest=sha256:5b080666e3b211a4ec13ca644df03fc8da2c8d22cc820cdc5aa84724a632b86c

Observation c9af338d-836f-4736-ac78-904d87b15541 · outbound

This paper cites Unsupervised feature learning via non-parametric instance discrimination,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Unsupervised feature learning via non-parametric instance discrimination,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.268598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.382343Z digest=sha256:6217aae19e98695e6320b786aac761ca48e07523f21d26df33c900acc6831b02

Observation 91992967-d5e3-4a92-b2e2-d45a6a180914 · outbound

This paper cites Cromm-vsr: Cross-modal memory augmented visual speech recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Cromm-vsr: Cross-modal memory augmented visual speech recognition,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.250036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.386674Z digest=sha256:7529f228d39873b024ca158a5af2e848373022580984163ca0e3d415002cc9bc

Observation f400ee2c-ab6b-4fc8-8a4a-bb3348d44d10 · outbound

This paper cites Multi-temporal lip-audio memory for visual speech recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multi-temporal lip-audio memory for visual speech recognition,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.235219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.391493Z digest=sha256:d7a02287984980f87c1f995e6ca794657817612fa729ab09be5c696f0446b5c6

Observation 3d307f7b-3734-4902-b509-f0f6e5b525b0 · outbound

This paper cites Multi-modality associative bridging through memory: Speech sound recollected from face video,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Multi-modality associative bridging through memory: Speech sound recollected from face video,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.220095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.395575Z digest=sha256:908f7a662d6d155c5c42447185c18e7a591f1fe58b34261a58524ddfbb825767

Observation b7540409-2519-41ea-b221-7034dc95db8b · outbound

This paper cites Akvsr: Audio knowledge empowered visual speech recognition by compressing audio knowledge of a pretrained model,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Akvsr: Audio knowledge empowered visual speech recognition by compressing audio knowledge of a pretrained model,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.202608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.400260Z digest=sha256:22a8c3d48a8ddb9f80f13c9fa2fbb9a5c63a741b1622fabdd3574d5e37fa46b7

Observation 62dbf9dc-1e99-402c-a311-d97e8e29a2b3 · outbound

This paper cites Distinguishing homophenes using multi-head visual-audio memory for lip reading,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Distinguishing homophenes using multi-head visual-audio memory for lip reading,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.184305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.405373Z digest=sha256:1506b7a1bb343f02c1976bf91a239b9cc6902acb70fa0d42024282a457601457

Observation 8b8a75df-6b28-4901-82c0-914c3a569f20 · outbound

This paper cites Speech reconstruction with reminiscent sound via visual voice memory,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Speech reconstruction with reminiscent sound via visual voice memory,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.166533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.409884Z digest=sha256:3e04ca661f1d9c6d077bf5018219ff78ac0b27ee52dab844af46a84aecb39dd3

Observation aec87788-c5ef-4974-a09f-11d8b2aa3a8e · outbound

This paper cites Modeling attention and memory for auditory selection in a cocktail party environment,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Modeling attention and memory for auditory selection in a cocktail party environment,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.151632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.414371Z digest=sha256:6f198cc34c3e78f02fb9f37048521bc8dde6c7385d0079c0984526e0793c8919

Observation 4057b1ff-ab00-44c2-a304-2d26fb5eb04f · outbound

This paper cites Explicit-memory multiresolution adaptive framework for speech and music separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Explicit-memory multiresolution adaptive framework for speech and music separation,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.135206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.419236Z digest=sha256:fdce5c2cb197853233d0aa9c9f0d4db440e1ea9c2c9bc8dbbd09697a2e255274

Observation f8dedcdf-0717-4d31-ace6-d546dc222950 · outbound

This paper cites On the effectiveness of enrollment speech augmentation for target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions On the effectiveness of enrollment speech augmentation for target speaker extraction,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.118919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.423757Z digest=sha256:a71a92ae87a8d7692bcd772f531ad10ae2682ec2b6e9130e7c4730f63d8debfd

Observation 09f89054-bda5-48b8-8c9f-f5e021deb3c4 · outbound

This paper cites Selective listening by synchronizing speech with lips,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Selective listening by synchronizing speech with lips,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.102594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.428367Z digest=sha256:22f1f02a4429a858141e44ab86e02f78317f424bed3063a21b155729f1dba197

Observation 02d088bf-9e49-468c-b71e-80c73883c938 · outbound

This paper cites Robust Speaker Extraction Network Based on Iterative Refined Adap- tation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Robust Speaker Extraction Network Based on Iterative Refined Adap- tation,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.086812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.432582Z digest=sha256:2ea94c1b24bd90f3d8bee6209fa0b842771bbd81abe9dfd739a2a7aeb3d181aa

Observation 7752d5ec-a8d0-4ad4-b71a-94b4668935d2 · outbound

This paper cites Listening and group- ing: an online autoregressive approach for monaural speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Listening and group- ing: an online autoregressive approach for monaural speech separation,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.070125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.437077Z digest=sha256:cca5b710df93acee7509a4daf7c9f2cb55bbf42d5145a3016c2ee957d34e959b

Observation ec8da356-41c7-4aa4-ad74-9509c544109d · outbound

This paper cites Source-aware context network for single-channel multi-speaker speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Source-aware context network for single-channel multi-speaker speech separation,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.053802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.441118Z digest=sha256:b93836befcd1fab326a04498f03a7d96a75aff1c91dbe907bdc956e53828d14e

Observation 7b63a849-45b7-4346-b50a-5dfc98bd0baa · outbound

This paper cites An online speaker-aware speech separation approach based on time-domain representation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions An online speaker-aware speech separation approach based on time-domain representation,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.037326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.445488Z digest=sha256:09685e9e182808b72d007bb98829a9f62acb9a291d2148a151a56792715ec944

Observation c75e5f9b-344c-48e1-a81c-cbb23b9a1946 · outbound

This paper cites Iterative autoregression: a novel trick to improve your low-latency speech enhancement model,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Iterative autoregression: a novel trick to improve your low-latency speech enhancement model,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.020906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.450282Z digest=sha256:04f82d464fec2900c45a58e062b92e716794ef5a0fa47565f3013fc739fe0f6c

Observation f3f3203c-8165-4df5-8d2f-c873e09cbeab · outbound

This paper cites Neuroheed: Neuro- steered speaker extraction using eeg signals,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Neuroheed: Neuro- steered speaker extraction using eeg signals,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:10.004328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.454651Z digest=sha256:7cbbbe50eefb6318b696f00fb50e94539c15a6d59c908f19224dfb5c393e0d60

Observation e6b96c13-e4cd-4786-a305-f25608f9ce1e · outbound

This paper cites Neuroheed+: Improving neuro-steered speaker extraction with joint auditory attention detection,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Neuroheed+: Improving neuro-steered speaker extraction with joint auditory attention detection,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.985128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.458874Z digest=sha256:15def655697ebb2311fa9f148f30b992add2b26238fe5c384b198adc66019798

Observation 0f2bcac3-be4f-490a-ab12-cecd3933a5e9 · outbound

This paper cites Paris: Pseudo-autoregressive siamese training for online speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Paris: Pseudo-autoregressive siamese training for online speech separation,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.968849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.463998Z digest=sha256:c54d6fa8081f06bf1e94c45938c6073d0303e0050c65b687bca556d921668c0d

Observation 900d7c57-838a-4484-a067-9c545a703533 · outbound

This paper cites On- line Audio-Visual Autoregressive Speaker Extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions On- line Audio-Visual Autoregressive Speaker Extraction,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.949809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.469257Z digest=sha256:253d938258c7b76b3982b648afe92f57c05b0aa6e12c3b7ce12da1a4dbb3f9ef

Observation d6aa6f40-20ba-4b58-94f3-1292d36b9aab · outbound

This paper cites Attention is all you need,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Attention is all you need,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.473959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.473959Z digest=sha256:88c66ea7049671b176e90764e942de4cf6db93b68521f3492ae8bbed0a50d685

Observation 81286334-2444-45bf-baeb-f40273a994ad · outbound

This paper cites Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Ecapa-tdnn: Em- phasized channel attention, propagation and aggregation in tdnn based speaker verification,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.916620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.478541Z digest=sha256:369563fe1a71e21ccd7f5d1167b10df6b72880274df31a953abef015f7050716

Observation f360c19f-93a5-4926-8a80-b8ebe629387f · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Wespeaker: A research and production oriented speaker embedding learning toolkit,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.897397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.483332Z digest=sha256:d083fd69a14e3c83e87ea48528e8c3abfbeb2e86411c71e562ff6f5fd542e523

Observation a5e18d5e-3975-41f5-83d8-e7782d961065 · outbound

This paper cites Curriculum learning,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Curriculum learning,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.488014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.488014Z digest=sha256:b612a014456467703c5c228f74974311798862708678336c2449b71777e29c6c

Observation c59cc09d-4efc-4930-8960-00e534280cab · outbound

This paper cites A learning algorithm for continually running fully recurrent neural networks,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions A learning algorithm for continually running fully recurrent neural networks,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.492392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.492392Z digest=sha256:f3be002460d4d1a67891de8a6e62bddf1624e2a5d391bdf35f24dd71233e8e17

Observation 3b0191e3-e40d-42e9-88a2-1557279844d5 · outbound

This paper cites Scheduled sampling for sequence prediction with recurrent neural networks,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Scheduled sampling for sequence prediction with recurrent neural networks,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.858309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.497009Z digest=sha256:25e7ccc7f224696e451c73e98e71906cdb823ee7b3eb32cb28705c784dca0599

Observation 17078865-8ccd-4862-b091-1f5803386d60 · outbound

This paper cites Delving into high-quality syn- thetic face occlusion segmentation datasets,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Delving into high-quality syn- thetic face occlusion segmentation datasets,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.842492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.502291Z digest=sha256:7f34ce1c41fc4df8a9607472f4c4d7558944b0cc2f6274df4c930e6fea63b915

Observation ac2c226e-f775-4761-9c08-30a9a67f8a2a · outbound

This paper cites Watch or listen: Robust audio-visual speech recognition with visual corruption modeling and reliability scoring,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Watch or listen: Robust audio-visual speech recognition with visual corruption modeling and reliability scoring,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.825636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.506931Z digest=sha256:d20582add53abea12250c849cd972f557355a47cb5a2f7d6739e8a25ce229fd7

Observation a38afe9a-a0f5-4d0f-b894-fb657b1d27c7 · outbound

This paper cites V oxceleb2: Deep speaker recognition,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions V oxceleb2: Deep speaker recognition,

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.806766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.511270Z digest=sha256:af1f6330aae15cf071f156bb7ff36c7aee0ec0e6b42407d26c97619aeee0b9bb

Observation 60b84ecc-6bd7-4797-a666-4cc5be932f54 · outbound

This paper cites Avhumar: Audio-visual target speech extraction with pre-trained av-hubert and mask-and-recover strategy,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Avhumar: Audio-visual target speech extraction with pre-trained av-hubert and mask-and-recover strategy,

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.787987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.515654Z digest=sha256:ea83f8157e99fab5ee7d8c032539a4acf9802fb979fc4097445442dc1793b27a

Observation 7e5d0481-f3bb-42df-8a8c-4fe34c912811 · outbound

This paper cites Target speech extraction with pre-trained av-hubert and mask-and-recover strategy,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Target speech extraction with pre-trained av-hubert and mask-and-recover strategy,

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.766513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.521222Z digest=sha256:3883e891cfd11e8467968d10cbdd7a3f06708a4df7300bcd37a9440a3de6b528

Observation 89aae3a2-1857-492b-bb66-7adf0a6cd7da · outbound

This paper cites Sdr–half-baked or well done?.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Sdr–half-baked or well done?

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.749611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.525852Z digest=sha256:7a7c0896906ff2ac67428bf36a974bf268a043ccb756d9e34d7801a085ec2f52

Observation 03e57a72-d5f5-46cf-8ff2-e8a5e8c6791b · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.733522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.530574Z digest=sha256:fb0a2c4b07bbefc52c889ad22fb29870c30fda150d5934f431d511e03aa47051

Observation c9560885-5aa4-4169-8ffe-8241c468c3dd · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions A short- time objective intelligibility measure for time-frequency weighted noisy speech,

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.713313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.535064Z digest=sha256:0bf421495cb1f276eadc2cdf1ad9657e4279e03077c12d53e1a92a3de2fe16be

Observation f1c2c8e9-7380-4dc0-aab9-0c7cabd69ed8 · outbound

This paper cites Time domain audio visual speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Time domain audio visual speech separation,

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.693485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.540459Z digest=sha256:0b6c1478c80605237571f8dc7c2f765c7187e0880ebcc00d2ab46e5366a55b40

Observation 7cf1d4a5-6276-456f-9ad5-bdb419330ec7 · outbound

This paper cites Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation,

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.545087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.545087Z digest=sha256:97d98e298cdf69011d4133e63228886ac80dc6f371e241647c2d807ed1c15028

Observation 4453c919-dd08-4af9-a213-2da60ed02d89 · outbound

This paper cites Music source separation with band-split rnn,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Music source separation with band-split rnn,

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.549514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.549514Z digest=sha256:8679fdca82ef19bde48555762741bc0095db34b7d691729040d642f529eb5e00

Observation 88f831c1-baf9-4a3f-9125-e6d48eb7bc4d · outbound

This paper cites Wesep: A scalable and flexible toolkit towards generalizable target speaker extraction,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Wesep: A scalable and flexible toolkit towards generalizable target speaker extraction,

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.648914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.554592Z digest=sha256:6ac5a52a2d795d076637d6c46fb608377453a81f055d5ad04ad5ead0f9cc1397

Observation 8e34c688-aeea-432b-b1fc-c5cc47cd47d5 · outbound

This paper cites Audio-visual target speaker extraction with selective auditory attention,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Audio-visual target speaker extraction with selective auditory attention,

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:40:09.631585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:40:09.559373Z digest=sha256:0edf217205337a4053321a54cd8a07a9cbf8350e07c79905e7937364dbc89c17

Observation aa5cd0ee-cf22-4bee-8798-2d7dd572e36e · outbound

This paper cites Attention is all you need in speech separation,.

MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions Attention is all you need in speech separation,

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T15:40:09.564250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:40:09.564250Z digest=sha256:8c2455dfa714b5b00b22e18b8b70d32f2b98ed725e0dcf51daf780f418f78af4

Pith citing papers

Observation a093622f-23ed-4eee-b9d4-df82da779cc2 · inbound

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction cites this paper.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.277622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.277622Z digest=sha256:7f82063b447f083bc6bfe04c27f539e3a33355123a9287ae75b39a8db7524b12