Pith. sign in

Paper Citation Record · LEDGER

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

As of 7 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2506.00466.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00466 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:08:17.836014Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:41:16.865462Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T18:41:18.737055Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact3
  • verified fuzzy31
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bc195b6e-d5d0-4f63-a9d0-04ac866c9326 · outbound

This paper cites Electrophysiological correlates of semantic dissimilarity reflect the comprehension of natural, narra- tive speech.Current Biology, 28(5):803–809,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Electrophysiological correlates of semantic dissimilarity reflect the comprehension of natural, narra- tive speech.Current Biology, 28(5):803–809,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.666999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:14.230964Z digest=sha256:4d67999e93dcb804ec738078df3440a4ff96100af8c79053d26ce2a729a0dbea

Observation 8f307be1-86c4-49ac-8956-6f9759ca9460 · outbound

This paper cites Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Improved Feature Extraction Network for Neuro-Oriented Target Speaker Extraction

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:08:18.114926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:14.691413Z digest=sha256:448fde63353c3848f1b44532deb0343c23aae4a9092b5588f6fc7430d8d01e20

Observation 13523b52-39f6-431a-b171-19ca9d351e34 · outbound

This paper cites L-spex: Localized target speaker extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction L-spex: Localized target speaker extraction

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.756034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:14.922931Z digest=sha256:409e455f8aa3d6a9190f78e6864fe60d3f1d45cfea5552791723cc220f62c766

Observation e2974a40-42d1-4de0-8110-e608a5038a5f · outbound

This paper cites The cocktail party problem.Neural computation, 17(9):1875– 1902,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction The cocktail party problem.Neural computation, 17(9):1875– 1902,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.856502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:15.321264Z digest=sha256:5f928816a3ba7e22987b80a39d10d80b8862d23b1478836b229188278e91cdf2

Observation df764bc2-ffa9-4846-8928-5a6b8931a09f · outbound

This paper cites Speaker-independent brain enhanced speech denoising.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Speaker-independent brain enhanced speech denoising

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.420567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:15.542157Z digest=sha256:045d8eff1970f4b98e3069a88ed050a9ba2dbc835c82122cc7d19914a07ee001

Observation 3c2dcafd-2de0-422b-af07-0b1cda3acfb3 · outbound

This paper cites Cross-modal global interaction and local alignment for audio-visual speech recognition.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Cross-modal global interaction and local alignment for audio-visual speech recognition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.007943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:15.738422Z digest=sha256:caf2e3bb90d4a4919a79272e7ea0bf69f163ed720915289950703d6dbe192c98

Observation 5fa6fbf7-6719-4168-b5c3-cc96e9102fa3 · outbound

This paper cites [Le Rouxet al., 2019 ] Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R Hershey.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction [Le Rouxet al., 2019 ] Jonathan Le Roux, Scott Wisdom, Hakan Erdogan, and John R Hershey

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.729131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:15.828302Z digest=sha256:31427f4ba6359c9516186adfc09a69d0173c35f8a239a4f9be4f55c3194cd24f

Observation efeb095b-13d3-4666-b5b3-3425874acfd2 · outbound

This paper cites Align before fuse: Vision and language representation learning with momentum distilla- tion.Advances in neural information processing systems, 34:9694–9705,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Align before fuse: Vision and language representation learning with momentum distilla- tion.Advances in neural information processing systems, 34:9694–9705,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:15.910755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:15.910755Z digest=sha256:7a5a459ab578a09a511c644f242a29118177ba173cc4dd4eb51d8ee49c527584

Observation 2efd5f25-00e7-4d3f-8874-eb720ed719a3 · outbound

This paper cites Audio-visual active speaker extraction for sparsely overlapped multi-talker speech.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Audio-visual active speaker extraction for sparsely overlapped multi-talker speech

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.476889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.002599Z digest=sha256:2fff7d0ead03d686c1f9a43ce7190e00c9e7d4d2f847cd9f4d44c60637f3dfb7

Observation 51088153-38b9-44cf-9d43-d86685d3973a · outbound

This paper cites Av- sepformer: Cross-attention sepformer for audio-visual tar- get speaker extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Av- sepformer: Cross-attention sepformer for audio-visual tar- get speaker extraction

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:21.207261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.099123Z digest=sha256:ebb6da808419ace0c88b239124374c062e67bbcdc2e44f624e8cce13362f8c06

Observation bafb6e57-f5e0-45ae-8f6c-cb2825084cc9 · outbound

This paper cites Development of the audi- tory system.Handbook of clinical neurology, 129:55–72,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Development of the audi- tory system.Handbook of clinical neurology, 129:55–72,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.995877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.177957Z digest=sha256:569a7adecbe33a1094c58d5691e0d38bd3663c61de36280961949d7d5146cc61

Observation 5bd4c207-95ca-4561-bd42-5119b2fd12e1 · outbound

This paper cites Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.631601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.365027Z digest=sha256:d28e349a0508dcee99b213b342ebe6a4f740f7814b47c65da07da9d52da555f1

Observation 1556424d-5a68-4f54-a3d0-172c329858fa · outbound

This paper cites Dbpnet: Dual- branch parallel network with temporal-frequency fusion for auditory attention detection.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Dbpnet: Dual- branch parallel network with temporal-frequency fusion for auditory attention detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.412865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.461928Z digest=sha256:d39725cbb0179668dbfe909babb6f59cbbf3472ef3ab6f51668fc83b011fb989

Observation 37e4535f-4d26-4e5c-9eeb-a1b405a74090 · outbound

This paper cites Attentional selection in a cocktail party environment can be decoded from single- trial eeg.Cerebral cortex, 25(7):1697–1706,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Attentional selection in a cocktail party environment can be decoded from single- trial eeg.Cerebral cortex, 25(7):1697–1706,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.240052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.586298Z digest=sha256:7c71f78ecda2dd6cd0fc5c897d8f8e0662d2edea7ad3deac25149516caa9e57d

Observation 6046fe9f-d1f6-4977-8a76-b8d3c9047317 · outbound

This paper cites Neural decoding of at- tentional selection in multi-speaker environments without access to clean sources.Journal of neural engineering, 14(5):056001,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Neural decoding of at- tentional selection in multi-speaker environments without access to clean sources.Journal of neural engineering, 14(5):056001,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.058858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.677988Z digest=sha256:ce6cebf5fc1a87d5e67a738d897c1c5be235ba5f31b0c017de5ae034f3fd733f

Observation b3e6997e-0ba2-4970-ac42-73566d32e77a · outbound

This paper cites Neu- roheed+: Improving neuro-steered speaker extraction with joint auditory attention detection.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Neu- roheed+: Improving neuro-steered speaker extraction with joint auditory attention detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.870687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.758386Z digest=sha256:3ccf52ca5ab57d0ea0d6dad7bff2b0e912c0c716e33ba056913a20e0a8757c24

Observation 20c99bcf-0126-4174-a528-1a75a08b5db6 · outbound

This paper cites Tf-nsse: A time–frequency domain neuro-steered speaker extractor.Applied Acous- tics, 211:109519,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Tf-nsse: A time–frequency domain neuro-steered speaker extractor.Applied Acous- tics, 211:109519,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.706750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.857130Z digest=sha256:06f8128e64b2f5fd2e0ce0c034f9ef366f751a95ffc895dae5329c15a091c412

Observation 6acfd61d-a7aa-40dc-a6c3-b881a3eb6a0e · outbound

This paper cites Phase space graph convolutional network for chaotic time series learning.IEEE Transactions on Industrial Informat- ics,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Phase space graph convolutional network for chaotic time series learning.IEEE Transactions on Industrial Informat- ics,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.492949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.950800Z digest=sha256:0aa815ffeeffb96e63068acaca4d60226039181f9b8cde2d1692f0dfd736be64

Observation 4beca81b-dc3b-475c-aa27-943dd8141bc5 · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.286102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.035539Z digest=sha256:49617777c7248082d43f04cd836f49092aa75301a250219656bb64076c0c77e3

Observation 8e2a4f39-bf0f-4ec0-ae23-f64f8f361103 · outbound

This paper cites An algorithm for intelligibil- ity prediction of time–frequency weighted noisy speech.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction An algorithm for intelligibil- ity prediction of time–frequency weighted noisy speech

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.136099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.277497Z digest=sha256:27dab63022348c095c22ae9d7ac6fc501a747b57ed1d6cbb88be77dd78beb16e

Observation 3cb4f928-39ea-4f02-acbd-1f4a5fbd31e0 · outbound

This paper cites A study of multichannel spatiotemporal features and knowledge distillation on robust target speaker extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction A study of multichannel spatiotemporal features and knowledge distillation on robust target speaker extraction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.950548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.447655Z digest=sha256:cbecfdcc4bac4992ee5e75d0c4bc2fa8c682479a837ca1d2e83dcc75b17ebc2b

Observation ff249c62-b548-44b8-aa3d-a2595cb51d0a · outbound

This paper cites Spex: Multi-scale time domain speaker extraction network.IEEE/ACM transactions on audio, speech, and language processing, 28:1370–1384,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Spex: Multi-scale time domain speaker extraction network.IEEE/ACM transactions on audio, speech, and language processing, 28:1370–1384,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.820826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.503332Z digest=sha256:3d93f81bcced4eeb88b7520f5c4d9c37ff3ac0c3e1c5860b08a18fc5a779c4f3

Observation 3ad76a10-e7c1-4e4d-8d48-72455b9e78c3 · outbound

This paper cites DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction DARNet: Dual Attention Refinement Network with Spatiotemporal Construction for Auditory Attention Detection

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:08:17.955113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.613996Z digest=sha256:a1213308b4f7301628bf5f4ec5c08f19505581081771f820f1d23bd88cdf6dcc

Observation bc8e5ec2-d71f-41e9-8950-eb2003d6de73 · outbound

This paper cites Basen: Time-domain brain-assisted speech enhancement network with convolutional cross at- tention in multi-talker conditions.Interspeech 2023,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Basen: Time-domain brain-assisted speech enhancement network with convolutional cross at- tention in multi-talker conditions.Interspeech 2023,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.705495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.692399Z digest=sha256:bdb65d85388c05dc99a2e969acc1481982226c0fe4698f7d0db122c5b76c2203

Observation 4c96733a-00b6-4f10-8b93-6354c97f32d8 · outbound

This paper cites Based on audio-video evoked auditory attention detection electroencephalogram dataset.Journal of Tsinghua University (Science and Tech- nology), 64(11):1919–1926,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Based on audio-video evoked auditory attention detection electroencephalogram dataset.Journal of Tsinghua University (Science and Tech- nology), 64(11):1919–1926,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.568754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.749495Z digest=sha256:b96f01742b6754caa270e4d1786ebbaad22355e84d0cabe9692707c36669011b

Observation 54d544ed-48a9-4f5b-b627-dfc15aff88a2 · outbound

This paper cites Neural target speech extraction: An overview.IEEE Signal Processing Magazine, 40(3):8–29, 2023.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Neural target speech extraction: An overview.IEEE Signal Processing Magazine, 40(3):8–29, 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:18.404454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.836014Z digest=sha256:982057ae2e82611d46b5a4fddb3edf7a0c6f6de0a74a7a12b1bf78ae1fce88e4

Observation cb6d813a-1da5-4dd1-82c2-343b6f83e2c9 · outbound

This paper cites GroupMamba: Efficient Group-Based Visual State Space Model.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction GroupMamba: Efficient Group-Based Visual State Space Model

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.151270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.151270Z digest=sha256:5baa92f03a37b492c9c9382d85797ba2c3d200aa665df0d2369335d594bc866b

Observation c4fc0c39-1540-4c36-83cd-f0c5369100b8 · outbound

This paper cites Centroid estimation with transformer-based speaker embedder for robust target speaker extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Centroid estimation with transformer-based speaker embedder for robust target speaker extraction

Reference 2005

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.610268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:15.440078Z digest=sha256:79aedb6b76d0b406d6c97b652f3807de3c9e90add233969abbcd4fbde7150746

Observation cc40db08-4309-4add-a24d-7547d36f37ff · outbound

This paper cites Speech intelligibility predicted from neural entrainment of the speech envelope.Journal of the Association for Re- search in Otolaryngology, 19:181–191,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Speech intelligibility predicted from neural entrainment of the speech envelope.Journal of the Association for Re- search in Otolaryngology, 19:181–191,

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:19.047429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:17.369243Z digest=sha256:41a89068edc4b4fc94ba5190a7388f89961c25f87090f9a64be23b2e99815d03

Observation 187771db-3a3e-4b65-a4cd-ec87fac5dad5 · outbound

This paper cites Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation.IEEE/ACM transactions on audio, speech, and language processing, 27(8):1256– 1266,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Conv-tasnet: Surpassing ideal time–frequency magnitude masking for speech separation.IEEE/ACM transactions on audio, speech, and language processing, 27(8):1256– 1266,

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:20.818684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:16.284920Z digest=sha256:4b5db07da93f1033ad2cffba534a90f23c1aca5ffbf07be192128282453ae3f6

Observation 493b7a6c-6905-47f2-88dd-dccd89182d9c · outbound

This paper cites Brain-informed speech separation (biss) for enhancement of target speaker in multitalker speech perception.NeuroImage, 223:117282,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Brain-informed speech separation (biss) for enhancement of target speaker in multitalker speech perception.NeuroImage, 223:117282,

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.313507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:14.335409Z digest=sha256:2d633841291b0e3fc903b2b93c82a76ccf67f0a99f2bc51c254a060f353a5956

Observation d0c08969-5790-4ea4-b3f3-29de1790de9c · outbound

This paper cites Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Typing to Listen at the Cocktail Party: Text-Guided Target Speaker Extraction

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:15.107005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:15.107005Z digest=sha256:7d4a28a21da280c43295cbedd513d386ad6b08773abab8c51cf7d73b0bd57c78

Observation 0cd41983-8063-42d2-adc1-94a39ac04e1c · outbound

This paper cites NeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction NeuroSpex: Neuro-Guided Speaker Extraction with Cross-Modal Attention

Reference 2020

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:08:18.245325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:14.465323Z digest=sha256:8f520574d257d4ccd55281f323081d5350d9d0b9c500706d7fa1e02e4a2018e2

Observation f3f7cfbf-b1ee-46ac-8306-2586aa49a79c · outbound

This paper cites End-to-end brain-driven speech enhance- ment in multi-talker conditions.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:1718– 1733,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction End-to-end brain-driven speech enhance- ment in multi-talker conditions.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:1718– 1733,

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:22.213036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:15.638340Z digest=sha256:111901a4ecc33d886177e706039037b4a2eec43d2d079641457285e6c5a114ba

Observation 65930c91-b637-452e-9df8-32bf922a2c89 · outbound

This paper cites Speaker-independent auditory attention decoding with- out access to clean speech sources.Science advances, 5(5):eaav6134,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Speaker-independent auditory attention decoding with- out access to clean speech sources.Science advances, 5(5):eaav6134,

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.495225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:15.015545Z digest=sha256:fff3bc8ec3aa922130bea0ecb37f3494212fed6f7c0142a87578027ab5fd9868

Observation f78d3635-634d-4020-9633-6f9d2e4ad98c · outbound

This paper cites X-tf-gridnet: A time–frequency domain target speaker extraction network with adaptive speaker embed- ding fusion.Information Fusion, 112:102550,.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction X-tf-gridnet: A time–frequency domain target speaker extraction network with adaptive speaker embed- ding fusion.Information Fusion, 112:102550,

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:23.157869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:15.228471Z digest=sha256:e92c737ac98ea72bba034febb86af7f31db2e614983f05b38245528517d7d9f3

Observation e8195b1e-72b8-4d7c-aa33-933a0806b837 · outbound

This paper cites Msfnet: Multi-scale fusion net- work for brain-controlled speaker extraction.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction Msfnet: Multi-scale fusion net- work for brain-controlled speaker extraction

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:08:24.082962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:08:14.580915Z digest=sha256:ae20bc801307d27c092a65983a50a4d4d630ee1984c17dbb6266c98fe91a5a4a

Observation b3fb19f1-d6c2-4a3c-ad58-871fcf6c22aa · outbound

This paper cites SpEx+: A Complete Time Domain Speaker Extraction Network.

M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction SpEx+: A Complete Time Domain Speaker Extraction Network

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:14.833000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:14.833000Z digest=sha256:bf5b1bc52444521c731f7f460ec488accd423aebb8cf10d9ecd24e4843d0d485

Pith citing papers

Observation 8de67a4d-dd98-4d1f-9a7e-c8cfb58efa2a · inbound

DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction cites this paper.

DMF2Mel: A Dynamic Multiscale Fusion Network for EEG-Driven Mel Spectrogram Reconstruction M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:41:18.762933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:41:16.865462Z digest=sha256:de8e0bd6987fbc2ae4d11de8b4beb74e857be9b85b9caf7e5443cdeeeae64f9c

Observation 36a42793-abee-433f-8157-6c307e1077df · inbound

Decoding Speech Envelopes from Electroencephalogram with a Contrastive Pearson Correlation Coefficient Loss cites this paper.

Decoding Speech Envelopes from Electroencephalogram with a Contrastive Pearson Correlation Coefficient Loss M3ANet: Multi-scale and Multi-Modal Alignment Network for Brain-Assisted Target Speaker Extraction

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T07:24:44.110448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:24:44.110448Z digest=sha256:af191fea5c5f8ab59f7f8216afa8334dd5a1255150e8d2d5abd9a318674cfc36