Pith. sign in

Paper Citation Record · LEDGER

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction

As of 23 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2603.01530.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.01530 v2

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T19:39:52.897328Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7e745d9b-7458-4b8b-bbe7-c2f906a187bf · outbound

This paper cites Late audio-visual fusion for in-the-wild speaker diarization,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Late audio-visual fusion for in-the-wild speaker diarization,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:47.651812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:47.651812Z digest=sha256:2f64eb06dc3d03160857829c188e582a6eaf0e8287c5db133e6f56528630e302

Observation 63a9d8fc-3871-4b88-9369-966523c65264 · outbound

This paper cites Predict-and-update network: Audio-visual speech recognition inspired by human speech perception,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Predict-and-update network: Audio-visual speech recognition inspired by human speech perception,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:47.707197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:47.707197Z digest=sha256:621549255aaaeb0a7490918a22227d5536ff1369bb937a2a7d31a2abcbbd5748

Observation cf4cf079-4ba3-4aa5-88d6-1ce30a822961 · outbound

This paper cites Audio-visual cross- attention network for robotic speaker tracking,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Audio-visual cross- attention network for robotic speaker tracking,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:47.781204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:47.781204Z digest=sha256:1b94829b5fd6de4ba78948d47d635816110ae9932921fdcd02a9c2342499cd46

Observation 53e1d0f9-1ca6-4203-9034-30adf9fb679f · outbound

This paper cites My lips are concealed: Audio-visual speech enhancement through obstructions,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction My lips are concealed: Audio-visual speech enhancement through obstructions,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:47.881501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:47.881501Z digest=sha256:cde43b547a4ea8a8cd5ea4d2eb631f9832f195fd0fb7fc0fc44946dd5d0e419d

Observation 987c4fe5-f154-4815-bf84-06d766542a17 · outbound

This paper cites Multimodal attention fusion for target speaker extraction,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Multimodal attention fusion for target speaker extraction,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:47.988063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:47.988063Z digest=sha256:23c582abac55cfba8ef7111c3383e7860d204586cfe0d2495f811da2a4c49ae2

Observation bad0a1cd-ae40-4ee0-b7c3-8891513ad009 · outbound

This paper cites Time-domain audio-visual speech separation on low quality videos,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Time-domain audio-visual speech separation on low quality videos,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:48.134760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:48.134760Z digest=sha256:ac4d492e7a966c2183f6e017854c024c83e46342315deaefd498d2857c7355bb

Observation 1bc86261-7476-4377-9659-3a978ba39533 · outbound

This paper cites A two-stage audio-visual speech separation method without visual signals for testing and tuples loss with dynamic margin,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction A two-stage audio-visual speech separation method without visual signals for testing and tuples loss with dynamic margin,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:48.239563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:48.239563Z digest=sha256:a9e5209a1016aa942818c5ba32fa67652af76e96953fd2b391d6136603fddabd

Observation d505beb3-6b6d-4c24-a05d-f3c6fb0c9083 · outbound

This paper cites Ravss: Robust audio- visual speech separation in multi-speaker scenarios with missing visual cues,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Ravss: Robust audio- visual speech separation in multi-speaker scenarios with missing visual cues,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:48.336598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:48.336598Z digest=sha256:d7c56408b05c26b3e7e04f1be0f1381c05dc1dd32c484aa496c04c66601e0e21

Observation 9fdabcf9-e31b-4906-97fe-bfc6b61e1ab9 · outbound

This paper cites Multi-modal multi-correlation learning for audio-visual speech separation,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Multi-modal multi-correlation learning for audio-visual speech separation,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:48.413660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:48.413660Z digest=sha256:fde923bb104f81ed56fbdf19d6a71ce660e8a5b475cc80c48bd7fefb12da5afd

Observation 8568c310-aa1a-4828-827b-73181b449ef7 · outbound

This paper cites Visualvoice: Audio-visual speech separa- tion with cross-modal consistency,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Visualvoice: Audio-visual speech separa- tion with cross-modal consistency,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:48.493817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:48.493817Z digest=sha256:a0f340acbfe3dea4a834774ce4dc5b33c9a572410dcfb3e597126b817fa35e2b

Observation d5ce327e-36a9-4e13-bac5-9532e4eee839 · outbound

This paper cites Explaining face-voice matching decisions: The contribution of mouth movements, stimulus effects and response biases,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Explaining face-voice matching decisions: The contribution of mouth movements, stimulus effects and response biases,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:48.591505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:48.591505Z digest=sha256:e512dad68341011c83b5ae8be4c63988afe6431dc048d0de237beb59783cd7fb

Observation 9bc51f8a-9fb8-49dd-b95a-b842d2d15a59 · outbound

This paper cites Rethinking the visual cues in audio-visual speaker extraction,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Rethinking the visual cues in audio-visual speaker extraction,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:48.661921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:48.661921Z digest=sha256:2a0fc5d29f39bb50bf0fb2c4e07758bb15a75e6261dce29ecfb53efdc85cced3

Observation f1cd3764-6ff2-4e55-bd54-65183bec8189 · outbound

This paper cites Muse: Multi-modal target speaker extraction with visual cues,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Muse: Multi-modal target speaker extraction with visual cues,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:48.701570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:48.701570Z digest=sha256:ab26b1284a769edc6c5e0142213ab4d8203740e7805fc99a9696767afb88e274

Observation e85590b2-679c-4c45-849d-ca11468bc116 · outbound

This paper cites Hearing lips and seeing voices,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Hearing lips and seeing voices,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:48.807239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:48.807239Z digest=sha256:5af7cb1394463d6976f3bef76fce46de9ac4db9cdfd7990e1562f4d671efb1d4

Observation 82f601e2-676a-44f0-a97b-9ed29faa003a · outbound

This paper cites The effect of speechreading on masked detection thresh- olds for filtered speech,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction The effect of speechreading on masked detection thresh- olds for filtered speech,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:48.904251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:48.904251Z digest=sha256:4bfecaaa5db4bcfef7737e791128936dbe9edb398b38421f51e848ac512c7f9e

Observation 85a12557-dfd5-4976-a272-f45ea94cdfbd · outbound

This paper cites Neural target speech extraction: An overview,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Neural target speech extraction: An overview,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:48.980472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:48.980472Z digest=sha256:2bea75106e05f436b2e11455e3229844cf0d3970588187c29aeb5f61477db131

Observation a04400b8-a94e-4b62-9326-6d13188619c0 · outbound

This paper cites Multi-cue guided semi-supervised learning toward target speaker separation in real environments,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Multi-cue guided semi-supervised learning toward target speaker separation in real environments,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:49.048817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:49.048817Z digest=sha256:5d9c7f98547821c58a28246dabaccbe562f005925129091fd04bbc1931e1f8df

Observation 0e4e957a-21e2-49a6-b2d8-f4e8c1e28713 · outbound

This paper cites Multi- level speaker representation for target speaker extraction,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Multi- level speaker representation for target speaker extraction,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:49.105895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:49.105895Z digest=sha256:a916573523bddb873c0564a302ec8ff9ac4c6788a31cb7e6d71608e0c1920bc3

Observation 88058dec-30e4-45d3-add3-c9134998d60e · outbound

This paper cites Usef-tse: Universal speaker embedding free target speaker extraction,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Usef-tse: Universal speaker embedding free target speaker extraction,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:49.191132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:49.191132Z digest=sha256:f660ddf1d9f659c7f285a675dddb6b00313ae0d50b252af6974249fa517d3041

Observation 4f2b76a7-4c7e-4336-bb74-c2254bf92e91 · outbound

This paper cites Contextual speech extraction: Leveraging textual history as an implicit cue for target speech extraction,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Contextual speech extraction: Leveraging textual history as an implicit cue for target speech extraction,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:49.277012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:49.277012Z digest=sha256:f6c0e59e1897bf484026a4caabbf9ad110a6206d2d18560f34390774c3a4cdd5

Observation 380459a4-3d15-4e84-88b7-4e652e6ac1cf · outbound

This paper cites Conceptbeam: Concept driven target speech extraction,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Conceptbeam: Concept driven target speech extraction,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:49.344715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:49.344715Z digest=sha256:0442af7c388017d477e99ec161dc6481032e62ac16aa6408198efb87372706a2

Observation 09c71ba3-7a9c-46af-b636-f08cbbeefda8 · outbound

This paper cites SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:49.452897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:49.452897Z digest=sha256:f71bd2997e9005bdefb25932c42ce22283ec3372826950ecb97b4494003bc993

Observation 1943fd70-11df-4ce3-82f5-4ec69984d198 · outbound

This paper cites Av-crossnet: An audiovisual complex spectral mapping network for speech separation by leveraging narrow-and cross-band modeling,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Av-crossnet: An audiovisual complex spectral mapping network for speech separation by leveraging narrow-and cross-band modeling,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:49.559335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:49.559335Z digest=sha256:76bc5576e9e103ea15010a4e84503ea0d50b994a806e3a9e86bc016de604d072

Observation 0e469bc5-a982-4c66-8de4-f280c8144b95 · outbound

This paper cites Audio-visual speech separation and dereverberation with a two-stage multimodal network,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Audio-visual speech separation and dereverberation with a two-stage multimodal network,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:49.608480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:49.608480Z digest=sha256:d78b4a76a03d8e9e79805aa9dbb1c6f7dea37f588064bf48895f68ba39421f90

Observation cfef3ce0-3b85-4080-be40-26b12408f5de · outbound

This paper cites Spatialnet: Extensively learning spatial information for multichannel joint speech separation, denoising and dereverberation,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Spatialnet: Extensively learning spatial information for multichannel joint speech separation, denoising and dereverberation,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:49.722228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:49.722228Z digest=sha256:cb3607e20181aa665085e0d5315a2fad43355abed494497af73d8ba5567a45ff

Observation 04a2a2f6-a6b0-4183-8185-b3d3a328c608 · outbound

This paper cites Fasnet: Low- latency adaptive beamforming for multi-microphone audio processing,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Fasnet: Low- latency adaptive beamforming for multi-microphone audio processing,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:49.781277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:49.781277Z digest=sha256:647b5e1fab9c45d654fdf38952493a08a5c03d5b20ee3734e27d146a8ff823e6

Observation 660cbdaa-f9f3-47ac-99ba-357a0294e03c · outbound

This paper cites Multi-modal multi-channel target speech separation,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Multi-modal multi-channel target speech separation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:49.841882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:49.841882Z digest=sha256:1635508eab901b0774787aa1ff469aa5ee1a11fc7f5bf0c31e594f824d5525a0

Observation 0e4f0049-db7f-4fe2-b9c2-759f3fc98d65 · outbound

This paper cites An overview of deep-learning-based audio-visual speech en- hancement and separation,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction An overview of deep-learning-based audio-visual speech en- hancement and separation,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:49.952505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:49.952505Z digest=sha256:e98b1a4c3d3f135224429c8471434ef1aa72cbd5abaf3f0303565c1bafe7a509

Observation 6ac8057b-44e4-40b1-a379-571d7678c9d6 · outbound

This paper cites Unified audio visual cues for target speaker extraction,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Unified audio visual cues for target speaker extraction,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.040323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.040323Z digest=sha256:018029c140d1845a7a3ba795599736d24974e32cd634fcccbd1a5bc056540ba4

Observation 7e7e147e-96ba-4759-a544-6ae1f7d213e7 · outbound

This paper cites MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual Cues

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.132135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.132135Z digest=sha256:ce5e5547c4d207264555be0443c1accd170968ce926b78e1354a5718c58ba254

Observation a093622f-23ed-4eee-b9d4-df82da779cc2 · outbound

This paper cites MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction MeMo: Attentional Momentum for Real-time Audio-visual Speaker Extraction under Impaired Visual Conditions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.277622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.277622Z digest=sha256:8379d067dcbcee4943e353228cd97ea8c48c580b7466477cee871d330dcc0d3d

Observation 9236990d-7209-4cd7-bfe7-6c87dcfc5928 · outbound

This paper cites Tf-gridnet: Making time-frequency domain models great again for monaural speaker separation,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Tf-gridnet: Making time-frequency domain models great again for monaural speaker separation,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.421972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.421972Z digest=sha256:3373c5c954d4958d5e1dff6fa646893ba476eb69e66ebeb69774ced08efe2c08

Observation 7691e0c9-511a-47ff-8a85-d36e2f90fdbc · outbound

This paper cites Separate in the Speech Chain: Cross-Modal Conditional Audio-Visual Target Speech Extraction.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Separate in the Speech Chain: Cross-Modal Conditional Audio-Visual Target Speech Extraction

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.510822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.510822Z digest=sha256:64d346644f2fc3a47db94599c9777367f8f5e28f4585261ec4721d3f9ccef943

Observation 02546d07-92d4-4255-a078-45e13b8b6517 · outbound

This paper cites ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction ClearerVoice-Studio: Bridging Advanced Speech Processing Research and Practical Deployment

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.604200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.604200Z digest=sha256:6032356c1b048a65f1260b17c579ccd9d6fc55faa996984ccc972674718bcdb2

Observation 6f81324b-2b19-405e-8ff6-4eefc7fed8ef · outbound

This paper cites Av-sepformer: Cross-attention sepformer for audio-visual target speaker extraction,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Av-sepformer: Cross-attention sepformer for audio-visual target speaker extraction,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.666512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.666512Z digest=sha256:7aec0e37d8c40cd9ab0400fd3a278bf7291a0894f7b596dc082c11ec1af69160

Observation 42cae12f-b223-4ce4-aed9-ad366c7f16b2 · outbound

This paper cites Deep residual learning for image recognition,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Deep residual learning for image recognition,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.710240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.710240Z digest=sha256:3cd86f291b22089b1385ba46b4edc256316ca9ad335753e11223f031046e985c

Observation 147f996f-0f6c-44d0-8946-9b8264e2aba8 · outbound

This paper cites Audio-visual target speaker extraction with selective auditory attention,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Audio-visual target speaker extraction with selective auditory attention,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.758566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.758566Z digest=sha256:51cf70c96a550d68e21f69c5e3d2922ce063b115b940de148e0df38b66e3b81f

Observation 9e040b87-7754-4ebd-ab87-558186acd7c4 · outbound

This paper cites Pitch range variations improve cognitive processing of audio messages,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Pitch range variations improve cognitive processing of audio messages,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.817421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.817421Z digest=sha256:e55111ddea06fc209cf83cf70d4c2c0fe5a2923b9adcf462d2ee791f8a991dcc

Observation 83d5ffcf-441a-4334-9951-63ea1fc05a65 · outbound

This paper cites Intonation and speaker identification,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Intonation and speaker identification,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.878247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.878247Z digest=sha256:d3541b3514f3e3ff7dc46d4e3b83510fe797dee7e5398f697cd82670d7e2ffbd

Observation a3c7590b-37bd-4fd8-b8f2-99b6028d3f0e · outbound

This paper cites Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Dual-path rnn: efficient long sequence modeling for time-domain single-channel speech separation,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:50.973379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:50.973379Z digest=sha256:3652ac2bec6ee34745954c620aa5c9d1ea7830c0bf5c6355c4d9f0ce9b4f96b5

Observation 3bde35da-e50b-4f1e-a4a4-b0cea45ceb1b · outbound

This paper cites Time domain audio visual speech separation,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Time domain audio visual speech separation,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:51.116351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:51.116351Z digest=sha256:b21ca498b6c886f464750b9607fdf51bd9e6c04ad6cc04c35126ec8a888d2adc

Observation 861f072b-09f5-4af3-b854-a775d81c64cd · outbound

This paper cites Learning audio- visual speech representation by masked multimodal cluster prediction,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Learning audio- visual speech representation by masked multimodal cluster prediction,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:51.297240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:51.297240Z digest=sha256:dfc5990dbae6d39a81c072f80268762ca8a382bfcac17a572cb71e91e849dd14

Observation 2ec3e718-24ea-4b47-a486-cf2ee07409a7 · outbound

This paper cites Rethinking the inception architecture for computer vision,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Rethinking the inception architecture for computer vision,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:51.421789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:51.421789Z digest=sha256:90412677f424f66f3a5b0fa6508912e04384f0ee61c7aff0a0021aa829321d76

Observation 99af4452-84f4-4a29-8fc8-7208625e9b3d · outbound

This paper cites Sdr–half-baked or well done?.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Sdr–half-baked or well done?

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:51.538493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:51.538493Z digest=sha256:acdc457a20dbac993decc5199eb0a6a3c893ffee89cbdcf58dcda34a8b8e1f60

Observation 2e4cbf8c-bce6-4d7c-8181-af5e0fd61cd4 · outbound

This paper cites Audio-Visual Speech Separation in Noisy Environments with a Lightweight Iterative Model,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Audio-Visual Speech Separation in Noisy Environments with a Lightweight Iterative Model,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:51.693426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:51.693426Z digest=sha256:ae536a3658fbf79a4039b99415b4b2c4626b2b72cc0d1253073bbacc44941eb2

Observation 3c3d8a28-7000-4038-adf8-e588e3a7538f · outbound

This paper cites An audio-visual speech separation model inspired by cortico-thalamo-cortical circuits,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction An audio-visual speech separation model inspired by cortico-thalamo-cortical circuits,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:51.835676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:51.835676Z digest=sha256:9e4e95e20a067fa84548ee492c37817e9fb32f503747ecdd0fce7d2f37bfe856

Observation 514335a0-d76a-4404-8576-ecdf84ed0ece · outbound

This paper cites Iianet: an intra-and inter-modality attention network for audio-visual speech separation,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Iianet: an intra-and inter-modality attention network for audio-visual speech separation,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:51.902230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:51.902230Z digest=sha256:9f35f5e229ce6e38117284c613f552185223fc3aa0774b73361c5c10086c59df

Observation b78ba67f-4f44-4135-abdf-d3d62d824dd6 · outbound

This paper cites Deep audio-visual speech recognition,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Deep audio-visual speech recognition,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:51.948455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:51.948455Z digest=sha256:3fff73adc47e08c2ec132fda2eac3f3d6757251239542cc05f143cc795b90265

Observation a667534c-8589-45d5-9bf7-6a3ec0025dae · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction VoxCeleb2: Deep Speaker Recognition

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:52.095802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:52.095802Z digest=sha256:886a0812943ab5d7f48637bb56b93f917df08e01e90b912479e2a5ef76d7eb29

Observation de8584db-4a6a-46fa-ba23-6aa2417334f6 · outbound

This paper cites Looking into your speech: Learning cross-modal affinity for audio-visual speech separation,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Looking into your speech: Learning cross-modal affinity for audio-visual speech separation,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:52.175405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:52.175405Z digest=sha256:204a9e73c3f3661c94c5d258247c24402e7d05b666a93f145fddd108f9ee3f72

Observation a7256a29-e690-49c8-b956-a3dd7f0a2d5b · outbound

This paper cites Watch or listen: Robust audio- visual speech recognition with visual corruption modeling and reliability scoring,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Watch or listen: Robust audio- visual speech recognition with visual corruption modeling and reliability scoring,

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:52.273806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:52.273806Z digest=sha256:f0382bca71d30ac06ca833783dab03fa1110081f46312c319da8615c300f3de6

Observation b334d924-65e3-47f9-b207-5d84923d7c9f · outbound

This paper cites Restoring speaking lips from occlusion for audio-visual speech recognition,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Restoring speaking lips from occlusion for audio-visual speech recognition,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:52.379271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:52.379271Z digest=sha256:a7d434680c7d89c1b13e687770aed6cc76caa766e9c7c1d3b446655da2c7a250

Observation dd4a8bae-a187-420a-a7a4-4568218dfee1 · outbound

This paper cites Delving into high-quality synthetic face occlusion segmentation datasets,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Delving into high-quality synthetic face occlusion segmentation datasets,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:52.500142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:52.500142Z digest=sha256:7cf93d99b69819c650c73f9b0d79899a056270dc1928481a9239e5ef113c1893

Observation e34e8283-d8b2-4aab-9303-c6a80775876a · outbound

This paper cites Data2vec: A general framework for self-supervised learning in speech, vision and language,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Data2vec: A general framework for self-supervised learning in speech, vision and language,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:52.616289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:52.616289Z digest=sha256:9a2b9c716ca661b50ff9a2ccf8adda8192c0ad6c68437ec8dd68313a8ee5144c

Observation 82b83497-9b72-49fa-b5a1-0ebe40e3426a · outbound

This paper cites Adam: A Method for Stochastic Optimization.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Adam: A Method for Stochastic Optimization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:52.723885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:52.723885Z digest=sha256:19b785275e4a5984c8e102369bcc2fadd958b395281911e53e4479ed68b69dfa

Observation cd1737d7-935f-4925-954b-b90a960cda81 · outbound

This paper cites Long short-term memory,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction Long short-term memory,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:52.814970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:52.814970Z digest=sha256:fea95596e842b2e9c47a202a5aaaad2858322bf82354c9830b39e07f6bbd842f

Observation 0edb5e94-dfa7-49ff-8040-5d7d0f070bfc · outbound

This paper cites How many phonemes does the english language have,.

CueNet: Robust Audio-Visual Speaker Extraction through Cross-Modal Cue Mining and Interaction How many phonemes does the english language have,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T19:39:52.897328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:39:52.897328Z digest=sha256:d3c703b1c0c7e4b783a5d56de4ac275b27ce8e31cd3bb9b70e1f60ad0a86dcb2

Pith citing papers

No inbound Pith citation observations are available.