Pith. sign in

Paper Citation Record · LEDGER

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation

As of 11 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2501.01518.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01518 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:30:41.135038Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact3
  • verified fuzzy42
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0ab3960e-a3fe-479b-b16c-aebb3609c7a2 · outbound

This paper cites The conversation: Deep audio-visual speech enhance- ment.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation The conversation: Deep audio-visual speech enhance- ment

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.824149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.906332Z digest=sha256:78a1419c3c1bb21eeb367232d19f75c93ea72cc59e3295281b4ead2e81bd7601

Observation b5d5e497-055a-4040-86ad-48fa0db08193 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation LRS3-TED: a large-scale dataset for visual speech recognition

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:40.911436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:40.911436Z digest=sha256:564d27616a96f409de545585702919a4dd1cc95466c0f6728dd9d95ea104467e

Observation 38034054-41b9-483e-9656-bf421910de83 · outbound

This paper cites My lips are concealed: Audio-visual speech enhance- ment through obstructions.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation My lips are concealed: Audio-visual speech enhance- ment through obstructions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.813131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.915374Z digest=sha256:7200b550fbc881ce0fbec8b6ad755393485d4677e8acaead7b1c50e18d0f6b38

Observation d99466b3-8a37-406c-aa62-f2b7d2139348 · outbound

This paper cites Self-supervised learning of audio-visual objects from video.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Self-supervised learning of audio-visual objects from video

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.802130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.919387Z digest=sha256:14bb3e964eb72f51655f98e8c21aafbc8c7e38dd30f19b7df7f0be07d125b93b

Observation db76f7ff-7fb8-4c11-b568-f39062179782 · outbound

This paper cites LipNet: End-to-End Sentence-level Lipreading.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation LipNet: End-to-End Sentence-level Lipreading

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:40.923188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:40.923188Z digest=sha256:6a8aa7706b2e442423151ab341777425e04788f7df6d962b3e77b30f1b643641

Observation 9b8835f8-a4dd-48dd-b9dc-840c56951f08 · outbound

This paper cites Phonemizer: Text to phones transcription for multiple languages in python.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Phonemizer: Text to phones transcription for multiple languages in python

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.791554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.927322Z digest=sha256:ba80790cf57fcb677bb3f3925ad7b3279b6acf6537b2dc99409131e67f3095a6

Observation fd460ee0-f4d2-429f-936d-921a0843a975 · outbound

This paper cites Audio-visual synchronisation in the wild.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Audio-visual synchronisation in the wild

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.780893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.931502Z digest=sha256:ec1cf11c5b382fe1d4822d6c817200f4e1b1f373c79d1c0b627fd3e0b4f7c61d

Observation afafdaec-dce2-499d-9b34-ac6934cf9b75 · outbound

This paper cites Deep attrac- tor network for single-microphone speaker separation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Deep attrac- tor network for single-microphone speaker separation

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.770243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.934940Z digest=sha256:1d5f67b7008874ae4cd3b14b50fc2884ae4397655f0b5b999ea76f9e3d136fe5

Observation 027513d9-d655-4d13-96b1-1b3ddbb520d1 · outbound

This paper cites Lip reading sentences in the wild.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Lip reading sentences in the wild

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.759334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.938885Z digest=sha256:6c77ab83c0eff65dbf37a68350fee7d4e5d711cefb6a00dea64058f15fbaed60

Observation dcd99952-8200-4af1-b432-b440c346d91a · outbound

This paper cites Lip reading in the wild.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Lip reading in the wild

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.749567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.942276Z digest=sha256:c6c6b93718e9fe18c44a6a4666eef05ff04a5fde51b741f538c6282675494d72

Observation 1f66ee20-3968-4ac0-9620-6a4d2d0c9e1f · outbound

This paper cites FaceFilter: Audio-visual speech separation using still images.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation FaceFilter: Audio-visual speech separation using still images

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:40.946090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:40.946090Z digest=sha256:d8466243e863f8ef73ea68dd28c54a7726b5f01a79db72961a1c98ea05a71ab2

Observation 57efc268-dfbc-4605-97cc-bc7a7ed3d7bf · outbound

This paper cites Real time speech enhancement in the waveform domain.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Real time speech enhancement in the waveform domain

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.738937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.949963Z digest=sha256:30d0ec63c599027b1a4d9e2d6d1e4f5e1a3546916da8cf59fbf290393f319875

Observation 7181ccb4-2d5e-4513-bc3d-2092414e194a · outbound

This paper cites Music Source Separation in the Waveform Domain.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Music Source Separation in the Waveform Domain

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:40.953728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:40.953728Z digest=sha256:69bf9516c40fb70b7e766d0ff03dcb57373d4e53b61c7744494960cd9a74f64e

Observation b2f14a6a-a786-4146-88cb-40414f437063 · outbound

This paper cites Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:40.957623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:40.957623Z digest=sha256:a0c4084c040925af4449a6aaffd4aab745f61bf92d6966672e42dfbbf868b8c6

Observation 4b307e38-25a9-40ea-bee5-f689fd46e2ba · outbound

This paper cites Learning joint statistical models for audio- visual fusion and segregation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Learning joint statistical models for audio- visual fusion and segregation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.726885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.961532Z digest=sha256:3656d1a88d398757791f918a0d411c5bf608a89513b46c6c18f6299eaab967ae

Observation 9a87c14a-fab0-42d1-b323-d0f98a8f47da · outbound

This paper cites Seeing through noise: Visually driven speaker separa- tion and enhancement.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Seeing through noise: Visually driven speaker separa- tion and enhancement

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.715075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.965596Z digest=sha256:8cb1ebedb7f0bfbecc86d4bf465ae32d06ce0406e3137f11e05176872c0e84a1

Observation 14e843c0-fdcf-4304-bb15-488a77c03f41 · outbound

This paper cites Visual Speech Enhancement.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Visual Speech Enhancement

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:40.969620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:40.969620Z digest=sha256:ca2fa6064a33f7459920c2125351ede30aafd7b1a4e245f7eb99dd9857c3a899

Observation 47642839-a43a-4d17-b1c7-4f6bd9bfce5d · outbound

This paper cites Tenen- baum, and Antonio Torralba.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Tenen- baum, and Antonio Torralba

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.703401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.973935Z digest=sha256:4e0d1ef671a077e81c9ae09739b1bc712d14fc36a9a2e9371d19d91a811594f3

Observation af3765e5-7ee4-4e70-be0b-70631630a711 · outbound

This paper cites Learning to Separate Object Sounds by Watching Unlabeled Video.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Learning to Separate Object Sounds by Watching Unlabeled Video

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:30:41.243456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.977615Z digest=sha256:9caa8e371274f6db30af5ec62e564b9a2382e2e8cd3702c9f1c49756c0772756

Observation 043b39e3-2d44-46ed-aec7-e68bb17ab545 · outbound

This paper cites Co-Separating Sounds of Visual Objects.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Co-Separating Sounds of Visual Objects

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:30:41.228155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.981589Z digest=sha256:081e15e4e6ad2a47655c6e59e858f90c955b25609730d4c942ced4477dc91321

Observation ea8cc40b-b7da-4e9e-a01d-f9f7e666a425 · outbound

This paper cites VisualV oice: Audio- Visual Speech Separation with Cross-Modal Consistency.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation VisualV oice: Audio- Visual Speech Separation with Cross-Modal Consistency

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.692101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.985549Z digest=sha256:365a1cf1f16a0ef8aa1390a88dca6e44fca18de20e4b0e5b58712b047a7f0e5c

Observation 95c0dff2-817e-4b81-a01c-9447e9188322 · outbound

This paper cites Multi-modal multi-channel tar- get speech separation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Multi-modal multi-channel tar- get speech separation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.681770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.989309Z digest=sha256:f28cd91b8df0a6ce217db1c89158a4e7262ad9244f24733124a0373b3f23ed42

Observation 895d562e-4cf0-488d-ac62-931e5016a1db · outbound

This paper cites Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Hershey, Zhuo Chen, Jonathan Le Roux, and Shinji Watanabe

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.670976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:40.993004Z digest=sha256:cab9fb03f49e0b20366fab84a110d86bf8261236de7ef40c7f9147ee670cf15c

Observation a6536841-966a-446e-9df7-b7a1a40b160d · outbound

This paper cites Perceiver: General Perception with Iterative Attention.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Perceiver: General Perception with Iterative Attention

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:40.996705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:40.996705Z digest=sha256:49cddd10edde7e4b076ecedb7eaa247108a7e5e2f00af77b3781acd90b6b3812

Observation 29573d01-f545-40e5-97fa-0338a57339cb · outbound

This paper cites Looking into your speech: Learning cross-modal affinity for audio-visual speech sepa- ration.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Looking into your speech: Learning cross-modal affinity for audio-visual speech sepa- ration

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.659982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.000900Z digest=sha256:57509a8150004182641e6c36f6cc4a9d9885236f7a597a4f605e0d5213ce8626

Observation b5081915-2b49-492b-be21-8e7cefaff36f · outbound

This paper cites Parameter efficient multimodal trans- formers for video representation learning.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Parameter efficient multimodal trans- formers for video representation learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.650080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.005406Z digest=sha256:b64dcb426e50f2096e44e44b9d1c5d9c09e577f3cb5e33e8dd035fc1a2034397

Observation 16cd0da0-dcf0-4149-be3f-f141833517bd · outbound

This paper cites Audiovisual trans- former with instance attention for audio-visual event local- ization.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Audiovisual trans- former with instance attention for audio-visual event local- ization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.638710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.009871Z digest=sha256:7421704afc1737a821b11a321119f1e0407b6d8aee4a35ee9894645ea9daddc3

Observation daeb5589-62dd-47b6-97e2-52ef461facb0 · outbound

This paper cites Speaker- independent speech separation with deep attractor network.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Speaker- independent speech separation with deep attractor network

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.627331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.013675Z digest=sha256:6e106e17799ca9e0351b2f55363250c0ae9a3cbc24733ddac96c850fafc1e97d

Observation fa8a0789-3ef8-4ad0-b9fa-870f2b0d0139 · outbound

This paper cites Attention in dichotic listening: Affective cues and the influence of instructions.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Attention in dichotic listening: Affective cues and the influence of instructions

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.614029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.017791Z digest=sha256:9d5fb47d1bd401ee8c3757427fe542ca021ad303bb280284c837a7b42f464abe

Observation 0027f422-4a53-40ae-83c9-28e7c7e097e5 · outbound

This paper cites Attention bottlenecks for multimodal fusion.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Attention bottlenecks for multimodal fusion

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.602676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.021525Z digest=sha256:1bd8e417d340a14137db4d35bace2d6e8b6d45274ee64013872c1d6bd1b7fe4a

Observation 96797977-ada7-4f82-9afd-12465ab51e46 · outbound

This paper cites an unresolved cited work.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:30:41.591631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.026076Z digest=sha256:f612404a509cb25862e13adceee6b33653045fd03ae54924baea8adef6ada1e8

Observation 5dbec9bf-5987-4585-93b1-1f46c80d6a0b · outbound

This paper cites an unresolved cited work.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:30:41.580932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.031979Z digest=sha256:5054ef0354c0ea3c4cba8c47b5bfbc9be6c81fd371e89cbf5209d21a9d26eb96

Observation 1b99f204-5e96-45ae-be0d-f747a5de7cff · outbound

This paper cites Audio-visual object localization and separation using low- rank and sparsity.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Audio-visual object localization and separation using low- rank and sparsity

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.569605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.035764Z digest=sha256:18d3e3abaa532722fa1cd7455b91148be3c31766f8cc55a64b292a0bfb3474fa

Observation 2ab1403e-7f12-4380-87af-c1df667eb2ee · outbound

This paper cites mir eval: A transparent implementation of common mir metrics.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation mir eval: A transparent implementation of common mir metrics

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.558567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.039549Z digest=sha256:2004ce495f3c8cd7d2bf190e3eb0305644f738d3ed28ccbad718f5a69b6d59da

Observation f89505bb-77e7-4f4f-913d-95f3628267a9 · outbound

This paper cites Interspeech 2021 deep noise suppression challenge.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Interspeech 2021 deep noise suppression challenge

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.547587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.043823Z digest=sha256:721eb4b51b91c592f94c122aa242019b763ffd1553b7b3f8f3c92a27c0929316

Observation 1605df20-c4fb-4ef3-9265-94cc66d4dc32 · outbound

This paper cites Visual keyword spotting with attention.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Visual keyword spotting with attention

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.535480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.048654Z digest=sha256:2b8c58096df9f3019d433c8b9c93fe0c51ae3ec91fb376fae98c7629c36a6351

Observation 98ba1576-386b-4233-8484-dee668835509 · outbound

This paper cites Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of tele- phone networks and codecs.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Perceptual evaluation of speech quality (pesq)-a new method for speech quality assessment of tele- phone networks and codecs

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.523260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.052999Z digest=sha256:e2599e0e00f83befa8890c29461f0979532755d6604312e92ae7c3ed85e56ba3

Observation 7bf38965-c9ea-4739-9ec7-3dfa4869b105 · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation U- net: Convolutional networks for biomedical image segmen- tation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:41.056750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:41.056750Z digest=sha256:dd4de3af6709b7eb1677b84959af6e1c6e3caeaca523249a61fef89b0cf6788a

Observation b46a5225-02d9-431c-8c60-1da3672a0ab8 · outbound

This paper cites Self-supervised audio-visual co-segmentation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Self-supervised audio-visual co-segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.504725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.060530Z digest=sha256:6ca72cdd9727a28ca006cb841124f1fbf9f556e8a874a9b56279a48f8e9e6305

Observation 8439aeb5-1892-4af5-93e5-8858be541a88 · outbound

This paper cites Audio-visual speech enhancement using conditional variational auto-encoders.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Audio-visual speech enhancement using conditional variational auto-encoders

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.493571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.064735Z digest=sha256:cc8b48d7a88855ea9bc2d7209c09f01583f59062143a52977dee3d3c4db2f9a9

Observation 8958c87f-7317-48ca-8abc-2a206c2c2e68 · outbound

This paper cites Seeing to hear better: evidence for early audio- visual interactions in speech identification.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Seeing to hear better: evidence for early audio- visual interactions in speech identification

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.482804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.068465Z digest=sha256:1e1afabc3afa1fc382566df8c64caa3202162fa3217a6a633b052532d1c88325

Observation 925342c2-ae96-4f7e-b01b-7464fd4eef87 · outbound

This paper cites Combining residual networks with lstms for lipreading.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Combining residual networks with lstms for lipreading

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.471859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.072889Z digest=sha256:09b8b97dd136d3f4f2cc5be18051f37cc5f26c9860a4bc00bc4ba1003a9c9a81

Observation f5d63be5-f809-4ebc-8f76-4623a4669aa5 · outbound

This paper cites An algorithm for intelligibility prediction of time- frequency weighted noisy speech.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation An algorithm for intelligibility prediction of time- frequency weighted noisy speech

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.459504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.076646Z digest=sha256:b9bba0f2c05f399b8cc43aa5c8065b489f1a482333aa0160d7966585c3d4bd3c

Observation bf6c4257-0306-4ed4-b4a4-53236c4e19cb · outbound

This paper cites Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.448453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.080374Z digest=sha256:b3241a74879e6a9258701f16722b8b97678f0bedd3852dc1cac337e6461280ef

Observation 01094ad2-422e-4d84-b36a-b305d6649124 · outbound

This paper cites Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Into the Wild with AudioScope: Unsupervised Audio-Visual Separation of On-Screen Sounds

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:41.084566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:41.084566Z digest=sha256:6603dbb6d7fd8c340947e723a757568e2a7e6feea93f231447241d83dccfc5cf

Observation 70ce8276-f361-4f3d-a88e-9a3837b1bd3c · outbound

This paper cites Improving On-Screen Sound Separation for Open-Domain Videos with Audio-Visual Self-Attention.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Improving On-Screen Sound Separation for Open-Domain Videos with Audio-Visual Self-Attention

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:30:41.188958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.088534Z digest=sha256:4ac0d8dce36fee623bc227d19fa8f5d979ec6775b5dd40cb7475808338de9d48

Observation 6c9a8694-407f-4356-aad8-6d1fba8e3278 · outbound

This paper cites BSS EV AL toolbox user guide.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation BSS EV AL toolbox user guide

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.437238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.092693Z digest=sha256:0a42ad38f01221c817e10f481f63db92383b076a9a88a0808afb27e0317f8916

Observation c4a5be8b-1a04-4b42-8acf-601a24046850 · outbound

This paper cites Supervised speech sep- aration based on deep learning: An overview.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Supervised speech sep- aration based on deep learning: An overview

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.413067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.100737Z digest=sha256:ffedf5ba5780b8033765b7f56f3a5e7a088c9e4c1149deda30d6db813a861130

Observation b40b4959-5bfd-4f45-bca7-24f196a58ba5 · outbound

This paper cites V oicefilter: Tar- geted voice separation by speaker-conditioned spectrogram masking.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation V oicefilter: Tar- geted voice separation by speaker-conditioned spectrogram masking

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.401200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.104155Z digest=sha256:0a5e0077accf815ec6e16e32c8ce31556d560f2ff9eab7ffa85f0ca09bac28fa

Observation 48bcb8d7-55c5-4300-a1c9-f683f67eb681 · outbound

This paper cites Combining spectral and spatial features for deep learning based blind speaker separation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Combining spectral and spatial features for deep learning based blind speaker separation

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.389835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.108114Z digest=sha256:5eba2c273fc136362d3db78e94e9352f054dd70f7ea034a63511f8114b51482d

Observation aec89adc-0dbb-4aa4-80d4-390ae8562903 · outbound

This paper cites Time domain audio visual speech separation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Time domain audio visual speech separation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.377627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.111482Z digest=sha256:c5f0640404d3ba05f20811af1b688f6ca23b3396f955e1634676afcfc1421c38

Observation 46297670-87a4-4164-844e-89649023ac2c · outbound

This paper cites Multilevel language and vision integration for text-to-clip retrieval.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Multilevel language and vision integration for text-to-clip retrieval

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.365307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.115377Z digest=sha256:1f4ca7bffd1e1eeaaf05092428a2e2a77f60dc4bee1a70b67169fb7e36a216aa

Observation 7c021b59-cba2-4afd-b9c2-b8154d58eb59 · outbound

This paper cites Recursive visual sound separation using minus-plus net.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Recursive visual sound separation using minus-plus net

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.354847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.118996Z digest=sha256:4e378452b01a62e17996e58dcdfe0459174b4a54774b6019fa551b10ecbbf1ec

Observation bd574924-a41b-45cd-87e1-d1041bfb2cfe · outbound

This paper cites Permutation invariant training of deep models for speaker-independent multi-talker speech separation.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Permutation invariant training of deep models for speaker-independent multi-talker speech separation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.344301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.123101Z digest=sha256:41357850786ecc01d3266e0b1cce810d4ac35a3898dcabe6eb2d71fa32b37afb

Observation 3a19ff94-23a7-4a91-b9ef-9504a7120d69 · outbound

This paper cites To find where you talk: Temporal sentence localization in video with attention based location regression.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation To find where you talk: Temporal sentence localization in video with attention based location regression

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.332912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.127305Z digest=sha256:9486555fcca8bc407db759e09528dd73707627ad62b5f4b09912ee42e57bd131

Observation 874a4fa0-1153-49f4-9d49-67239dd48ae4 · outbound

This paper cites The sound of motions.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation The sound of motions

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:30:41.321307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.131259Z digest=sha256:64aa93051661847a8bbbf3aadb9b6a0d94c25b20ef9cf0ff7c7b00d63b3f7d5a

Observation de3c0aa3-f74f-45a9-9044-40555c7f02a1 · outbound

This paper cites The Sound of Pixels.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation The Sound of Pixels

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T22:30:41.135038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:30:41.135038Z digest=sha256:ccd9db175d3aa63b12dae8599b2e7af5e2feeee924aa0f64a13f55d3c461b6ef

Observation 6aec83be-3d62-49aa-9182-461b5a26161e · outbound

This paper cites an unresolved cited work.

Reading to Listen at the Cocktail Party: Multi-Modal Speech Separation Unresolved cited work

Reference 1706

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:30:41.425188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T22:30:41.097036Z digest=sha256:f7bb02d2beb62f6efb19be13d5686a5e1cc4794946c4237a401f0d9298572742

Pith citing papers

No inbound Pith citation observations are available.