Pith. sign in

Paper Citation Record · LEDGER

Cocktail-Party Audio-Visual Speech Recognition

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2506.02178.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02178 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:33:52.053922Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact1
  • verified fuzzy25
  • unresolved15
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69629e47-2e64-441f-9e27-0d3ef9f2be0e · outbound

This paper cites an unresolved cited work.

Cocktail-Party Audio-Visual Speech Recognition Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:34:00.070508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:46.680231Z digest=sha256:08a2d5e58c0dd6e4a601a1970ee8dfd92021349ca7a18002c81669dd6b51c58a

Observation 09abe8b3-a228-4a21-9d8a-51e66e20fe00 · outbound

This paper cites Task definition Given an input sequence of audio A = {a1, a2,.

Cocktail-Party Audio-Visual Speech Recognition Task definition Given an input sequence of audio A = {a1, a2,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:59.796475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:46.797891Z digest=sha256:ff75c8bff53ced5369db2d6a0fe275a5ccc0c77da082e636754f3bcc88930538

Observation 89a15ed6-0bee-4b18-b1f9-1b79cf888d84 · outbound

This paper cites For training, we use LRS2 (train and pretrain sets), V ox2 (train set), and A VYT.

Cocktail-Party Audio-Visual Speech Recognition For training, we use LRS2 (train and pretrain sets), V ox2 (train set), and A VYT

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:33:59.322712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:47.064621Z digest=sha256:74307b063fd54b3ec79799a9eff6186370204bde45712eec94cc3ebe6becd016

Observation d388c060-6fb0-47e8-a097-b47628032925 · outbound

This paper cites The A V-HuBERT CTC/Attention (A V1) model uses the A V-HuBERT large [12] as the encoder, which has 24 transformer blocks, each with 16 attention heads.

Cocktail-Party Audio-Visual Speech Recognition The A V-HuBERT CTC/Attention (A V1) model uses the A V-HuBERT large [12] as the encoder, which has 24 transformer blocks, each with 16 attention heads

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:59.076103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:47.194615Z digest=sha256:e132ad434adffba9e24b3492116f8b8375eef6f6ca1635ec5d48fb40e51221f3

Observation 544ba7dc-756e-42e9-b7c2-76383fb4e87e · outbound

This paper cites The WERs for models evaluated on the original LRS2 test set are shown in the column where SNR = ∞.

Cocktail-Party Audio-Visual Speech Recognition The WERs for models evaluated on the original LRS2 test set are shown in the column where SNR = ∞

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:33:58.810431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:47.330476Z digest=sha256:e2b0e15de7f34dbaa16e9a1099b15f80bf1d31e032272da1c0d32aa1deb5554c

Observation 0b47fdb0-1712-40ff-9aec-a3e62004d8d3 · outbound

This paper cites We highlighted the gap between conventional datasets and real-world cocktail- party scenarios, where target speakers are not always ac- tive.

Cocktail-Party Audio-Visual Speech Recognition We highlighted the gap between conventional datasets and real-world cocktail- party scenarios, where target speakers are not always ac- tive

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:58.566966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:47.459176Z digest=sha256:476d9e11b7bbdfc898512e75a5e76437b2c92d6097e774e5f7fcf6002d1cc15a

Observation 8aa6569e-f973-45b0-9770-ef4ae4cfc25b · outbound

This paper cites How is AI Changing Science? Research in the Era of Learning Algorithms.

Cocktail-Party Audio-Visual Speech Recognition How is AI Changing Science? Research in the Era of Learning Algorithms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:58.280233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:47.577534Z digest=sha256:ac43ee940c7b7023709a1499edc5efedd6bf4c942e2156af8990b3fd0ddfdc62

Observation 000f8e70-32b7-4932-bca3-bdb3b8fcf14e · outbound

This paper cites Knowing who to listen to in speech recognition: Visually guided beamforming,.

Cocktail-Party Audio-Visual Speech Recognition Knowing who to listen to in speech recognition: Visually guided beamforming,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:57.228249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:48.648069Z digest=sha256:9b84b62a9308fa0ade97020ed1fd0975caa7fd1290db4bb706afe115070077ad

Observation 49f73fe2-8684-49d0-a265-347a8aa056a7 · outbound

This paper cites Hearing lips and seeing voices,.

Cocktail-Party Audio-Visual Speech Recognition Hearing lips and seeing voices,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:47.738240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:47.738240Z digest=sha256:c645841cfdf9dd9baf99884b69e1aeaca1ce81fabe8838a3ef538f5767eff182

Observation 0e0c50f4-0308-40d5-8c9d-1ad72f60d392 · outbound

This paper cites Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation,.

Cocktail-Party Audio-Visual Speech Recognition Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:47.925263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:47.925263Z digest=sha256:e1de58539f440e2403cdef0ab4ad5989977b5593487c2edc32031b3c8d248c06

Observation 571c6281-057d-4e5e-a4bb-7ac627c59fdc · outbound

This paper cites Super-Human Performance in Online Low-latency Recognition of Conversational Speech.

Cocktail-Party Audio-Visual Speech Recognition Super-Human Performance in Online Low-latency Recognition of Conversational Speech

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:33:52.430647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:48.048385Z digest=sha256:8e5f1802b3b215812b47bf4bafc26be04a247b55a3e33b2c3968ea965f78d44c

Observation 66de8a82-0444-497e-80e7-a82766e5d749 · outbound

This paper cites an unresolved cited work.

Cocktail-Party Audio-Visual Speech Recognition Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:33:59.539666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:46.922739Z digest=sha256:f8b9d2325f866113b7ecbe309a5ac7f230a345100cfab5fb4d3367aa48ebfa7e

Observation e4af78f2-617d-404a-ab3a-0487f7473dba · outbound

This paper cites Recognition of conversational telephone speech using the janus speech engine,.

Cocktail-Party Audio-Visual Speech Recognition Recognition of conversational telephone speech using the janus speech engine,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:58.018024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:48.166647Z digest=sha256:95fa9aeb3313fcf013b60448c4ac7acf918c89ebb5dc476d4f4a675c714f89b6

Observation d0750410-d6e9-4e47-84c5-4340803780ca · outbound

This paper cites See me, hear me: inte- grating automatic speech recognition and lip-reading,.

Cocktail-Party Audio-Visual Speech Recognition See me, hear me: inte- grating automatic speech recognition and lip-reading,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:48.303473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:48.303473Z digest=sha256:c16a9743b8a596b73f3934ceeeeee661006b705f4159547cd6a717f146d30fb2

Observation 4d39485d-cc8e-4944-80d8-c594e9dc7e1c · outbound

This paper cites Multimodal interfaces,.

Cocktail-Party Audio-Visual Speech Recognition Multimodal interfaces,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:57.723392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:48.418256Z digest=sha256:119ab4fb04036a280b892c9021b5af231a1c41b0af0bd71cf2691614decac848

Observation a64442cf-b5ec-473d-8a67-87295b2d83c8 · outbound

This paper cites Modeling focus of at- tention for meeting indexing,.

Cocktail-Party Audio-Visual Speech Recognition Modeling focus of at- tention for meeting indexing,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:57.500393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:48.541362Z digest=sha256:79c8a1f28b8260590ea907b63ebcefa2300c6d6443d7efd0a9426c41aec6c67c

Observation 895b5d0d-c927-4c2e-9035-fa9af05703ee · outbound

This paper cites Visual track- ing for multimodal human computer interaction,.

Cocktail-Party Audio-Visual Speech Recognition Visual track- ing for multimodal human computer interaction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:56.972059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:48.752559Z digest=sha256:328c3f75d6524d2099b5e31e6616c44a0d7c68b954015c991548901c549b110a

Observation 2fb81885-5478-49af-a611-a35f13ec12ad · outbound

This paper cites Estimating focus of attention based on gaze and sound,.

Cocktail-Party Audio-Visual Speech Recognition Estimating focus of attention based on gaze and sound,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:56.728640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:48.882045Z digest=sha256:043e4d4fe8f6c95e6e7dcf760b65447ecc956663c478f7ac13f1aa4831620e51

Observation 449b42b3-e185-4d2f-8aa5-487dbb195c09 · outbound

This paper cites Chil: Computers in the human interaction loop,.

Cocktail-Party Audio-Visual Speech Recognition Chil: Computers in the human interaction loop,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:56.459235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:49.010398Z digest=sha256:be3478de84b3ceae7723d9d060342b4b686a9fa946754e811844564cfc8e8862

Observation 2d352c00-65e7-45db-9210-f041cb23e2c1 · outbound

This paper cites Robust self-supervised audio-visual speech recognition,.

Cocktail-Party Audio-Visual Speech Recognition Robust self-supervised audio-visual speech recognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:56.121810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:49.140475Z digest=sha256:cc49fa565841fe038c78a4122cde9df1af8911632140bf566ff622d6254a26eb

Observation 417a4001-a7f1-4c34-bba2-e274b8007a83 · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels,.

Cocktail-Party Audio-Visual Speech Recognition Auto-avsr: Audio-visual speech recognition with automatic labels,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:55.873050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:49.233772Z digest=sha256:1145a9e9c90486de8fcd0977cada5af32adc246af47c292e2c1247c12d9b5872

Observation c68e33f3-e420-4d17-8582-18862c7170c1 · outbound

This paper cites Whisper-flamingo: Integrating visual fea- tures into whisper for audio-visual speech recognition and trans- lation,.

Cocktail-Party Audio-Visual Speech Recognition Whisper-flamingo: Integrating visual fea- tures into whisper for audio-visual speech recognition and trans- lation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:49.361747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:49.361747Z digest=sha256:2ea3219684659ff770cdb2e6441bbae5363cf84b9b0238c89419313fd6b4768d

Observation ce23936d-fff7-4fcc-8e82-78f2229ee0c5 · outbound

This paper cites Speaker-targeted audio-visual models for speech recognition in cocktail-party environments,.

Cocktail-Party Audio-Visual Speech Recognition Speaker-targeted audio-visual models for speech recognition in cocktail-party environments,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:55.584216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:49.466179Z digest=sha256:2cfb028c0412c97de49dc6bd52553c718b2d08cc8f8c3f07bd63a43f4276ea9d

Observation 616c2b1a-4dc2-409b-b6f2-cbab93fdb9a8 · outbound

This paper cites Audio-visual multi-talker speech recognition in a cocktail party,.

Cocktail-Party Audio-Visual Speech Recognition Audio-visual multi-talker speech recognition in a cocktail party,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:55.337694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:49.588319Z digest=sha256:c0fe72aa0c5fb80767807dd70700a19ab2f58281d8ea0e722bc608958b701714

Observation 5412a3ce-84d3-4df4-9229-aed1624a57c4 · outbound

This paper cites Robust audio-visual asr with unified cross-modal attention,.

Cocktail-Party Audio-Visual Speech Recognition Robust audio-visual asr with unified cross-modal attention,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:55.055313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:49.736454Z digest=sha256:063460d6f8db1e43d1f47265e051ecdb32eddc5d47b53a60f60ee121e7eb564b

Observation a4257471-b8c4-48c8-98e7-4c55494e0247 · outbound

This paper cites Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation,.

Cocktail-Party Audio-Visual Speech Recognition Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:49.876226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:49.876226Z digest=sha256:14311241428a8484f934fa16517f282edf3b5ae9804941ee6def629e8b29fdcf

Observation 86c0b254-0e57-4c40-aea5-d2534ecd2d31 · outbound

This paper cites Visualvoice: Audio-visual speech sep- aration with cross-modal consistency,.

Cocktail-Party Audio-Visual Speech Recognition Visualvoice: Audio-visual speech sep- aration with cross-modal consistency,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:54.796426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:50.019620Z digest=sha256:11d4a1972bb6a1f1d3ecfac333c2c5b9d786291bff1cee75f4f61e482c402c90

Observation a96f7f83-0fed-4d3a-b55b-31f2ebeec776 · outbound

This paper cites Seeing through the conversation: Audio-visual speech separation based on diffu- sion model,.

Cocktail-Party Audio-Visual Speech Recognition Seeing through the conversation: Audio-visual speech separation based on diffu- sion model,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:54.504909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:50.095686Z digest=sha256:287e8ad79cff9b5bbebd3d08d51dd7c5553e6c34af39e13a3435627a091aece9

Observation 16f1a9e5-2c18-4ed6-9d0c-d6b801f490dc · outbound

This paper cites Lip read- ing sentences in the wild,.

Cocktail-Party Audio-Visual Speech Recognition Lip read- ing sentences in the wild,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:50.237457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:50.237457Z digest=sha256:44695475464a469f1baa2a0b09427d6f1e5138ad1ca2453e38ab2a5ead458173

Observation b7ad5a0f-a6ca-467e-b941-c5f4ef9d3d18 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Cocktail-Party Audio-Visual Speech Recognition LRS3-TED: a large-scale dataset for visual speech recognition

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:50.389831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:50.389831Z digest=sha256:30e30838f49fc13797d0e6fb2d93f79885cb0bd53bf101a84817f8c16515168d

Observation 754f3861-d6af-4fd6-a7d5-268aafc9688d · outbound

This paper cites V oxceleb2: Deep speaker recognition,.

Cocktail-Party Audio-Visual Speech Recognition V oxceleb2: Deep speaker recognition,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:50.516594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:50.516594Z digest=sha256:a87f2118fafba3bbdcbce324412cb63875d1bd2bd4ee201f5c91d88376b73b06

Observation d149db78-4009-4645-a845-b1eaa41a963e · outbound

This paper cites Unified cross-modal at- tention: Robust audio-visual speech recognition and beyond,.

Cocktail-Party Audio-Visual Speech Recognition Unified cross-modal at- tention: Robust audio-visual speech recognition and beyond,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:54.253078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:50.613225Z digest=sha256:a89b577b80a0bf4d43d687693161d5a65fc5ff6c8d3e886f4871a6dbf147e0d2

Observation 5671cdd1-37fa-430c-bbd5-13b737cd4ac9 · outbound

This paper cites The first multimodal information based speech processing (misp) challenge: Data, tasks, baselines and results,.

Cocktail-Party Audio-Visual Speech Recognition The first multimodal information based speech processing (misp) challenge: Data, tasks, baselines and results,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:53.977114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:50.729030Z digest=sha256:b7396d943b9fb37341ae8453df68a239eade60bc7ed02ac197ae8705756fbf3e

Observation 85d7a2be-fdbc-4a69-964e-a9c1d392050d · outbound

This paper cites Summary on the multimodal information based speech processing (misp) 2022 challenge,.

Cocktail-Party Audio-Visual Speech Recognition Summary on the multimodal information based speech processing (misp) 2022 challenge,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:53.723238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:50.893618Z digest=sha256:12fdd40a6e602ee73162d6dcf60b7bad5858795f3494d8a50ba1074608f5e2f7

Observation 3b46bbb2-dd24-4a08-b344-7fac75f5ebd8 · outbound

This paper cites Summary on the multimodal information- based speech processing (misp) 2023 challenge,.

Cocktail-Party Audio-Visual Speech Recognition Summary on the multimodal information- based speech processing (misp) 2023 challenge,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:53.446222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:51.016640Z digest=sha256:56ae61af58a2a6822b1751deb3f866f02c960fab987a7cb3980b38c50c5dca31

Observation a5122e6c-0b43-417c-9412-5550698e7071 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Cocktail-Party Audio-Visual Speech Recognition Robust Speech Recognition via Large-Scale Weak Supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:51.141038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:51.141038Z digest=sha256:184ecfd7bfdc24df27043b03932a1e46125b2261881cd39efdada1afb470aedf

Observation 044d076e-366b-4787-98de-26b2f36dca3a · outbound

This paper cites Hy- brid ctc/attention architecture for end-to-end speech recognition,.

Cocktail-Party Audio-Visual Speech Recognition Hy- brid ctc/attention architecture for end-to-end speech recognition,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:51.258214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:51.258214Z digest=sha256:c2387b34cafef42d6b7a868ba5543531aa47a71bde9936d3cca3c6f1d94bea6d

Observation 4f5a526e-c436-4455-a64f-17f06af5a54a · outbound

This paper cites End-to-end audio-visual speech recognition with conformers,.

Cocktail-Party Audio-Visual Speech Recognition End-to-end audio-visual speech recognition with conformers,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:53.196111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:51.412311Z digest=sha256:17af976d7705adbd5cec9dc284f066c9e5d95928ba0eabeefac6e06d5ebd5fd2

Observation 2e3ed08e-49da-42ae-bd49-95707080c23d · outbound

This paper cites From text segmentation to smart chaptering: A novel benchmark for structuring video transcriptions,.

Cocktail-Party Audio-Visual Speech Recognition From text segmentation to smart chaptering: A novel benchmark for structuring video transcriptions,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:52.906716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:51.563545Z digest=sha256:b7a117eea1f44b40de822e50a329090d637207f7710c0f92018d71f9c1f4cac2

Observation b61790ee-525c-413c-be7c-5c2b623b9c80 · outbound

This paper cites A light weight model for active speaker detection,.

Cocktail-Party Audio-Visual Speech Recognition A light weight model for active speaker detection,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:51.689641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:51.689641Z digest=sha256:be20a35f179296615c956f4e933b24f2f565bccf36697a64d0ad048bc202edd6

Observation e55c6a2c-03d5-4e7e-86af-cc3423175638 · outbound

This paper cites Out of time: Automated lip sync in the wild,.

Cocktail-Party Audio-Visual Speech Recognition Out of time: Automated lip sync in the wild,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:51.824718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:51.824718Z digest=sha256:81d8f298e53eeb4468b853d7449bc3fef02b413c43b7bfc760f7b6ed61933a16

Observation ad33d507-50b3-411d-8a83-3b4fc66134b2 · outbound

This paper cites Synthetic conversations improve multi-talker asr,.

Cocktail-Party Audio-Visual Speech Recognition Synthetic conversations improve multi-talker asr,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:51.949505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:51.949505Z digest=sha256:1fec2fd792874e6757dcdd9cd8801e5c9731a5daa3c5dca8dfa38f77ebb3040e

Observation 07a0f896-1316-491a-a9f5-6c8445a1a8ee · outbound

This paper cites Msa-asr: Efficient multilingual speaker attribution with frozen asr models,.

Cocktail-Party Audio-Visual Speech Recognition Msa-asr: Efficient multilingual speaker attribution with frozen asr models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:52.651124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T11:33:52.053922Z digest=sha256:12391f06cc97bdefab0a5d5edb98740c1b0a15e2c08bc777b7d693baedb797aa

Pith citing papers

No inbound Pith citation observations are available.