Pith. sign in

Paper Citation Record · LEDGER

Cocktail-Party Audio-Visual Speech Recognition

As of 8 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2506.02178.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02178 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:33:52.053922Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact1
  • verified fuzzy25
  • unresolved15
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69629e47-2e64-441f-9e27-0d3ef9f2be0e · outbound

This paper cites an unresolved cited work.

Cocktail-Party Audio-Visual Speech Recognition Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:34:00.070508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:46.680231Z digest=sha256:8be3f8c1537377bc424b21242cadd4b2406afc2659da7eb1d5613507ca30f895

Observation 09abe8b3-a228-4a21-9d8a-51e66e20fe00 · outbound

This paper cites Task definition Given an input sequence of audio A = {a1, a2,.

Cocktail-Party Audio-Visual Speech Recognition Task definition Given an input sequence of audio A = {a1, a2,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:59.796475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:46.797891Z digest=sha256:5060c81dbb1c0f7d49c8f57e95b85bdbfef3b45f52247218e9f78e99f0d67ca8

Observation 89a15ed6-0bee-4b18-b1f9-1b79cf888d84 · outbound

This paper cites For training, we use LRS2 (train and pretrain sets), V ox2 (train set), and A VYT.

Cocktail-Party Audio-Visual Speech Recognition For training, we use LRS2 (train and pretrain sets), V ox2 (train set), and A VYT

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:33:59.322712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:47.064621Z digest=sha256:d127f614980c0f7759e7c3062d0a0e3dae1fe509b3f6241fbcd65bcdbb57d666

Observation d388c060-6fb0-47e8-a097-b47628032925 · outbound

This paper cites The A V-HuBERT CTC/Attention (A V1) model uses the A V-HuBERT large [12] as the encoder, which has 24 transformer blocks, each with 16 attention heads.

Cocktail-Party Audio-Visual Speech Recognition The A V-HuBERT CTC/Attention (A V1) model uses the A V-HuBERT large [12] as the encoder, which has 24 transformer blocks, each with 16 attention heads

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:59.076103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:47.194615Z digest=sha256:7f7b941104b509dc725132b24da485f7748ffaabc9002fd22788ead30e0fc3ba

Observation 544ba7dc-756e-42e9-b7c2-76383fb4e87e · outbound

This paper cites The WERs for models evaluated on the original LRS2 test set are shown in the column where SNR = ∞.

Cocktail-Party Audio-Visual Speech Recognition The WERs for models evaluated on the original LRS2 test set are shown in the column where SNR = ∞

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:33:58.810431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:47.330476Z digest=sha256:437a5846f024287583de0fc97671e98a369242efc513e23770cc95304f723aed

Observation 0b47fdb0-1712-40ff-9aec-a3e62004d8d3 · outbound

This paper cites We highlighted the gap between conventional datasets and real-world cocktail- party scenarios, where target speakers are not always ac- tive.

Cocktail-Party Audio-Visual Speech Recognition We highlighted the gap between conventional datasets and real-world cocktail- party scenarios, where target speakers are not always ac- tive

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:58.566966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:47.459176Z digest=sha256:8cb6f74b77946fdc18a4ebbe003d18466354c6d813691c9891edcdc5ee44c6cf

Observation 8aa6569e-f973-45b0-9770-ef4ae4cfc25b · outbound

This paper cites How is AI Changing Science? Research in the Era of Learning Algorithms.

Cocktail-Party Audio-Visual Speech Recognition How is AI Changing Science? Research in the Era of Learning Algorithms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:58.280233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:47.577534Z digest=sha256:1cb24c1b710b91eaadef9097445c75781bb3f4930170f3519f213cf4836b8f5f

Observation 000f8e70-32b7-4932-bca3-bdb3b8fcf14e · outbound

This paper cites Knowing who to listen to in speech recognition: Visually guided beamforming,.

Cocktail-Party Audio-Visual Speech Recognition Knowing who to listen to in speech recognition: Visually guided beamforming,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:57.228249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:48.648069Z digest=sha256:da288d3832164740bc0d49bb96bc5a972ff5a23b302abb2b56000010e652f81d

Observation 49f73fe2-8684-49d0-a265-347a8aa056a7 · outbound

This paper cites Hearing lips and seeing voices,.

Cocktail-Party Audio-Visual Speech Recognition Hearing lips and seeing voices,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:47.738240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:47.738240Z digest=sha256:c645841cfdf9dd9baf99884b69e1aeaca1ce81fabe8838a3ef538f5767eff182

Observation 0e0c50f4-0308-40d5-8c9d-1ad72f60d392 · outbound

This paper cites Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation,.

Cocktail-Party Audio-Visual Speech Recognition Muavic: A multilingual audio-visual corpus for robust speech recognition and robust speech-to-text translation,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:47.925263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:47.925263Z digest=sha256:e1de58539f440e2403cdef0ab4ad5989977b5593487c2edc32031b3c8d248c06

Observation 571c6281-057d-4e5e-a4bb-7ac627c59fdc · outbound

This paper cites Super-Human Performance in Online Low-latency Recognition of Conversational Speech.

Cocktail-Party Audio-Visual Speech Recognition Super-Human Performance in Online Low-latency Recognition of Conversational Speech

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:33:52.430647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:48.048385Z digest=sha256:bb698de43dd3cc293342462a47ceba8c4ec692dc665bd66016ad4cfdb11db009

Observation 66de8a82-0444-497e-80e7-a82766e5d749 · outbound

This paper cites an unresolved cited work.

Cocktail-Party Audio-Visual Speech Recognition Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:33:59.539666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:46.922739Z digest=sha256:b17e14e51667af00adb37cff97ad9166cae0e46ae2e9e020c4a42d6306b67159

Observation e4af78f2-617d-404a-ab3a-0487f7473dba · outbound

This paper cites Recognition of conversational telephone speech using the janus speech engine,.

Cocktail-Party Audio-Visual Speech Recognition Recognition of conversational telephone speech using the janus speech engine,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:58.018024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:48.166647Z digest=sha256:3ab2c59408f2014934c02c08d9ff1a8a36713491798dc7dc009f105360dbf11a

Observation d0750410-d6e9-4e47-84c5-4340803780ca · outbound

This paper cites See me, hear me: inte- grating automatic speech recognition and lip-reading,.

Cocktail-Party Audio-Visual Speech Recognition See me, hear me: inte- grating automatic speech recognition and lip-reading,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:48.303473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:48.303473Z digest=sha256:c16a9743b8a596b73f3934ceeeeee661006b705f4159547cd6a717f146d30fb2

Observation 4d39485d-cc8e-4944-80d8-c594e9dc7e1c · outbound

This paper cites Multimodal interfaces,.

Cocktail-Party Audio-Visual Speech Recognition Multimodal interfaces,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:57.723392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:48.418256Z digest=sha256:5e6558f7abee472a8856dbfada7504d1c4a96bd2e9b5a1a04faab0ded2111b58

Observation a64442cf-b5ec-473d-8a67-87295b2d83c8 · outbound

This paper cites Modeling focus of at- tention for meeting indexing,.

Cocktail-Party Audio-Visual Speech Recognition Modeling focus of at- tention for meeting indexing,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:57.500393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:48.541362Z digest=sha256:563a54915f87a5b79aae116e4b5f4e129b7e263070a6edb50190c35b3ab1fe9a

Observation 895b5d0d-c927-4c2e-9035-fa9af05703ee · outbound

This paper cites Visual track- ing for multimodal human computer interaction,.

Cocktail-Party Audio-Visual Speech Recognition Visual track- ing for multimodal human computer interaction,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:56.972059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:48.752559Z digest=sha256:f16938e1a670c40cf68969cbba81665d2a935b14700e340c1d6f4e898ed199eb

Observation 2fb81885-5478-49af-a611-a35f13ec12ad · outbound

This paper cites Estimating focus of attention based on gaze and sound,.

Cocktail-Party Audio-Visual Speech Recognition Estimating focus of attention based on gaze and sound,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:56.728640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:48.882045Z digest=sha256:57b6fa8cecd1dedbbaa8a2281309d6bd62167e18b0a270bee88141949c6a9223

Observation 449b42b3-e185-4d2f-8aa5-487dbb195c09 · outbound

This paper cites Chil: Computers in the human interaction loop,.

Cocktail-Party Audio-Visual Speech Recognition Chil: Computers in the human interaction loop,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:56.459235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:49.010398Z digest=sha256:a7f5cd3507dbc4e20d6596374ace53774dd1051129eb33ae63723f103b2dda2c

Observation 2d352c00-65e7-45db-9210-f041cb23e2c1 · outbound

This paper cites Robust self-supervised audio-visual speech recognition,.

Cocktail-Party Audio-Visual Speech Recognition Robust self-supervised audio-visual speech recognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:56.121810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:49.140475Z digest=sha256:9b3060fa161ddf35554d0f75b33efe0fc115a021df6e6fc54e406ac631d908d5

Observation 417a4001-a7f1-4c34-bba2-e274b8007a83 · outbound

This paper cites Auto-avsr: Audio-visual speech recognition with automatic labels,.

Cocktail-Party Audio-Visual Speech Recognition Auto-avsr: Audio-visual speech recognition with automatic labels,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:55.873050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:49.233772Z digest=sha256:59b11ffa29843d44c9b85edc743e2395b6fea77b22ed6f601ebc1426ed561082

Observation c68e33f3-e420-4d17-8582-18862c7170c1 · outbound

This paper cites Whisper-flamingo: Integrating visual fea- tures into whisper for audio-visual speech recognition and trans- lation,.

Cocktail-Party Audio-Visual Speech Recognition Whisper-flamingo: Integrating visual fea- tures into whisper for audio-visual speech recognition and trans- lation,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:49.361747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:49.361747Z digest=sha256:2ea3219684659ff770cdb2e6441bbae5363cf84b9b0238c89419313fd6b4768d

Observation ce23936d-fff7-4fcc-8e82-78f2229ee0c5 · outbound

This paper cites Speaker-targeted audio-visual models for speech recognition in cocktail-party environments,.

Cocktail-Party Audio-Visual Speech Recognition Speaker-targeted audio-visual models for speech recognition in cocktail-party environments,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:55.584216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:49.466179Z digest=sha256:8e18d13c060d25bbff03255da5a3dd91474eddae1284975d2c7d2a19f5a62263

Observation 616c2b1a-4dc2-409b-b6f2-cbab93fdb9a8 · outbound

This paper cites Audio-visual multi-talker speech recognition in a cocktail party,.

Cocktail-Party Audio-Visual Speech Recognition Audio-visual multi-talker speech recognition in a cocktail party,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:55.337694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:49.588319Z digest=sha256:d4867eefea137637512c135dcf6db6e534c8de15b018e7b931742590b2559a5f

Observation 5412a3ce-84d3-4df4-9229-aed1624a57c4 · outbound

This paper cites Robust audio-visual asr with unified cross-modal attention,.

Cocktail-Party Audio-Visual Speech Recognition Robust audio-visual asr with unified cross-modal attention,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:55.055313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:49.736454Z digest=sha256:2ab7a36614758d99052046f332254321fc1fb6acd3d12d90f1f01d66bf4132ff

Observation a4257471-b8c4-48c8-98e7-4c55494e0247 · outbound

This paper cites Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation,.

Cocktail-Party Audio-Visual Speech Recognition Looking to listen at the cocktail party: a speaker-independent audio-visual model for speech separation,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:49.876226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:49.876226Z digest=sha256:14311241428a8484f934fa16517f282edf3b5ae9804941ee6def629e8b29fdcf

Observation 86c0b254-0e57-4c40-aea5-d2534ecd2d31 · outbound

This paper cites Visualvoice: Audio-visual speech sep- aration with cross-modal consistency,.

Cocktail-Party Audio-Visual Speech Recognition Visualvoice: Audio-visual speech sep- aration with cross-modal consistency,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:54.796426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:50.019620Z digest=sha256:d2c77ab2fe8dca745872b16043fc33d08375cea76d61b2eff7de0a2dff03bc44

Observation a96f7f83-0fed-4d3a-b55b-31f2ebeec776 · outbound

This paper cites Seeing through the conversation: Audio-visual speech separation based on diffu- sion model,.

Cocktail-Party Audio-Visual Speech Recognition Seeing through the conversation: Audio-visual speech separation based on diffu- sion model,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:54.504909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:50.095686Z digest=sha256:5911469462757279ef5296b0fe20186d75f57cd6aceb55b218ef96adf6562e93

Observation 16f1a9e5-2c18-4ed6-9d0c-d6b801f490dc · outbound

This paper cites Lip read- ing sentences in the wild,.

Cocktail-Party Audio-Visual Speech Recognition Lip read- ing sentences in the wild,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:50.237457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:50.237457Z digest=sha256:44695475464a469f1baa2a0b09427d6f1e5138ad1ca2453e38ab2a5ead458173

Observation b7ad5a0f-a6ca-467e-b941-c5f4ef9d3d18 · outbound

This paper cites LRS3-TED: a large-scale dataset for visual speech recognition.

Cocktail-Party Audio-Visual Speech Recognition LRS3-TED: a large-scale dataset for visual speech recognition

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:50.389831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:50.389831Z digest=sha256:30e30838f49fc13797d0e6fb2d93f79885cb0bd53bf101a84817f8c16515168d

Observation 754f3861-d6af-4fd6-a7d5-268aafc9688d · outbound

This paper cites V oxceleb2: Deep speaker recognition,.

Cocktail-Party Audio-Visual Speech Recognition V oxceleb2: Deep speaker recognition,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:50.516594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:50.516594Z digest=sha256:a87f2118fafba3bbdcbce324412cb63875d1bd2bd4ee201f5c91d88376b73b06

Observation d149db78-4009-4645-a845-b1eaa41a963e · outbound

This paper cites Unified cross-modal at- tention: Robust audio-visual speech recognition and beyond,.

Cocktail-Party Audio-Visual Speech Recognition Unified cross-modal at- tention: Robust audio-visual speech recognition and beyond,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:54.253078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:50.613225Z digest=sha256:ff18916eb210c15847cd5a053a9c4c68eff16bd70c1179ba5b62af6678f25bed

Observation 5671cdd1-37fa-430c-bbd5-13b737cd4ac9 · outbound

This paper cites The first multimodal information based speech processing (misp) challenge: Data, tasks, baselines and results,.

Cocktail-Party Audio-Visual Speech Recognition The first multimodal information based speech processing (misp) challenge: Data, tasks, baselines and results,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:53.977114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:50.729030Z digest=sha256:7fd642d74543b3b392e9c189f46ed76fb8033719fb8d5ea0648c87401b294b6a

Observation 85d7a2be-fdbc-4a69-964e-a9c1d392050d · outbound

This paper cites Summary on the multimodal information based speech processing (misp) 2022 challenge,.

Cocktail-Party Audio-Visual Speech Recognition Summary on the multimodal information based speech processing (misp) 2022 challenge,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:53.723238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:50.893618Z digest=sha256:083878afdae8ad1b9814da5ac223c1dc9ec6ac11f77b9e6993760af550982570

Observation 3b46bbb2-dd24-4a08-b344-7fac75f5ebd8 · outbound

This paper cites Summary on the multimodal information- based speech processing (misp) 2023 challenge,.

Cocktail-Party Audio-Visual Speech Recognition Summary on the multimodal information- based speech processing (misp) 2023 challenge,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:53.446222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:51.016640Z digest=sha256:134592988a9480062b140a76c8c2232197e7698432f7ee17645c089af02b37b2

Observation a5122e6c-0b43-417c-9412-5550698e7071 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Cocktail-Party Audio-Visual Speech Recognition Robust Speech Recognition via Large-Scale Weak Supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:51.141038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:51.141038Z digest=sha256:184ecfd7bfdc24df27043b03932a1e46125b2261881cd39efdada1afb470aedf

Observation 044d076e-366b-4787-98de-26b2f36dca3a · outbound

This paper cites Hy- brid ctc/attention architecture for end-to-end speech recognition,.

Cocktail-Party Audio-Visual Speech Recognition Hy- brid ctc/attention architecture for end-to-end speech recognition,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:51.258214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:51.258214Z digest=sha256:c2387b34cafef42d6b7a868ba5543531aa47a71bde9936d3cca3c6f1d94bea6d

Observation 4f5a526e-c436-4455-a64f-17f06af5a54a · outbound

This paper cites End-to-end audio-visual speech recognition with conformers,.

Cocktail-Party Audio-Visual Speech Recognition End-to-end audio-visual speech recognition with conformers,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:53.196111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:51.412311Z digest=sha256:2579b644a79f32e607103769b67814e4236106bf41af82605992f75407e4dad4

Observation 2e3ed08e-49da-42ae-bd49-95707080c23d · outbound

This paper cites From text segmentation to smart chaptering: A novel benchmark for structuring video transcriptions,.

Cocktail-Party Audio-Visual Speech Recognition From text segmentation to smart chaptering: A novel benchmark for structuring video transcriptions,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:52.906716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:51.563545Z digest=sha256:fae34d5b690b7ee45b470d4dda72a2a696dcb0343cb5010d44f150e3f2c14458

Observation b61790ee-525c-413c-be7c-5c2b623b9c80 · outbound

This paper cites A light weight model for active speaker detection,.

Cocktail-Party Audio-Visual Speech Recognition A light weight model for active speaker detection,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:51.689641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:51.689641Z digest=sha256:be20a35f179296615c956f4e933b24f2f565bccf36697a64d0ad048bc202edd6

Observation e55c6a2c-03d5-4e7e-86af-cc3423175638 · outbound

This paper cites Out of time: Automated lip sync in the wild,.

Cocktail-Party Audio-Visual Speech Recognition Out of time: Automated lip sync in the wild,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:51.824718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:51.824718Z digest=sha256:81d8f298e53eeb4468b853d7449bc3fef02b413c43b7bfc760f7b6ed61933a16

Observation ad33d507-50b3-411d-8a83-3b4fc66134b2 · outbound

This paper cites Synthetic conversations improve multi-talker asr,.

Cocktail-Party Audio-Visual Speech Recognition Synthetic conversations improve multi-talker asr,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:51.949505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:51.949505Z digest=sha256:1fec2fd792874e6757dcdd9cd8801e5c9731a5daa3c5dca8dfa38f77ebb3040e

Observation 07a0f896-1316-491a-a9f5-6c8445a1a8ee · outbound

This paper cites Msa-asr: Efficient multilingual speaker attribution with frozen asr models,.

Cocktail-Party Audio-Visual Speech Recognition Msa-asr: Efficient multilingual speaker attribution with frozen asr models,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:33:52.651124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:33:52.053922Z digest=sha256:c6787f5f7c54a3522c59f5d591c8026f0c547c20fa2af57dd7f428a52c56b1bb

Pith citing papers

No inbound Pith citation observations are available.