Pith. sign in

Paper Citation Record · LEDGER

Robust Speech Recognition via Large-Scale Weak Supervision

As of 10 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 100 inbound Pith citation observations for arXiv:2212.04356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.04356 v1

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-24T10:13:51.001946Z

measured 116 of 116 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 276 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:52:42.964721Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact1
  • verified fuzzy2
  • unresolved11
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

1161
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 61781080-c448-447f-8fc7-026510d051c5 · outbound

This paper cites file number.

Robust Speech Recognition via Large-Scale Weak Supervision file number

Reference 1

Resolution
verified exact
doi, observed 2026-05-24T10:14:18.668420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:e4926fd8b90ad2c4c6e5efcf894ef77fcb2716ea577a1f5b8adb08f95781d694

Observation 5f4d4acb-031d-40db-980f-cd98729ed52f · outbound

This paper cites an unresolved cited work.

Robust Speech Recognition via Large-Scale Weak Supervision Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:14:19.076523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:11602c47e7ba3ae4ed64ec21c49ff29d97c5cf9dc089ba685717e7af756ca348

Observation 757b85d9-7666-4058-8a6a-fd54a6a1eb16 · outbound

This paper cites an unresolved cited work.

Robust Speech Recognition via Large-Scale Weak Supervision Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:14:19.080390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:1d6230d192f21185f814a14673f3041776fd4ac9e1182106c1279351ab622a8a

Observation 86368a5c-e694-4932-9ae7-cc2bc0394a6b · outbound

This paper cites an unresolved cited work.

Robust Speech Recognition via Large-Scale Weak Supervision Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:14:19.065035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:5a2817bd4342f9bc303c2c9b48dba9fad914e413a026c7e8b7b96c1a87132699

Observation 2cfb5b40-3877-4fd3-8de6-53a9468b4172 · outbound

This paper cites an unresolved cited work.

Robust Speech Recognition via Large-Scale Weak Supervision Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:14:19.072429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:9782a1f6a05f5e111feb06cbd3a5c6cf079c8b7f82f35bb0758c42f038f4f6cd

Observation ae76955c-c168-43e0-bedf-48187814673c · outbound

This paper cites an unresolved cited work.

Robust Speech Recognition via Large-Scale Weak Supervision Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:14:19.068617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:275eb5fef224d32d942e192617d87fe6fdbcb68d4208ae75367c09eddef4890e

Observation f4aaf036-1856-4510-9e90-ead9e6f40212 · outbound

This paper cites an unresolved cited work.

Robust Speech Recognition via Large-Scale Weak Supervision Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:14:19.094208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:63d8bb4421422053f402131c7c98694ed8e6243bace19f1fd5b93cc506cff138

Observation 78e3db12-ddb6-4054-8746-9d49369dd073 · outbound

This paper cites Ten thousand dollars.

Robust Speech Recognition via Large-Scale Weak Supervision Ten thousand dollars

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T10:14:19.097962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:cf8e61a97e2de6a803927c0d738fb73e2a9c1803e8c92c06dd0ee0a0ba061c19

Observation f6be1ece-56f6-499a-b792-a5b6824af10f · outbound

This paper cites an unresolved cited work.

Robust Speech Recognition via Large-Scale Weak Supervision Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:14:19.061475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:17ebbe6f942f64d61f0d1eee681fecca71dbdd880c77bee6dc485574f861d57a

Observation fac8fe65-1d88-46cd-9160-810fde870e25 · outbound

This paper cites an unresolved cited work.

Robust Speech Recognition via Large-Scale Weak Supervision Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:14:19.091159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:24a2786898e28d89e1c1af4dcfc58a97ed51d76640e5536479863ad353c0773c

Observation 3478da73-e72c-4689-9d70-d740b11e45e2 · outbound

This paper cites an unresolved cited work.

Robust Speech Recognition via Large-Scale Weak Supervision Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:14:19.084234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:8160f19fc831f63102f94aef20d524c9a5cc209fa8d0761f94d94afa194edcea

Observation 4e2f6698-1e71-4cc2-abc4-ad326a53951d · outbound

This paper cites an unresolved cited work.

Robust Speech Recognition via Large-Scale Weak Supervision Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:14:19.046686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:bc1f310ba719dd71fa8be9cec028267e3753dd124ce3d7ceba2c050eea074058

Observation cece95c8-ffc8-40af-9c5d-9438189560a2 · outbound

This paper cites an unresolved cited work.

Robust Speech Recognition via Large-Scale Weak Supervision Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-05-24T10:14:19.087489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:ae7290d16181311e7fd826e6fa9697333fa34ca95111e79f7c65eb694646db23

Observation f84b9375-4576-450f-b101-fa52eb5668c9 · outbound

This paper cites when the Unicode category of each character in the NFKC-normalized string starts with M, S, or P.

Robust Speech Recognition via Large-Scale Weak Supervision when the Unicode category of each character in the NFKC-normalized string starts with M, S, or P

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-24T10:14:19.056636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:6e16d69cdc0d1d5ac574a17678954d87cdb8a27f210cacb80c8212458564f662

Observation 3f3bbae0-3ce6-4c80-9cd1-98b0d28333b9 · outbound

This paper cites an unresolved cited work.

Robust Speech Recognition via Large-Scale Weak Supervision Unresolved cited work

Reference 17

Resolution
parse uncertain
raw_fallback, observed 2026-05-24T10:14:19.049852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:fd819ad6e8a624e026d855103aa1fcd6479f61f046570fd711316ae9b33470dd

Observation 51a45dff-b5ce-4992-ad0c-3ce67a589859 · outbound

This paper cites an unresolved cited work.

Robust Speech Recognition via Large-Scale Weak Supervision Unresolved cited work

Reference 18

Resolution
malformed identifier
raw_fallback, observed 2026-05-24T10:14:19.053117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T10:13:51.001946Z digest=sha256:e28a7fafb96a03919178ca8bc46225d36f67d5e16f05d5512de6e8c9304fd51c

Pith citing papers

Observation 25aebbe4-9b63-4c96-a5ff-564402859396 · inbound

Basic syntax from speech: Spontaneous concatenation in unsupervised deep neural networks cites this paper.

Basic syntax from speech: Spontaneous concatenation in unsupervised deep neural networks Robust Speech Recognition via Large-Scale Weak Supervision

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:09:15.624681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-24T09:06:55.725579Z digest=sha256:7b67a892156966d77346d1a953dfebcd5627b049493d0d80fe9d39a0c8c926d2

Observation 1f5592a9-5b6c-4a81-bcce-e25bde195f02 · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding Robust Speech Recognition via Large-Scale Weak Supervision

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:30:00.609860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:7208406ac673292c78e2b392746f7fe2e66c793f899c82fe1adf823d9ae4dff3

Observation 9c409c77-d2f6-46a8-bb54-13ea20ac87b2 · inbound

AudioPaLM: A Large Language Model That Can Speak and Listen cites this paper.

AudioPaLM: A Large Language Model That Can Speak and Listen Robust Speech Recognition via Large-Scale Weak Supervision

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:07:57.890817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T07:07:57.800866Z digest=sha256:c4781aa70cf39c3c69bb44f282519e201081b3f7126a430c0a796ab550732de7

Observation 0d8fddea-ab3a-4da0-9d3c-1bb5ecce36fb · inbound

Towards General Text Embeddings with Multi-stage Contrastive Learning cites this paper.

Towards General Text Embeddings with Multi-stage Contrastive Learning Robust Speech Recognition via Large-Scale Weak Supervision

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-12T03:33:45.943920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T03:33:45.855974Z digest=sha256:83fd1a78e3ba19ef01f3d13e2a6a0b4398ad23801ef35f0c08b08da3ee6ac3ab

Observation e00c5b17-6604-43f4-a113-790b109ab2ac · inbound

Community-Informed AI Models for Police Accountability cites this paper.

Community-Informed AI Models for Police Accountability Robust Speech Recognition via Large-Scale Weak Supervision

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-24T05:08:54.647010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T05:08:34.362690Z digest=sha256:bd484b63ff88a079263ae6f26c1b4aea8062116e37cad590dac0f84dfd34f1db

Observation e3f41e3d-44d2-462e-8ef7-e2a26197d103 · inbound

DASB - Discrete Audio and Speech Benchmark cites this paper.

DASB - Discrete Audio and Speech Benchmark Robust Speech Recognition via Large-Scale Weak Supervision

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-24T00:28:39.523430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T00:26:57.419537Z digest=sha256:3894597975d44f10f2949fd11c509390070ff7fd4398c1815bf704f258a24496

Observation 90d6252e-f3ee-4b63-b42e-fcf338e96f23 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output Robust Speech Recognition via Large-Scale Weak Supervision

Reference 124

Resolution
verified exact
local_arxiv, observed 2026-05-17T10:46:28.621293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:e73beaac097ff683d2fd524fe0523cc1aa609c2a9f0e7cdb969f3b67abc80ce5

Observation aed2f2ed-ae40-44fd-8d88-2056d9e2154c · inbound

Multimodal Contextualized Support for Enhancing Video Retrieval System cites this paper.

Multimodal Contextualized Support for Enhancing Video Retrieval System Robust Speech Recognition via Large-Scale Weak Supervision

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:15:28.460774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T07:14:34.843867Z digest=sha256:38e199e8a38bfa7b03b33cd60bb7e490d6675058751b419f69a0f8e2e5ec7919

Observation 6407b19f-7730-4a5f-ba5b-821e9f6ef611 · inbound

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) cites this paper.

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) Robust Speech Recognition via Large-Scale Weak Supervision

Reference 104

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T05:52:37.599659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T05:47:48.488826Z digest=sha256:42305f51b3265a157b4af24d2a93516f622cf2bad39b3c3407dcc4dbadf31cab

Observation add67a61-7082-413d-94da-588cddc480b8 · inbound

Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data cites this paper.

Leveraging Chain of Thought towards Empathetic Spoken Dialogue without Corresponding Question-Answering Data Robust Speech Recognition via Large-Scale Weak Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T18:52:42.964721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:52:42.964721Z digest=sha256:7e5d34d1b7200abe860355aae8952507db9ee77ce99e3a6d09ef444f97ced29e

Observation 9fb5f113-5f3e-4810-a7c2-1fa40851ed81 · inbound

LLM-Assisted Knowledge Graph Completion for Curriculum and Domain Modelling in Personalized Higher Education Recommendations cites this paper.

LLM-Assisted Knowledge Graph Completion for Curriculum and Domain Modelling in Personalized Higher Education Recommendations Robust Speech Recognition via Large-Scale Weak Supervision

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T17:22:05.514989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:22:05.514989Z digest=sha256:ee1252a7c1d3ba2f72467bdf699f93e99ef81c4d8d935fc7f0b9a034893aa165

Observation 2f341de6-15a3-4f6b-9f02-40afc73fd34a · inbound

Toxicity Begets Toxicity: Unraveling Conversational Chains in Political Podcasts cites this paper.

Toxicity Begets Toxicity: Unraveling Conversational Chains in Political Podcasts Robust Speech Recognition via Large-Scale Weak Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T17:02:51.148066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:02:51.148066Z digest=sha256:ad67cd50fdbc44cf03ad2e548cd75ec696afe30a4619cb1571b401e317f53b4b

Observation 04d38a44-17df-41f3-ab92-77c111efba33 · inbound

Collective Memory and Narrative Cohesion: A Computational Study of Palestinian Refugee Oral Histories in Lebanon cites this paper.

Collective Memory and Narrative Cohesion: A Computational Study of Palestinian Refugee Oral Histories in Lebanon Robust Speech Recognition via Large-Scale Weak Supervision

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-10T15:46:12.170322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:46:12.170322Z digest=sha256:41cd827ddbc5062c20ead11d881069674c5948f1344b37ca036b0f3dd5c25376

Observation 07987b6a-79e5-4c70-adfd-fdbf8e344ddf · inbound

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos cites this paper.

Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos Robust Speech Recognition via Large-Scale Weak Supervision

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-14T00:32:41.244010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T00:32:41.059558Z digest=sha256:823496efe08ef57a6e7f2218331fcc4ef204d8c86181632e3a4d301116add4af

Observation 8535ee46-ff4d-438e-8ffd-bf24a9da38b1 · inbound

DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images cites this paper.

DrawEduMath: Evaluating Vision Language Models with Expert-Annotated Students' Hand-Drawn Math Images Robust Speech Recognition via Large-Scale Weak Supervision

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T14:54:33.838786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T14:54:33.838786Z digest=sha256:d6a2e949a2910d682aa901e96d53a719c184e00b4a7cc0f12ab5974799adc033

Observation 32a94e51-607c-42ff-8a59-9bd688dd324c · inbound

Profiling Apple Silicon Performance for ML Training cites this paper.

Profiling Apple Silicon Performance for ML Training Robust Speech Recognition via Large-Scale Weak Supervision

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:49:55.588901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:49:55.588901Z digest=sha256:a23b2130ee0ec0f44363ddc585aded9cde9389fe67f69c8b2e4bcef30cb2f6b4

Observation 0b56fcd9-dbb9-434f-9f7b-2f1ec0e7c116 · inbound

SELMA: A Speech-Enabled Language Model for Virtual Assistant Interactions cites this paper.

SELMA: A Speech-Enabled Language Model for Virtual Assistant Interactions Robust Speech Recognition via Large-Scale Weak Supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T20:20:58.946422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T20:20:58.946422Z digest=sha256:399c06773bb30746a90ec08b2d760b17d06e0afa91e8cce3ec8e80def8519ec6

Observation 2beee3f1-c15f-4a57-ac43-cf17a5a27df7 · inbound

Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language cites this paper.

Sagalee: an Open Source Automatic Speech Recognition Dataset for Oromo Language Robust Speech Recognition via Large-Scale Weak Supervision

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T19:07:33.651511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T19:07:33.651511Z digest=sha256:52a4479a650b6f8137947ddbc20717072e6e6e171f0cc5914000ed5f100ed8ba

Observation 14eb5793-877f-4d05-99f9-9a7cadc400e0 · inbound

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models cites this paper.

"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models Robust Speech Recognition via Large-Scale Weak Supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T18:05:52.341242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:05:52.341242Z digest=sha256:9d6e7743568a2f980cb5f7306f7c002629f2fdd0d8444030db40e68e7dd316a6

Observation e5c149c6-979a-4f34-b0d0-349eee89c42f · inbound

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis cites this paper.

Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis Robust Speech Recognition via Large-Scale Weak Supervision

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T16:43:08.831448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T16:43:08.831448Z digest=sha256:7636fe1bcb8253573d3024662597eae2146ec4c918e1d822ae7ac3dabe442616

Observation d0a90cce-4323-417b-94d6-5ff980e2f47f · inbound

Secure & Personalized Music-to-Video Generation via CHARCHA cites this paper.

Secure & Personalized Music-to-Video Generation via CHARCHA Robust Speech Recognition via Large-Scale Weak Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T17:06:31.329295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T17:06:31.329295Z digest=sha256:126a33c29c94ea21793325c5552bf3e99029c75ea3320eb3ce54d80b7e094b77

Observation 6c11505b-32de-4a0c-995f-7148d96d07bb · inbound

NUTSHELL: A Dataset for Abstract Generation from Scientific Talks cites this paper.

NUTSHELL: A Dataset for Abstract Generation from Scientific Talks Robust Speech Recognition via Large-Scale Weak Supervision

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-23T02:52:26.539614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T02:50:10.307310Z digest=sha256:1c8f5864295ae79aa1fccb256ba0ce1476cc47a21dae8f882438dcfe4f1f3a69

Observation cf7e1600-435f-49ed-9a59-da803e35217e · inbound

ModRWKV: Transformer Multimodality in Linear Time cites this paper.

ModRWKV: Transformer Multimodality in Linear Time Robust Speech Recognition via Large-Scale Weak Supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:35.723258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:35.723258Z digest=sha256:1f21c899374501bc91f0fe126db3270be091ed06f0c6a649c30d7231f0aa077a

Observation 092bb8b5-818c-403a-9762-39b6ad937b9a · inbound

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing cites this paper.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Robust Speech Recognition via Large-Scale Weak Supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.363576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.363576Z digest=sha256:d63e69f0677ffd7215251e7447c04b53495dcdff4fe4aa0a0c66f3bc51642489

Observation 7489012e-c329-4f5a-8be7-8331d9ac307c · inbound

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition cites this paper.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Robust Speech Recognition via Large-Scale Weak Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:56.715907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:56.715907Z digest=sha256:fd06c935ade6d9d1c2e3c854403d9e0ebaa7bd33614e86569752ec02ac8ed0f7

Observation 362e459a-056b-41d1-892a-a9e44a7ccfaa · inbound

VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering cites this paper.

VoxRAG: A Step Toward Transcription-Free RAG Systems in Spoken Question Answering Robust Speech Recognition via Large-Scale Weak Supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:01.681010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:01.681010Z digest=sha256:8ead8407a1c807960e84a2d49e5c4139032bf3f03b9346acabee1bb856b1874c

Observation 669ef1d5-6dae-491c-9720-dd9879aa6fa2 · inbound

Voice of a Continent: Mapping Africa's Speech Technology Frontier cites this paper.

Voice of a Continent: Mapping Africa's Speech Technology Frontier Robust Speech Recognition via Large-Scale Weak Supervision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:34:18.559801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:34:18.559801Z digest=sha256:fcb5462221165566a7c16b72e5d26334b269e5a15d0a47887eb52084548e6c16

Observation 9ab42704-7e69-4852-a0a3-61943c0b3599 · inbound

Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving cites this paper.

Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving Robust Speech Recognition via Large-Scale Weak Supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:33.802965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:33.802965Z digest=sha256:8d0ef3c5ad73cece881d060961a2335952741d5540eb19ae9ee88294e806756d

Observation eac08d46-0bd1-4f26-8de5-760e16456aaf · inbound

WHISTRESS: Enriching Transcriptions with Sentence Stress Detection cites this paper.

WHISTRESS: Enriching Transcriptions with Sentence Stress Detection Robust Speech Recognition via Large-Scale Weak Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:23:38.902233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:23:38.902233Z digest=sha256:2952d3c79f24e72c7e7245ea737357b61f4ddc25a9af6cea6f6e5613ca9a4045

Observation 7ffd8915-540a-457e-ad2a-feebe127b0e5 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:58.084628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:58.084628Z digest=sha256:eba823c336eb84ea16177d00cde8fa8f3d74d7e92ccdb33098a1cca89df4c7ee

Observation 82083e53-5fa3-4c13-a1a8-e6e616108ce4 · inbound

FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching cites this paper.

FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching Robust Speech Recognition via Large-Scale Weak Supervision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:27.221397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:27.221397Z digest=sha256:8d64d3cb142606c60125ca2c31b658c8abd89c1f727434e80c503b2eaadce366

Observation ac4fef8f-9040-40f2-8b58-b96e78ba18bf · inbound

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation cites this paper.

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:47.158251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:47.158251Z digest=sha256:bb103ac0ab4bab8d7ae9647ef120ce6a44b1ba21444075546e897e715219da79

Observation 653a531d-516d-43c4-a81d-4a0d21144b4e · inbound

ALAS: An Automatic Latent Alignment Score for Audio Language Models cites this paper.

ALAS: An Automatic Latent Alignment Score for Audio Language Models Robust Speech Recognition via Large-Scale Weak Supervision

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:07:23.097224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:07:23.097224Z digest=sha256:05bbc36c7ceba7a8a20036addf1fdcc0d532fd6ea696c999c2197ae08454e3de

Observation 5d37df95-1f99-4f47-b8ae-994a436c9dc8 · inbound

Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation cites this paper.

Towards Pretraining Robust ASR Foundation Model with Acoustic-Aware Data Augmentation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:56:55.026428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:56:55.026428Z digest=sha256:9cf4099e84fdc4963dece3170803d73d9b7e3823d71271b14a4643f2fcba06be

Observation 00cfc314-ae5e-4872-803a-f1b002791421 · inbound

Tell me Habibi, is it Real or Fake? cites this paper.

Tell me Habibi, is it Real or Fake? Robust Speech Recognition via Large-Scale Weak Supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:17.805464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:17.805464Z digest=sha256:d7e1c11fe5b7d4de15f7516e2362a0a0f77e32ebf8844b123dded9d015c31784

Observation fd8df946-d27d-4c42-842c-6f7511a71823 · inbound

Evaluating AI capabilities in detecting conspiracy theories on YouTube cites this paper.

Evaluating AI capabilities in detecting conspiracy theories on YouTube Robust Speech Recognition via Large-Scale Weak Supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:44.309947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:44.309947Z digest=sha256:357c36267569a301f6d1db82f2299b37bcf9fb1b24529f7aefa34cfc606c8568

Observation 510a616b-305f-42a4-9ceb-3b4a2bf17860 · inbound

Automatic classification of stop realisation with wav2vec2.0 cites this paper.

Automatic classification of stop realisation with wav2vec2.0 Robust Speech Recognition via Large-Scale Weak Supervision

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:41.505406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:41.505406Z digest=sha256:cd10b82d9740d3b53e66586f7d32b7700f69489da033b38e09113b821a4aa49f

Observation 67a4ed0c-25ad-4b1f-8b11-03838c63e7c6 · inbound

Speech as a Multimodal Digital Phenotype for Multi-Task LLM-based Mental Health Prediction cites this paper.

Speech as a Multimodal Digital Phenotype for Multi-Task LLM-based Mental Health Prediction Robust Speech Recognition via Large-Scale Weak Supervision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:23:36.500911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:23:36.500911Z digest=sha256:764c72b64f837646eaf2f45a0b092f5e776c1c30d8f6ddb913466993e9eb680a

Observation 927629a9-723c-4a3a-b10d-8afa9b6fcabb · inbound

BeaverTalk: Oregon State University's IWSLT 2025 Simultaneous Speech Translation System cites this paper.

BeaverTalk: Oregon State University's IWSLT 2025 Simultaneous Speech Translation System Robust Speech Recognition via Large-Scale Weak Supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:01.114392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:43:01.114392Z digest=sha256:561850f3171e247708e92f2f452ea13b07eb3b5120b8fd68302802e842ed9b6e

Observation 7ccbdb48-8c5d-4768-a10f-c3d2414dfe97 · inbound

Improving Language and Modality Transfer in Translation by Character-level Modeling cites this paper.

Improving Language and Modality Transfer in Translation by Character-level Modeling Robust Speech Recognition via Large-Scale Weak Supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.601428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.601428Z digest=sha256:b3807465447cac80363bc311cffff6d08995e959e940f88e04e8a63f40c1b19c

Observation fe169053-4210-4de0-8198-e5079bc077c8 · inbound

CASPER: A Large Scale Spontaneous Speech Dataset cites this paper.

CASPER: A Large Scale Spontaneous Speech Dataset Robust Speech Recognition via Large-Scale Weak Supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:11:47.662234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:11:47.662234Z digest=sha256:74bf7014453511a08e2d2ae5a33f63d177b742c6ba78cd0f1d3c1e8499f9b9aa

Observation 0e18c3fc-2abf-449a-b237-3e7684be45e7 · inbound

MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation cites this paper.

MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:29.194153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:29.194153Z digest=sha256:582251de80e8bd6806e170b566c9504b96bda0249b53f4102b0e1cdac058bf2e

Observation 8f669cfc-e9bd-4201-bfe5-7cf4c0b33446 · inbound

Towards Temporally Explainable Dysarthric Speech Clarity Assessment cites this paper.

Towards Temporally Explainable Dysarthric Speech Clarity Assessment Robust Speech Recognition via Large-Scale Weak Supervision

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:16.251129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:16.251129Z digest=sha256:95a6343c99fd070b2bdcaefe14bc0fa039f1d9d5b839722a2fff5d074042e153

Observation b3c97642-4e2f-49ba-b054-61c624db5296 · inbound

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction cites this paper.

CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction Robust Speech Recognition via Large-Scale Weak Supervision

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T11:51:05.676589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:51:05.676589Z digest=sha256:3caf99c3ee2005433438ff130a745771ca006d39fd65140f7d244b740a3594a4

Observation 5b7dbf67-aec0-49ee-a7c1-e096206ce4e3 · inbound

Speech-to-Speech Translation Pipelines for Conversations in Low-Resource Languages cites this paper.

Speech-to-Speech Translation Pipelines for Conversations in Low-Resource Languages Robust Speech Recognition via Large-Scale Weak Supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:50:22.032115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:50:22.032115Z digest=sha256:4b39e27b7af233fdcf4ef7d32ddc95389da118e10999346b32bc97a7ca42e3b9

Observation a5122e6c-0b43-417c-9412-5550698e7071 · inbound

Cocktail-Party Audio-Visual Speech Recognition cites this paper.

Cocktail-Party Audio-Visual Speech Recognition Robust Speech Recognition via Large-Scale Weak Supervision

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:33:51.141038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:33:51.141038Z digest=sha256:948b0513395621e805e14150034c30300e164b90c4e21bcba26a675a9977ab2b

Observation e411ff1b-a951-471b-ac16-e942d99d6bfc · inbound

SingaKids: A Multilingual Multimodal Dialogic Tutor for Language Learning cites this paper.

SingaKids: A Multilingual Multimodal Dialogic Tutor for Language Learning Robust Speech Recognition via Large-Scale Weak Supervision

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:27:12.982988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:27:12.982988Z digest=sha256:2b2245bbc4bf34a84438f7ccaea86bff4586ebb51f24cc4ae7fb386f7b8e618d

Observation 000021e7-1a98-488e-a767-fbe79ea35c0a · inbound

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions cites this paper.

Prompt-Unseen-Emotion: Zero-shot Expressive Speech Synthesis with Prompt-LLM Contextual Knowledge for Mixed Emotions Robust Speech Recognition via Large-Scale Weak Supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:21:11.541084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:21:11.541084Z digest=sha256:a976781b7d8ece21f8f9789953eaae907e6431a15096ac846ea06dac597a1c0b

Observation 9bab065b-80a5-40ab-8ae1-e532b1240ee6 · inbound

It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text Systems cites this paper.

It's Not a Walk in the Park! Challenges of Idiom Translation in Speech-to-text Systems Robust Speech Recognition via Large-Scale Weak Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:00.179089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:16:00.179089Z digest=sha256:f78458dfa789cf73f96458aa4971ae64032420a795957aa43fd42e4f1facd7f4

Observation c89fd632-0615-417d-824b-23b60a90d21b · inbound

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation cites this paper.

Controllable Text-to-Speech Synthesis with Masked-Autoencoded Style-Rich Representation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:16:32.166827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:16:32.166827Z digest=sha256:dca65be6a4bd63f6118e5f1eb88d4106ed8f8fe0242eafdbb79918814701d814

Observation 7561d1b5-c72f-4b58-9ef6-34f8a48cddea · inbound

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text cites this paper.

The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text Robust Speech Recognition via Large-Scale Weak Supervision

Reference 138

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:44.929007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:44.929007Z digest=sha256:33dadfa1559c55cd7c57eb4ba46e5dbb28c77ee4781e029994e585bc6cc18071

Observation 053d37e6-e36e-4348-a95b-a9e45a799d34 · inbound

SocialDF: Benchmark Dataset and Detection Model for Mitigating Harmful Deepfake Content on Social Media Platforms cites this paper.

SocialDF: Benchmark Dataset and Detection Model for Mitigating Harmful Deepfake Content on Social Media Platforms Robust Speech Recognition via Large-Scale Weak Supervision

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.738562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.738562Z digest=sha256:46ce37c06e0985306b055c5c200819dbc0227fbe5d40790f0d3dd9756cb532b5

Observation 9b849b07-d899-4c0d-aed9-622c927b7025 · inbound

SocialDF: Benchmark Dataset and Detection Model for Mitigating Harmful Deepfake Content on Social Media Platforms cites this paper.

SocialDF: Benchmark Dataset and Detection Model for Mitigating Harmful Deepfake Content on Social Media Platforms Robust Speech Recognition via Large-Scale Weak Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:24:24.751517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:24:24.751517Z digest=sha256:b0e28b11673c9e7bcebe9b53e937080ab6cabcc2317ca786a01c81de7de24ed7

Observation 63cd450c-52ea-467c-b862-0b91c26bfd4f · inbound

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval cites this paper.

CLaMR: Contextualized Late-Interaction for Multimodal Content Retrieval Robust Speech Recognition via Large-Scale Weak Supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T06:05:31.772749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:05:31.772749Z digest=sha256:3e689c5733d933e15b85dc8d9ce589afbb10825f6f31ffcfea1be65de2a305f7

Observation 95e1751b-1236-4cd7-9a5b-858ef3a85bc0 · inbound

Rhythm Features for Speaker Identification cites this paper.

Rhythm Features for Speaker Identification Robust Speech Recognition via Large-Scale Weak Supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:06.944151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:54:06.944151Z digest=sha256:b2acb3d7acf4b24406bd1f837f02470769204f1b6a7873c487311981a1c182e3

Observation 780ec326-efaf-43c1-b0ed-fc11a23069e1 · inbound

DeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error Correction cites this paper.

DeRAGEC: Denoising Named Entity Candidates with Synthetic Rationale for ASR Error Correction Robust Speech Recognition via Large-Scale Weak Supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:37:55.521083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:37:55.521083Z digest=sha256:46972676235e36690e7ca0c92bef65bdc27e231e7bef017c61bdd3418aa48dd4

Observation 8de39f9e-f8cc-4735-a072-778fc06aeeb6 · inbound

"I Said Things I Needed to Hear Myself": Peer Support as an Emotional, Organisational, and Sociotechnical Practice in Singapore cites this paper.

"I Said Things I Needed to Hear Myself": Peer Support as an Emotional, Organisational, and Sociotechnical Practice in Singapore Robust Speech Recognition via Large-Scale Weak Supervision

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-19T10:32:14.540239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T10:28:12.506891Z digest=sha256:96dfeeef5456819706414e0634c915a39bbcf037dba4c09a068d93b9492e6690

Observation f5fc0067-badf-4842-bce7-67fd00d00cac · inbound

"I Said Things I Needed to Hear Myself": Peer Support as an Emotional, Organisational, and Sociotechnical Practice in Singapore cites this paper.

"I Said Things I Needed to Hear Myself": Peer Support as an Emotional, Organisational, and Sociotechnical Practice in Singapore Robust Speech Recognition via Large-Scale Weak Supervision

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:52.220765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:52.220765Z digest=sha256:4a34c6d4adda8fd512f91ec3367460c03dc69da1b1f83ef411fed5c3da44b69c

Observation 8e062e4f-a776-4688-8bab-4b1ca7efad0f · inbound

MAGMaR Shared Task System Description: Video Retrieval with OmniEmbed cites this paper.

MAGMaR Shared Task System Description: Video Retrieval with OmniEmbed Robust Speech Recognition via Large-Scale Weak Supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:44.054699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:53:44.054699Z digest=sha256:9bf7a1c0b128d032bc61e42437a8d7f8581f7115fd6b18f149985fa08ba4ffe9

Observation 3e43e5e7-da8b-4a6a-ac04-6751ecdffe79 · inbound

Hidden in Plain Sight: Evaluation of the Deception Detection Capabilities of LLMs in Multimodal Settings cites this paper.

Hidden in Plain Sight: Evaluation of the Deception Detection Capabilities of LLMs in Multimodal Settings Robust Speech Recognition via Large-Scale Weak Supervision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:39.096796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:53:39.096796Z digest=sha256:120f90e4eec7ef4165b8843609d04e19cd93a84bb159f691f113c0f296cd9daf

Observation 5092846e-ce14-438f-b405-42cea1176d02 · inbound

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models cites this paper.

Do We Still Need Audio? Rethinking Speaker Diarization with a Text-Based Approach Using Multiple Prediction Models Robust Speech Recognition via Large-Scale Weak Supervision

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:15:56.689663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:15:56.689663Z digest=sha256:701b0f394d2b98ef4abe6abf423a4fd8406d45c0d7f55bb7b7756f627b63d055

Observation 1cdb8260-90bf-493d-bc95-08b50280516b · inbound

CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following cites this paper.

CMI-Bench: A Comprehensive Benchmark for Evaluating Music Instruction Following Robust Speech Recognition via Large-Scale Weak Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:59:49.340105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:59:49.340105Z digest=sha256:6840aa2a50cc1b7b04b6719bb8910a2d88405b7e3a493e3de67fa6584757b796

Observation ef7dbf33-7411-4e07-87b9-bc4450a22762 · inbound

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling cites this paper.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Robust Speech Recognition via Large-Scale Weak Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:03.933734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:03.933734Z digest=sha256:61eb8a3af3d63b19bd11595db3d056626e6f8eb23ed98da820e7c05718c4993c

Observation 561a7fc8-18a6-4e7b-a1ae-42a92ef577a4 · inbound

AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR cites this paper.

AsyncSwitch: Asynchronous Text-Speech Adaptation for Code-Switched ASR Robust Speech Recognition via Large-Scale Weak Supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:13.221764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:23:13.221764Z digest=sha256:11336860d177e4ec3ca58d45bd1b937ae8b168dfd4c6991d697b94b289d83863

Observation ec36eb69-f540-4f2e-9e9d-2ae8e6b8fabd · inbound

Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion cites this paper.

Double Entendre: Robust Audio-Based AI-Generated Lyrics Detection via Multi-View Fusion Robust Speech Recognition via Large-Scale Weak Supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T23:48:39.299238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:48:39.299238Z digest=sha256:2ba40fd6f1937e8c76f584efcc7bc5531ea85bcad5503540a7b34f6d06668894

Observation 9b8ec22b-69c3-4c6b-9053-dbe902e1692e · inbound

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM cites this paper.

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM Robust Speech Recognition via Large-Scale Weak Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:49.574295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:49.574295Z digest=sha256:136a0849844dda6fbe2aeeed5c53cfc59be55c358ff3c42874845df56b603484

Observation 547ac7e2-02e2-4d3e-a840-c1a1a4ba4b50 · inbound

AI-Generated Song Detection via Lyrics Transcripts cites this paper.

AI-Generated Song Detection via Lyrics Transcripts Robust Speech Recognition via Large-Scale Weak Supervision

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:34.685007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:21:34.685007Z digest=sha256:30c78db4752ee226bff3f5bf254b3bb2598b86e79c5ea8485ade466b00ea89d0

Observation fb4c00f9-49d5-4e61-b4cb-9c792d554f9f · inbound

The Anatomy of Speech Persuasion: Linguistic Shifts in LLM-Modified Speeches cites this paper.

The Anatomy of Speech Persuasion: Linguistic Shifts in LLM-Modified Speeches Robust Speech Recognition via Large-Scale Weak Supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:14.347829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:20:14.347829Z digest=sha256:7a8ae553c3afaa6c87a7e896a79eec4fe067fdc52e3ef881c079b333266f1456

Observation 30b1a339-1ffb-4fff-b9a5-23f5fe2b993d · inbound

FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Models cites this paper.

FormosanBench: Benchmarking Low-Resource Austronesian Languages in the Era of Large Language Models Robust Speech Recognition via Large-Scale Weak Supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:31:32.218638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:31:32.218638Z digest=sha256:af5c447406c2b3b42fcce7f37e72e55d065b6a8ccefaf11387c2e1a3628793fb

Observation 4f04612a-9ca2-4627-be6c-f3e8eaa918b1 · inbound

Conversations with Andrea: Visitors' Opinions on Android Robots in a Museum cites this paper.

Conversations with Andrea: Visitors' Opinions on Android Robots in a Museum Robust Speech Recognition via Large-Scale Weak Supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T23:57:16.980912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:57:16.980912Z digest=sha256:b33db7f4d96cc0e5ef6b0a244e0c724cc7e9a2355ebd6f8e3515bec28e845513

Observation a54531ca-2667-4c2e-8e03-44a0c2ebe267 · inbound

Boosting CTC-Based ASR Using LLM-Based Intermediate Loss Regularization cites this paper.

Boosting CTC-Based ASR Using LLM-Based Intermediate Loss Regularization Robust Speech Recognition via Large-Scale Weak Supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:18.895300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:03:18.895300Z digest=sha256:99eeb965e35af42220900ff8dd96d5667a7ad39869e5aa2816ca3c50a169ce58

Observation 137ae469-8815-4f1c-9b43-7c514257771e · inbound

Mind the Gap: Entity-Preserved Context-Aware ASR Structured Transcriptions cites this paper.

Mind the Gap: Entity-Preserved Context-Aware ASR Structured Transcriptions Robust Speech Recognition via Large-Scale Weak Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:03:39.096442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:03:39.096442Z digest=sha256:f0399de1161bb5a358f4273e1d87696a834f373ec73c8126b3026032eb02c8ba

Observation 13a632cc-e96a-46d8-820f-e45d952320fc · inbound

AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks cites this paper.

AURA: Agent for Understanding, Reasoning, and Automated Tool Use in Voice-Driven Tasks Robust Speech Recognition via Large-Scale Weak Supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:54.575759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:55:54.575759Z digest=sha256:452eac42bf4f3e3c05b1d67f74848c9af1570a8fe2d9b58b52e2fba2975d881c

Observation 155da7b2-5768-40fa-a998-96514840291a · inbound

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs cites this paper.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Robust Speech Recognition via Large-Scale Weak Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:35.718054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:35.718054Z digest=sha256:12266022c16640d3694f5e939eb1048d45d1c342e2712fb116ebda36f2a7787d

Observation 454caf69-4c5d-4748-8eef-257f5696f746 · inbound

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization cites this paper.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Robust Speech Recognition via Large-Scale Weak Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:44.378445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:44.378445Z digest=sha256:da476e99771b8fdf029d70cabd16b1be3d2b9d47e0a9859ad92fefd92648577d

Observation 430e6464-070f-4b65-97e9-5592c608a025 · inbound

AI Meets Maritime Training: Precision Analytics for Enhanced Safety and Performance cites this paper.

AI Meets Maritime Training: Precision Analytics for Enhanced Safety and Performance Robust Speech Recognition via Large-Scale Weak Supervision

Reference 109

Resolution
unresolved
no resolver link, observed 2026-08-06T20:59:47.174249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:59:47.174249Z digest=sha256:2f7f563d21e8e409beadbb69e8a82e7113a26fb0d55449f232d1d2d897bd1703

Observation eb8410f3-ebf4-48a8-935b-b5172942e000 · inbound

Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla cites this paper.

Adaptability of ASR Models on Low-Resource Language: A Comparative Study of Whisper and Wav2Vec-BERT on Bangla Robust Speech Recognition via Large-Scale Weak Supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:39.740292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:39.740292Z digest=sha256:a8771817940c3284550c164b9ad9f8cb7d635b1ca91f9442226cf38539797b66

Observation d8254148-5783-49b6-b7a7-bfba91d4b099 · inbound

A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations cites this paper.

A Unified Speech LLM for Diarization and Speech Recognition in Multilingual Conversations Robust Speech Recognition via Large-Scale Weak Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:42:23.077634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:42:23.077634Z digest=sha256:78e5f8c6af84b4d5ac6ce16cf870106f99ddc41a2ef4d56d3277f3a21c63f36d

Observation d6393925-b0eb-4ac8-8324-8fe9fdeb1b18 · inbound

A Multistakeholder Approach to Value-Driven Co-Design of Recommender System Evaluation Metrics in Digital Archives cites this paper.

A Multistakeholder Approach to Value-Driven Co-Design of Recommender System Evaluation Metrics in Digital Archives Robust Speech Recognition via Large-Scale Weak Supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:10:38.210134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:10:38.210134Z digest=sha256:2ed12517219355273e660afcab16052648e1d000b317bcef32dcc7a5069de1c2

Observation a707dcb3-367b-43f9-80ab-79c83a0489a6 · inbound

Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis cites this paper.

Multi-Step Prediction and Control of Hierarchical Emotion Distribution in Text-to-Speech Synthesis Robust Speech Recognition via Large-Scale Weak Supervision

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:49:49.081002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:49:49.081002Z digest=sha256:3b496f429178d495392748be8172cc7d240b4d111cae3005f5534adaf5b3d06d

Observation 7ae0f467-f322-4e49-9d98-7ef88b704a61 · inbound

Fine-tuning on simulated data outperforms prompting for agent tone of voice cites this paper.

Fine-tuning on simulated data outperforms prompting for agent tone of voice Robust Speech Recognition via Large-Scale Weak Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:42:02.363973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:42:02.363973Z digest=sha256:ca09fc1fc839df577114f01a7075c5d72fac51633be8b1b693fb646cceba8854

Observation ac687549-46e1-44fe-a6a3-79d5849eca00 · inbound

Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners cites this paper.

Non-Intrusive Binaural Speech Intelligibility Prediction Using Mamba for Hearing-Impaired Listeners Robust Speech Recognition via Large-Scale Weak Supervision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:22:46.219396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:22:46.219396Z digest=sha256:71462b869ab11da87583b3c845458cf513375c2c0d95ef33420fe8513d411270

Observation 5f89e426-80c3-4f7b-889d-7f0670e5d05c · inbound

Differentiable Reward Optimization for LLM based TTS system cites this paper.

Differentiable Reward Optimization for LLM based TTS system Robust Speech Recognition via Large-Scale Weak Supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:04.004618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:04.004618Z digest=sha256:db7370d767e69c46baee6955344f946b30fef833705db4ae51e13ef3d3df0b71

Observation 25e6e748-9c27-4427-be00-9150bb52ee63 · inbound

Pluri-perspectivism in Human-robot Co-creativity with Older Adults cites this paper.

Pluri-perspectivism in Human-robot Co-creativity with Older Adults Robust Speech Recognition via Large-Scale Weak Supervision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:41:36.801092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:41:36.801092Z digest=sha256:8b87bd38746028658b03f578a3e40bb976f93449748974c5b59d5d1652cbf0ce

Observation dfd33c6d-b74f-4e6d-bd96-6d64dce2979e · inbound

Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models cites this paper.

Edge-ASR: Towards Low-Bit Quantization of Automatic Speech Recognition Models Robust Speech Recognition via Large-Scale Weak Supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:33:18.191288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:33:18.191288Z digest=sha256:807c057ce14969b43eb0277e38a2baeb1d744f94021cd275ada0d89b23c1b60b

Observation af316edf-f4b4-4b89-a2d7-6e53281a4ac4 · inbound

Unlocking Speech Instruction Data Potential with Query Rewriting cites this paper.

Unlocking Speech Instruction Data Potential with Query Rewriting Robust Speech Recognition via Large-Scale Weak Supervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:21:34.303507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:21:34.303507Z digest=sha256:4597a45081db68ae25f9880ef93fb7457bcc50e2db8028880b4c98d46fbd6875

Observation 68529c00-d533-40f3-8fdb-19e8ed6916bb · inbound

The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents cites this paper.

The Man Behind the Sound: Demystifying Audio Private Attribute Profiling via Multimodal Large Language Model Agents Robust Speech Recognition via Large-Scale Weak Supervision

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:45:45.153762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:45:45.153762Z digest=sha256:13f09b020f6bb5257fdbb42b7c30ac4e3ba401124511126ba93e05942f60eb06

Observation a0760c96-3840-4fce-95bc-eaa5a0ac0585 · inbound

WhisperKit: On-device Real-time ASR with Billion-Scale Transformers cites this paper.

WhisperKit: On-device Real-time ASR with Billion-Scale Transformers Robust Speech Recognition via Large-Scale Weak Supervision

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T17:29:00.894627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:29:00.894627Z digest=sha256:9b0a70492216709f793833def3800ef45ecd435ab9e361ca7ccea9d1725d7430

Observation b5f830ea-7fcd-4892-abf6-5845fed1ad37 · inbound

Detecting In-Person Conversations in Noisy Real-World Environments with Smartwatch Audio and Motion Sensing cites this paper.

Detecting In-Person Conversations in Noisy Real-World Environments with Smartwatch Audio and Motion Sensing Robust Speech Recognition via Large-Scale Weak Supervision

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:02:58.662409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T05:02:30.603869Z digest=sha256:557983958a16e2492a5c618ec2249970429a0f21999fcf540f8c7ff211500aa5

Observation 574eee85-20ac-447c-a5d2-438fb3272c5f · inbound

Improving Contextual ASR via Multi-grained Fusion with Large Language Models cites this paper.

Improving Contextual ASR via Multi-grained Fusion with Large Language Models Robust Speech Recognition via Large-Scale Weak Supervision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:56:53.397036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:56:53.397036Z digest=sha256:c5f2db5682dd2d561f33a5cf53777592b576c66f898834b9ac2b029fd76d62ad

Observation 42d7cc22-0da8-414d-a31f-95b64b67886d · inbound

Autoregressive Speech Enhancement via Acoustic Tokens cites this paper.

Autoregressive Speech Enhancement via Acoustic Tokens Robust Speech Recognition via Large-Scale Weak Supervision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:24.784828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:24.784828Z digest=sha256:710ccf62b340deeab2cd6e2e5352a3745dae40f72fc0b2e51db9dcd50d77fb47

Observation 67803a5b-440f-4008-8c96-5f47244dae01 · inbound

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech cites this paper.

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech Robust Speech Recognition via Large-Scale Weak Supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T16:33:52.525067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:33:52.525067Z digest=sha256:5861b4b515f9f25ad1a5dfa7b5ba28343e6c620343e20f35c52b2073e12436cd

Observation b28d93b5-22fe-4116-87a3-e81d8192e779 · inbound

FocusView: Understanding and Customizing Informational Video Watching Experiences for Viewers with ADHD cites this paper.

FocusView: Understanding and Customizing Informational Video Watching Experiences for Viewers with ADHD Robust Speech Recognition via Large-Scale Weak Supervision

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T16:29:07.212256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:29:07.212256Z digest=sha256:79c1f2b4b1d0043dfe28ae21c1c23f43a53fbdb7604147892c398e639e977d68

Observation 98eb2d34-7471-4222-b28b-50ef5d3d83d4 · inbound

Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech cites this paper.

Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech Robust Speech Recognition via Large-Scale Weak Supervision

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:42:57.335437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T03:42:32.087391Z digest=sha256:5582d979b222b2c2d64bbf2a2d86acd262cbc17c2699a64c831dfe3980eb99ce

Observation 58adfa33-8e07-49cd-a60c-1edb99d576ea · inbound

From Black Box to Biomarker: Sparse Autoencoders for Interpreting Speech Models of Parkinson's Disease cites this paper.

From Black Box to Biomarker: Sparse Autoencoders for Interpreting Speech Models of Parkinson's Disease Robust Speech Recognition via Large-Scale Weak Supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:53:12.420252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:53:12.420252Z digest=sha256:55d11ba7721655afb411d3348b34bbe39f73bea93514f7f085a5b70eff0960e4

Observation 20b59c6f-0807-417e-80d1-a2def0804837 · inbound

Weak Supervision Techniques towards Enhanced ASR Models in Industry-level CRM Systems cites this paper.

Weak Supervision Techniques towards Enhanced ASR Models in Industry-level CRM Systems Robust Speech Recognition via Large-Scale Weak Supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:53:05.725398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:53:05.725398Z digest=sha256:606b9245147cf4cc0d64f424da2717e145243064ed464ed981f4257f2cb31981

Observation 0584ce77-8125-4cba-948e-db65c007505f · inbound

VIBE: Video-Input Brain Encoder for fMRI Response Modeling cites this paper.

VIBE: Video-Input Brain Encoder for fMRI Response Modeling Robust Speech Recognition via Large-Scale Weak Supervision

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:29.601462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:43:29.601462Z digest=sha256:e6a688c36bc367d35f2fbb0ecfd63778a75d1dd3a5dce77faa3a29f078105c77

Observation 4d150ed8-8054-4d2f-bd69-4896eac0b6c2 · inbound

ProMemAssist: Exploring Timely Proactive Assistance Through Working Memory Modeling in Multi-Modal Wearable Devices cites this paper.

ProMemAssist: Exploring Timely Proactive Assistance Through Working Memory Modeling in Multi-Modal Wearable Devices Robust Speech Recognition via Large-Scale Weak Supervision

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T12:54:29.501960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:54:29.501960Z digest=sha256:9f63dd6992406c17d8f59094d904e6188ca2797d1c071abfc4ef23893ece3515

Observation a5588664-bf72-424d-8ec0-75dcfeda24da · inbound

A blessing or a burden? Exploring worker perspectives of using a social robot in a church cites this paper.

A blessing or a burden? Exploring worker perspectives of using a social robot in a church Robust Speech Recognition via Large-Scale Weak Supervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:58:53.145841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:58:53.145841Z digest=sha256:a5f0436afcb80464a70e4ff08c21a3df659a51fff52336b165dffff5601fa12d

Observation 965fedae-912b-4a28-88f9-092edd4ab9c8 · inbound

Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods cites this paper.

Investigating the Invertibility of Multimodal Latent Spaces: Limitations of Optimization-Based Methods Robust Speech Recognition via Large-Scale Weak Supervision

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T11:12:43.108116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:12:43.108116Z digest=sha256:3e0c250d8b2b51fafc539e93ac74c734191f5f2abb90d6d19dc3f5a9bffadc73