Pith. sign in

Paper Citation Record · LEDGER

Deep Speech: Scaling up end-to-end speech recognition

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:1412.5567.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1412.5567 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:04:11.978406Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T06:15:00.866473Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 446fe6cf-28f0-4211-887b-4ee6cdc0b47e · inbound

Mixed Precision Training cites this paper.

Mixed Precision Training Deep Speech: Scaling up end-to-end speech recognition

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:47:18.502759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-12T10:47:18.426912Z digest=sha256:cbebbad8d7105e68408db3628d88994a78f24340589232ef36960a7162c7d5cc

Observation d98064fe-8d22-4197-85d0-03e411a22c8b · inbound

Deep Learning Scaling is Predictable, Empirically cites this paper.

Deep Learning Scaling is Predictable, Empirically Deep Speech: Scaling up end-to-end speech recognition

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:01:58.419869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:01:58.384077Z digest=sha256:a1a3b55cf0cd3fe3bb8cf1debeaf60650f2bfefb7cf0330d356702207e90386a

Observation 2d7b3c58-e652-4406-889c-3900656cca5b · inbound

Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion cites this paper.

Gated Embeddings in End-to-End Speech Recognition for Conversational-Context Fusion Deep Speech: Scaling up end-to-end speech recognition

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-25T15:05:57.754837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-25T15:04:37.455896Z digest=sha256:c52d50697cda54a333c2502817ae48a65bdafb341b05cfe251e4a90b2e65910b

Observation 51ff81b0-f8f9-42b3-8f3c-dd60da89375a · inbound

Fine-grained robust prosody transfer for single-speaker neural text-to-speech cites this paper.

Fine-grained robust prosody transfer for single-speaker neural text-to-speech Deep Speech: Scaling up end-to-end speech recognition

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:30:32.083727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T08:27:16.910352Z digest=sha256:06cb29eba5556bf207af15a95b637e751da8daaf4be2f463fb9407e735d3b5cc

Observation 82ec4563-a7b2-428f-8d6f-ab4c8f21d13c · inbound

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models cites this paper.

Deepfake audio as a data augmentation technique for training automatic speech to text transcription models Deep Speech: Scaling up end-to-end speech recognition

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-24T07:04:03.119378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T07:01:04.631071Z digest=sha256:b39d3e9d62f918b50ceeb92d1d379cb4799d23baa58a57d5de45a7857904240a

Observation 2271d272-fc1d-4434-95cb-db8f2ab6596e · inbound

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive cites this paper.

Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive Deep Speech: Scaling up end-to-end speech recognition

Reference 239

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T23:04:44.622504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-17T23:04:44.287660Z digest=sha256:e531d8feeaa3d619d2dd35ea67e70cd4dc138f1394b3fd9155a9b85ae9f803d3

Observation 468f7e96-22df-46a9-b9b1-43c97c53aeec · inbound

AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition cites this paper.

AD-AVSR: Asymmetric Dual-stream Enhancement for Robust Audio-Visual Speech Recognition Deep Speech: Scaling up end-to-end speech recognition

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-05T22:04:11.978406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:04:11.978406Z digest=sha256:d2c98236c3e8be8bfc67213769e865ed0a0a3b388e16d93bf1efce64c7fc5661

Observation f4b98128-934c-44e5-b2b2-5a6d7dd6dc0c · inbound

HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis cites this paper.

HM-Talker: Hybrid Motion Modeling for High-Fidelity Talking Head Synthesis Deep Speech: Scaling up end-to-end speech recognition

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T20:23:51.658851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:23:51.658851Z digest=sha256:98f8518381a66b5e08954e5c2a42a7fa55f0067f8d9efac8fd7a5f2db4f9bbb9

Observation 1708bf61-afa7-4ad7-bacb-9eb6202d5a89 · inbound

D^3-Talker: Dual-Branch Decoupled Deformation Fields for Few-Shot 3D Talking Head Synthesis cites this paper.

D^3-Talker: Dual-Branch Decoupled Deformation Fields for Few-Shot 3D Talking Head Synthesis Deep Speech: Scaling up end-to-end speech recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T18:37:00.706707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:37:00.706707Z digest=sha256:01e3aec19ba128cd38ad90a86c0997b7a04631660570984757001d2072f858d2

Observation 319bc9a0-f60c-46cb-bc65-4564129e8a14 · inbound

Gauge-covariant stochastic neural fields: Stability and finite-width effects cites this paper.

Gauge-covariant stochastic neural fields: Stability and finite-width effects Deep Speech: Scaling up end-to-end speech recognition

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T21:46:51.886918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T21:45:08.119232Z digest=sha256:0e01a8a6526101d4fdcf7cebc9f85feb87e31e0292da2da814828de7c616d93b

Observation 409dd226-30de-4ad4-9fab-f88e40ba717c · inbound

Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting? cites this paper.

Revisiting Direct Speech-to-Text Translation with Speech LLMs: Better Scaling than CoT Prompting? Deep Speech: Scaling up end-to-end speech recognition

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:31:14.892893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:28:16.040337Z digest=sha256:4fe28ec3e4de93387815b81013d40ca00a7889ecdd22430885433c53724bdbe7

Observation 98865472-96d1-4f83-9c81-e9aaf9c35a3d · inbound

Sink or SWIM: Tackling Real-Time ASR at Scale cites this paper.

Sink or SWIM: Tackling Real-Time ASR at Scale Deep Speech: Scaling up end-to-end speech recognition

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-16T12:10:53.557431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T12:09:17.087477Z digest=sha256:3c5dbbda80baba466e58ed599cda999b19f946ef1b3446b996bd4d0985844643

Observation 5834117f-1eac-402b-9d87-ed9af41d2a03 · inbound

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation cites this paper.

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation Deep Speech: Scaling up end-to-end speech recognition

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:26.316355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:30:36.462567Z digest=sha256:cd17aef3fcf57b3ba888a8c1dbce58e259086f1560a4c013a340e25ff0b2af28

Observation 1c93f416-1030-47c2-80c0-27d8280e0d47 · inbound

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation cites this paper.

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation Deep Speech: Scaling up end-to-end speech recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T16:38:16.145548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:38:16.145548Z digest=sha256:c492db8421fcff4b5148d843bae821301e7edefcb7ae5ea70b1427e751a7e286

Observation 3de0fc65-9b4e-4993-8c10-bf722f69e27a · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Deep Speech: Scaling up end-to-end speech recognition

Reference 152

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:56.910165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:931c2fb8f0a1dcd8fea0e8ae93117e8ac5160ce0983ba2db4d6af3d62c686143

Observation 80b95312-4df9-4ba7-af79-1b6243102680 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Deep Speech: Scaling up end-to-end speech recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:e2aaf59faac6eee4dfa811f859418d7858f6593f452583753d57f2ec236dd355

Observation 684a4787-3093-4075-aba3-ac871ed0b664 · inbound

SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation cites this paper.

SpeechParaling-Bench: A Comprehensive Benchmark for Paralinguistic-Aware Speech Generation Deep Speech: Scaling up end-to-end speech recognition

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:39:48.467145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T00:39:04.303837Z digest=sha256:a8fe523fe802ece8860426a32f8cf3ffb06486a7d845470c40e3161a5d1f5d2b

Observation 46bcce13-80b7-437c-9177-c3ad1d7c7184 · inbound

End-to-End Intracortical Speech Decoding from Neural Activity cites this paper.

End-to-End Intracortical Speech Decoding from Neural Activity Deep Speech: Scaling up end-to-end speech recognition

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-06-30T14:04:44.710471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T13:58:02.170701Z digest=sha256:37c7d66cfcb3543027d8e7e31c2813ef6c54412aa12e55ee1fb78c033a69e374

Observation 1b17bb28-12cc-4f6a-ad52-05a35e2b0321 · inbound

Data-Efficient On-Policy Distillation for Automatic Speech Recognition cites this paper.

Data-Efficient On-Policy Distillation for Automatic Speech Recognition Deep Speech: Scaling up end-to-end speech recognition

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T12:03:23.194853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T12:02:35.332633Z digest=sha256:90516b7911767ebec68623589d277b4d945289d7f6db0c9230cd5bde44689d2a

Observation b91ef777-9229-463f-887b-e11c6daed347 · inbound

Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition cites this paper.

Beyond Waveform Robustness: Robust Feature-Vocoder Adversarial Attacks on Automatic Speech Recognition Deep Speech: Scaling up end-to-end speech recognition

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T14:47:04.341763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T00:09:19.291564Z digest=sha256:38755ef52b805b9429e8b041ec1dd21a4dcd5d7c7cf44d2d6d0b6109b7fb00a7

Observation 521abf9f-fdb3-4f0c-a2f5-ebdef0e21926 · inbound

Synthetic Audio Generation Framework for Air Traffic Control Speech Recognition cites this paper.

Synthetic Audio Generation Framework for Air Traffic Control Speech Recognition Deep Speech: Scaling up end-to-end speech recognition

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-04T06:19:38.617334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T14:34:38.316387Z digest=sha256:9becb0790797991aa15e721750e498285a00d3b5226624754af26282df8f66b8

Observation ce450b9b-0032-46fa-9cc6-042fa3747e6e · inbound

Spatial Speech Perception Systems: A Survey of Sound Source Localization, Directional Enhancement, and Speech Recognition cites this paper.

Spatial Speech Perception Systems: A Survey of Sound Source Localization, Directional Enhancement, and Speech Recognition Deep Speech: Scaling up end-to-end speech recognition

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-07-03T04:57:37.928585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-03T04:52:36.712797Z digest=sha256:a64b620978955722092ede425899d7228210f9b79093f267f292accd52618ac9

Observation 35ffceb3-8834-404f-9733-71d2c5ca0930 · inbound

Provably Lossless Acceleration of DNN Mutation Testing via Memoization cites this paper.

Provably Lossless Acceleration of DNN Mutation Testing via Memoization Deep Speech: Scaling up end-to-end speech recognition

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T15:08:19.307496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:08:19.307496Z digest=sha256:c9e0cd3dfe216c72d9a6a7238408e4005385c1b6db4536bca1aa10596dd605b6

Observation d5e2d7ca-a4b1-4c47-8366-e319a60e3036 · inbound

Greedy dynamical meta-learning cites this paper.

Greedy dynamical meta-learning Deep Speech: Scaling up end-to-end speech recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T23:37:11.363931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T23:37:11.363931Z digest=sha256:9fa265bb96da452a7ed43800964f8d1e1fb9dd68f273c45dcb30ff853c69c61d

Observation 3104b090-de81-49f1-8ca1-699f4617a492 · inbound

Towards High-Level Semantic Intelligence cites this paper.

Towards High-Level Semantic Intelligence Deep Speech: Scaling up end-to-end speech recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T23:06:30.462774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:06:30.462774Z digest=sha256:a39f71b77bbf6334be685cb75d3c3d33dd031aa5a40d2d5996115e7f72a1fec1

Observation 6a842f46-9408-43b3-b110-82b136ebbb8b · inbound

Automated Numerical Stability Analysis of Deep Learning Operators cites this paper.

Automated Numerical Stability Analysis of Deep Learning Operators Deep Speech: Scaling up end-to-end speech recognition

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T02:20:39.882725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:20:39.882725Z digest=sha256:0666802f93f8c274801149b735b327c2184c4e55e9a4b3930de146353532b656