Pith. sign in

Paper Citation Record · LEDGER

Listen, Attend and Spell

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:1508.01211.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1508.01211 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T10:34:37.884102Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T13:29:51.152577Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2aaef789-604b-41d6-9fca-64ddabb807fe · inbound

Self Multi-Head Attention for Speaker Recognition cites this paper.

Self Multi-Head Attention for Speaker Recognition Listen, Attend and Spell

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-25T17:16:04.540716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T17:14:01.218607Z digest=sha256:2d91422a7b0a1bbe2d8ea084418e4603dbd493dc9c2b95a0f19b2d3f19a57ca0

Observation ac6ba8a7-c512-4ddd-ba61-07ec11a13021 · inbound

NIESR: Nuisance Invariant End-to-end Speech Recognition cites this paper.

NIESR: Nuisance Invariant End-to-end Speech Recognition Listen, Attend and Spell

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-25T01:50:11.188236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T01:46:36.471156Z digest=sha256:4047b9ee8e199772ea74f556e2c3447738a5ca2b8e15cad243a2738f73d0ab1f

Observation 0e9a48a5-485f-41b5-82bc-a8ccb9f8ace4 · inbound

Hierarchical Sequence to Sequence Voice Conversion with Limited Data cites this paper.

Hierarchical Sequence to Sequence Voice Conversion with Limited Data Listen, Attend and Spell

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-24T21:34:58.941375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T21:30:16.923017Z digest=sha256:d72c8126c37fa81009b7ccef74e758107ec9bebf78c298ec714d5c169c049827

Observation 71f04163-c72a-4ec0-bb69-a9fb53957058 · inbound

Cross-Attention End-to-End ASR for Two-Party Conversations cites this paper.

Cross-Attention End-to-End ASR for Two-Party Conversations Listen, Attend and Spell

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-24T16:26:15.890289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T16:26:01.154655Z digest=sha256:5d7479283610d22413d0e7782daf79474cc63c82e9c9e9ec47d80d50b87da4b1

Observation bd56e8e4-e5b7-4071-ba58-22ffc2120c10 · inbound

Two-Pass End-to-End Speech Recognition cites this paper.

Two-Pass End-to-End Speech Recognition Listen, Attend and Spell

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T10:34:37.884102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:34:37.884102Z digest=sha256:760c0c183673154942202bc3bf3d102e3cb776acdf2d5689152334a818b3178a

Observation 511c37b9-a53a-49c5-b473-8ffcfcebd524 · inbound

In-context Learning and Induction Heads cites this paper.

In-context Learning and Induction Heads Listen, Attend and Spell

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:49:09.898808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T03:49:09.374351Z digest=sha256:055d1bc9e985eb72985c0746cc977a9a4feef41fb913f3fb0305873fe876e753

Observation c57b96bd-cf23-4998-bd38-e087f706bb27 · inbound

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison cites this paper.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Listen, Attend and Spell

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.505886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.505886Z digest=sha256:ff1dd8aad77165c6483d96fd5d6058027a5ab4a0831c51ee25478a1b8433f119

Observation de5af010-494a-43ef-a812-c3d97c274977 · inbound

Optimizing Speech Multi-View Feature Fusion through Conditional Computation cites this paper.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Listen, Attend and Spell

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:12.962138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:12.962138Z digest=sha256:a9fb6ba42b50734b3a7fd515bebb7f97990506e09c154e9ec04e6c3a1e501335

Observation 8e632cc7-4c56-4e4a-a7aa-50db407fd231 · inbound

Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation cites this paper.

Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation Listen, Attend and Spell

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T11:31:04.655943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:31:04.655943Z digest=sha256:fdb1d8119764bc371fc6c7464d738e52b2dd875eca2d797905ede91b309e029f

Observation 19e81f0b-d8f4-4201-ad39-e6447b7be2ea · inbound

Aligner-Encoders: Self-Attention Transformers Can Be Self-Transducers cites this paper.

Aligner-Encoders: Self-Attention Transformers Can Be Self-Transducers Listen, Attend and Spell

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:38.698823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:31:38.698823Z digest=sha256:685c12568099e46f0ff9634da36ecd753e99b6bc163ce1b574414d592525b721

Observation b7dc1772-fa5f-4ae5-8a4c-75c29f254617 · inbound

Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation cites this paper.

Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation Listen, Attend and Spell

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:34.281997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:34.281997Z digest=sha256:2336a6f21960363b22b8605ceeae5ddb6faab0fc4856b329463e6fed995e2ff3

Observation 80548279-b02e-4b6a-bfcc-ff886bf7da1c · inbound

Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR cites this paper.

Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR Listen, Attend and Spell

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:28.216789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:28.216789Z digest=sha256:81c6d1c49e0f19f4d36e6f7885c5b7af863e3c0999590a9e5767211349c1725f

Observation 5b655266-39e7-4ef1-8c54-145239a38cc2 · inbound

Improving Speech Recognition of Named Entities in Classroom Speech with LLM Revision and Phonetic-Semantic Context cites this paper.

Improving Speech Recognition of Named Entities in Classroom Speech with LLM Revision and Phonetic-Semantic Context Listen, Attend and Spell

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:17:13.864764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T09:16:48.512860Z digest=sha256:6d3c74c14311e3dbbd00c7dde882e931911e8549e7ca0e0328ae998d0914f846

Observation 0cb7a811-d6ee-4412-a40a-921c664d0241 · inbound

Learning to See Inside Opaque Liquid Containers using Speckle Vibrometry cites this paper.

Learning to See Inside Opaque Liquid Containers using Speckle Vibrometry Listen, Attend and Spell

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:22:32.429081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:22:32.429081Z digest=sha256:c0d96d73ac6944ffca102dd372e7ff89630b0ff63ee0f9d67ce347f40a47b5e6

Observation 1847bb47-c10b-4e4a-89d5-e65b12cd1fc9 · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization Listen, Attend and Spell

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.806418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T12:25:52.847432Z digest=sha256:9a8bb04ff999e108532f77ab928506d5332566aa14edb02af3d44b85dd81dc92

Observation a676cc31-cd1e-4fa4-8156-c44e4ffadf32 · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization Listen, Attend and Spell

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:15:07.882041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T23:14:32.494076Z digest=sha256:f7998940498388e2035392fb5c672563a012ae9f06273852c2903f10179ec994

Observation 5f871b2b-3415-4625-96eb-6522fa53b207 · inbound

MedASR: An Open-Source Model for High-Accuracy Medical Dictation cites this paper.

MedASR: An Open-Source Model for High-Accuracy Medical Dictation Listen, Attend and Spell

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:48.086856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T21:18:56.084055Z digest=sha256:e306ac705e01a9ce63cccadfeebd1e57e25f1ff3526ce7fafa7a2bd3f3c70676

Observation 6eab5b32-aab8-4e2a-a3fe-7d9fdbfc6b2a · inbound

StepAudio 2.5 Technical Report cites this paper.

StepAudio 2.5 Technical Report Listen, Attend and Spell

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.509973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:a7b83b0a4afc15e4be26c5589e54e58e2291e18aab4b5e60fb5b589255809376

Observation ac47e242-1940-42d7-9f82-5bf8a7f5de21 · inbound

Structure Before Collapse: Transient semantic geometry in next-token prediction cites this paper.

Structure Before Collapse: Transient semantic geometry in next-token prediction Listen, Attend and Spell

Reference 189

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:29:51.153703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T05:14:07.208255Z digest=sha256:a197099fe947e627fe7a607ccd84e3631a3b1dff42be42fa1f9701f4fc8ed74b

Observation c521aeb4-ade8-4f27-87a8-e077145e6a75 · inbound

Generative Testing of Automated Speech Recognition Systems cites this paper.

Generative Testing of Automated Speech Recognition Systems Listen, Attend and Spell

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T15:12:47.255094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:12:47.255094Z digest=sha256:c0eb80803a5269eabb6105db6ba63d2a84df3d592035886b71dc1f6039ee5413

Observation c4d5d563-3613-49ad-aad0-996cf83555a0 · inbound

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition cites this paper.

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition Listen, Attend and Spell

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:54.677464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:54.677464Z digest=sha256:940c142e819d3ace550a7282c2f8b1e540aba439f46eb9df2dbfee4ad8bfc0f6