Pith. sign in

Paper Citation Record · LEDGER

Listen, Attend and Spell

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 21 inbound Pith citation observations for arXiv:1508.01211.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1508.01211 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 21 of 21 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T10:34:37.884102Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T13:29:51.152577Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2aaef789-604b-41d6-9fca-64ddabb807fe · inbound

Self Multi-Head Attention for Speaker Recognition cites this paper.

Self Multi-Head Attention for Speaker Recognition Listen, Attend and Spell

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-25T17:16:04.540716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T17:14:01.218607Z digest=sha256:dc0c8bdcefe7ca8ef39e8348bde7e94b59639f360c646a7a5d2b24c6ef2c6767

Observation ac6ba8a7-c512-4ddd-ba61-07ec11a13021 · inbound

NIESR: Nuisance Invariant End-to-end Speech Recognition cites this paper.

NIESR: Nuisance Invariant End-to-end Speech Recognition Listen, Attend and Spell

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-25T01:50:11.188236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T01:46:36.471156Z digest=sha256:b5a0a4a117e40cbc66afd00b855b88253d7a5c72e85600a53a9ddd2e77525ab5

Observation 0e9a48a5-485f-41b5-82bc-a8ccb9f8ace4 · inbound

Hierarchical Sequence to Sequence Voice Conversion with Limited Data cites this paper.

Hierarchical Sequence to Sequence Voice Conversion with Limited Data Listen, Attend and Spell

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-24T21:34:58.941375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T21:30:16.923017Z digest=sha256:ecaa6598b972a4e32b9bed881fc30d377f3f7709a4025e49c32011953b72dee2

Observation 71f04163-c72a-4ec0-bb69-a9fb53957058 · inbound

Cross-Attention End-to-End ASR for Two-Party Conversations cites this paper.

Cross-Attention End-to-End ASR for Two-Party Conversations Listen, Attend and Spell

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-24T16:26:15.890289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-24T16:26:01.154655Z digest=sha256:b1b74b66548a06197aec698cdcb0ad76c22f0b985ff76f694f2ce61af0d3bc40

Observation bd56e8e4-e5b7-4071-ba58-22ffc2120c10 · inbound

Two-Pass End-to-End Speech Recognition cites this paper.

Two-Pass End-to-End Speech Recognition Listen, Attend and Spell

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T10:34:37.884102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T10:34:37.884102Z digest=sha256:7a26062be757dbce3536a33e3418088fbaa34d70779151cba062aba2fb12c304

Observation 511c37b9-a53a-49c5-b473-8ffcfcebd524 · inbound

In-context Learning and Induction Heads cites this paper.

In-context Learning and Induction Heads Listen, Attend and Spell

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:49:09.898808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T03:49:09.374351Z digest=sha256:dad3a26e22fb26b81249ba742d3084f82e87a0ab18f02e1739127f80c5a0edc5

Observation c57b96bd-cf23-4998-bd38-e087f706bb27 · inbound

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison cites this paper.

Prepending or Cross-Attention for Speech-to-Text? An Empirical Comparison Listen, Attend and Spell

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:18:01.505886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T22:18:01.505886Z digest=sha256:9bd2ef4605978435fe5e1138b95657839a475a0f6ea5037ee8c25449a1f39d20

Observation de5af010-494a-43ef-a812-c3d97c274977 · inbound

Optimizing Speech Multi-View Feature Fusion through Conditional Computation cites this paper.

Optimizing Speech Multi-View Feature Fusion through Conditional Computation Listen, Attend and Spell

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:35:12.962138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:35:12.962138Z digest=sha256:fd046f3ea057b93966320131b1dbfba39931d3f992642bbb4366e4e499f93afd

Observation 8e632cc7-4c56-4e4a-a7aa-50db407fd231 · inbound

Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation cites this paper.

Streaming Speaker Change Detection and Gender Classification for Transducer-Based Multi-Talker Speech Translation Listen, Attend and Spell

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T11:31:04.655943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:31:04.655943Z digest=sha256:31657f1408b051afa12064c416a4139a425938bcb3f408b952d4ac71184c7418

Observation 19e81f0b-d8f4-4201-ad39-e6447b7be2ea · inbound

Aligner-Encoders: Self-Attention Transformers Can Be Self-Transducers cites this paper.

Aligner-Encoders: Self-Attention Transformers Can Be Self-Transducers Listen, Attend and Spell

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T22:31:38.698823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:31:38.698823Z digest=sha256:ad4c38f9b350973044c7ffd701758f94442296e65a3a4bcd2a5a56e1f208de87

Observation b7dc1772-fa5f-4ae5-8a4c-75c29f254617 · inbound

Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation cites this paper.

Leveraging Unit Language Guidance to Advance Speech Modeling in Textless Speech-to-Speech Translation Listen, Attend and Spell

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:34.281997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:34.281997Z digest=sha256:2105932a9bdc4ecf6e4bf10877154689ff7933985d7ebbe03546781f41cba017

Observation 80548279-b02e-4b6a-bfcc-ff886bf7da1c · inbound

Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR cites this paper.

Delayed-KD: Delayed Knowledge Distillation based CTC for Low-Latency Streaming ASR Listen, Attend and Spell

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:28.216789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:28.216789Z digest=sha256:804aebdc71530e3d494c5fbb25148352f691406e1c06c5f7b1e9fc1663df8eb4

Observation 5b655266-39e7-4ef1-8c54-145239a38cc2 · inbound

Improving Speech Recognition of Named Entities in Classroom Speech with LLM Revision and Phonetic-Semantic Context cites this paper.

Improving Speech Recognition of Named Entities in Classroom Speech with LLM Revision and Phonetic-Semantic Context Listen, Attend and Spell

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-19T09:17:13.864764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T09:16:48.512860Z digest=sha256:2a7847926a8ce84dd8a347bc3562d9bc963ebce0e640df5c3eeba90313ba2b6e

Observation 0cb7a811-d6ee-4412-a40a-921c664d0241 · inbound

Learning to See Inside Opaque Liquid Containers using Speckle Vibrometry cites this paper.

Learning to See Inside Opaque Liquid Containers using Speckle Vibrometry Listen, Attend and Spell

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:22:32.429081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:22:32.429081Z digest=sha256:ea02a73600db2f32a124e51ca4d0c1d4b28d7a8614a34c49495933aed4ced56b

Observation 1847bb47-c10b-4e4a-89d5-e65b12cd1fc9 · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization Listen, Attend and Spell

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.806418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-08T12:25:52.847432Z digest=sha256:9de270f95e35806fd143d7377d42b4e31c08bc346282ea45f10d06abd6123a5d

Observation a676cc31-cd1e-4fa4-8156-c44e4ffadf32 · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization Listen, Attend and Spell

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:15:07.882041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T23:14:32.494076Z digest=sha256:d96d255d5574af4e81ffe61630489d32e598e731dcb0f31d2c6d3fed6638ddb0

Observation 5f871b2b-3415-4625-96eb-6522fa53b207 · inbound

MedASR: An Open-Source Model for High-Accuracy Medical Dictation cites this paper.

MedASR: An Open-Source Model for High-Accuracy Medical Dictation Listen, Attend and Spell

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:48.086856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T21:18:56.084055Z digest=sha256:a4976c22c903d7480c7db10cb51a810b4daa75759aab6e6862d50a9859c12ff0

Observation 6eab5b32-aab8-4e2a-a3fe-7d9fdbfc6b2a · inbound

StepAudio 2.5 Technical Report cites this paper.

StepAudio 2.5 Technical Report Listen, Attend and Spell

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.509973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:457d06c83ee856e85cfef2695cef9e463ee3367eee6eefd1a14562f94347ec71

Observation ac47e242-1940-42d7-9f82-5bf8a7f5de21 · inbound

Structure Before Collapse: Transient semantic geometry in next-token prediction cites this paper.

Structure Before Collapse: Transient semantic geometry in next-token prediction Listen, Attend and Spell

Reference 189

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:29:51.153703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-26T05:14:07.208255Z digest=sha256:9e7469e40638a68928ddc764fd91a14fab85f062d2cdd04c36d15e34fdf473f9

Observation c521aeb4-ade8-4f27-87a8-e077145e6a75 · inbound

Generative Testing of Automated Speech Recognition Systems cites this paper.

Generative Testing of Automated Speech Recognition Systems Listen, Attend and Spell

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T15:12:47.255094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:12:47.255094Z digest=sha256:20ec17127d1c9a96b0434423e457514edfe9d8c47682c1c21fe4078098dc5aaf

Observation c4d5d563-3613-49ad-aad0-996cf83555a0 · inbound

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition cites this paper.

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition Listen, Attend and Spell

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:54.677464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:54.677464Z digest=sha256:102de1c45f3ee85fd3a21d0314ffec47701dc91904a73660ef2752dbcacf210d