Pith. sign in

Paper Citation Record · LEDGER

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition

As of 19 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2412.15415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15415 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:30:54.980427Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 86e5cc7c-7bd3-496a-9681-17db7bf020ab · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:54.873361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:54.873361Z digest=sha256:dc34ec5ed0095d155eeeacb4b1a0aa8cc0de38964d21887a045eb0dc8e0baddf

Observation cc5cbc68-ec10-4db1-872f-5d5f83056e33 · outbound

This paper cites Strea- mAtt: Direct streaming speech-to-text translation with attention-based audio history selection,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Strea- mAtt: Direct streaming speech-to-text translation with attention-based audio history selection,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.342257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.878094Z digest=sha256:d47a7f169bf4c124ca8b992ff4ae68f388eefb9ae7405a22b6244ebf9efd21ba

Observation 14728ac5-8083-417b-826f-e79a91f97b8a · outbound

This paper cites StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition StreamSpeech: Simultaneous Speech-to-Speech Translation with Multi-task Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:54.882243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:54.882243Z digest=sha256:a68ae07792bddcada2752b20626756a07b214bbd1f3268fb1ee9a1fe36aa8032

Observation 7fea460e-6a89-4646-8529-81cf2d14f195 · outbound

This paper cites A comparative study on end-to-end speech to text translation,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition A comparative study on end-to-end speech to text translation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.330533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.886399Z digest=sha256:d16c18a29eabe338882dd079f3023c1dea7f5c32a6bcf3a1938a1d2ed704dce0

Observation 35307ee9-8d49-4272-b0ac-81e13a8fd4a9 · outbound

This paper cites Navigating the Minefield of MT Beam Search in Cascaded Streaming Speech Translation.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Navigating the Minefield of MT Beam Search in Cascaded Streaming Speech Translation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:54.891114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:54.891114Z digest=sha256:a80375ceb49b24fee079c217f8a42c52dd7187e6422c09846a0d87f38f63cbd1

Observation ebead768-663a-4dac-8207-ba460127a9bb · outbound

This paper cites End-to-end speech translation with knowledge distillation,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition End-to-end speech translation with knowledge distillation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.320074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.895771Z digest=sha256:a11f23a59e0e58d36fad8becbc30fd20d75974094a5821842934eb793a02edde

Observation 6f2d8820-edab-458e-b88c-1ca44e040fad · outbound

This paper cites Multilingual speech translation from efficient finetuning of pretrained models,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Multilingual speech translation from efficient finetuning of pretrained models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.309002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.899883Z digest=sha256:7459305156d6482299566991124ff0abb723409fede0da5226bb8a49da28677e

Observation 8d1beb8e-1fc8-40e6-926b-4f9c09278418 · outbound

This paper cites ComSl: A composite speech-language model for end-to-end speech-to-text translation,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition ComSl: A composite speech-language model for end-to-end speech-to-text translation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.297698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.903817Z digest=sha256:969ff11e11a68e2797cee4ffa491ecfbd80308e551fdb7737cfc2d695078aa8a

Observation 512bad9d-3990-4e07-b1fb-2fb75932973d · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Sequence Transduction with Recurrent Neural Networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:54.907705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:54.907705Z digest=sha256:1e419c102ec46ef88e317937079ee77cb8b674d5dc6fd9027c127439eeb211b8

Observation 0bd826f4-a64a-492e-96bd-c3ac2130af05 · outbound

This paper cites Streaming end-to-end speech recognition for mobile devices,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Streaming end-to-end speech recognition for mobile devices,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.284377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.911672Z digest=sha256:df88456245b8c198eedc9ff07218fbdbcd8eff422c1fbcba32848a4e16e409e1

Observation 5cbcbe64-14b7-4719-b0cc-cd36732ae98c · outbound

This paper cites Dual causal/non- causal self-attention for streaming end-to-end speech recognition,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Dual causal/non- causal self-attention for streaming end-to-end speech recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.273040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.915476Z digest=sha256:d42fe46cc9f22c975e87dd48e982270870696976538287980b09bfc5efc2c6cf

Observation 424d7360-f75c-4898-876a-426b3685ce42 · outbound

This paper cites Token-level serialized output training for joint streaming asr and st leveraging textual alignments,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Token-level serialized output training for joint streaming asr and st leveraging textual alignments,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.261972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.919238Z digest=sha256:a67ca26ae3d771b5986d26fd3f25675b71bbe6a5495ea0c298d73b414cf6811b

Observation 6f2c32db-c960-4501-a99d-2e92998482be · outbound

This paper cites Extended graph temporal classification for multi- speaker end-to-end ASR,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Extended graph temporal classification for multi- speaker end-to-end ASR,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.249537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.922840Z digest=sha256:33ca25319f02b503c087794286415d446dcfddfd0ac4055603baf606a60e2e3c

Observation ad8dca00-33b7-4fc8-8e42-b17d811a1ef8 · outbound

This paper cites Streaming Multi-Talker ASR with Token-Level Serialized Output Training.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Streaming Multi-Talker ASR with Token-Level Serialized Output Training

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:54.925968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:54.925968Z digest=sha256:02282d9706b4da5d76e59c2c0dcb1f8dfc5cac174709486f3a6174cb7501d3d9

Observation 93d656f7-3c2a-4d4d-8737-a6dd803d6adc · outbound

This paper cites Cross attention augmented transducer networks for simultaneous translation,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Cross attention augmented transducer networks for simultaneous translation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.238838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.929095Z digest=sha256:542e6b95d9150e20e1ff2b8462f8a126b7f9417effbad13ee173af3bee246797

Observation 106e0c4a-3475-47ec-b7f3-da04155d88ff · outbound

This paper cites Large- scale streaming end-to-end speech translation with neural transducers,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Large- scale streaming end-to-end speech translation with neural transducers,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.228418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.931776Z digest=sha256:3f8e67f1e76f766b68d84eef2659f0f6e000c65fb903c3d18a1177418895e5f9

Observation 696c84c9-ea74-48a6-b27b-cd479236b3bc · outbound

This paper cites Lamassu: Streaming language-agnostic mul- tilingual speech recognition and translation using neural transducers,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Lamassu: Streaming language-agnostic mul- tilingual speech recognition and translation using neural transducers,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.217231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.934652Z digest=sha256:73185601fb65a36dbea9b7c81e08d6a3332d080a2ba520264632ed6e20084eff

Observation ad90d283-29ac-45ee-a699-17b2ce6f1e76 · outbound

This paper cites Streaming parallel transducer beam search with fast-slow cascaded encoders,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Streaming parallel transducer beam search with fast-slow cascaded encoders,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.206064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.937728Z digest=sha256:d3d97addda099169888ff987bcd845c8696dc1fd04d6936a34796293081d5541

Observation f35e9cf1-7b9f-44d1-99ba-849f091e97d3 · outbound

This paper cites Streaming transformer transducer based speech recognition using non-causal convolution,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Streaming transformer transducer based speech recognition using non-causal convolution,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.195364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.940878Z digest=sha256:c42c2c36b402bf005166f948659a1136f13f044f4165dfd98ffd8ebe194020a8

Observation f8e47b8b-a212-4cbe-80f7-b8ce029c421f · outbound

This paper cites AGADIR: Towards array-geometry agnostic directional speech recognition,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition AGADIR: Towards array-geometry agnostic directional speech recognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.184371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.943885Z digest=sha256:7171d65a457549e53c24396fc0318457df3c2c823b90e5c6299477ecdb950f8e

Observation 8e53f27c-3b32-4ae0-8f62-6da2071d6fff · outbound

This paper cites Directional speech recognition for speaker disambiguation and cross-talk suppression,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Directional speech recognition for speaker disambiguation and cross-talk suppression,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.171162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.946785Z digest=sha256:3e754666720dd767fa1d141bb643e50a83110ac0fde738fb9be4c24238d69b64

Observation 92d476c5-10ee-462a-9040-1b29b027c162 · outbound

This paper cites FLEURS: Few-shot learning evaluation of universal representations of speech,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition FLEURS: Few-shot learning evaluation of universal representations of speech,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.158349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.949920Z digest=sha256:4d83d456786881ea32dba4b47aa09afbccac26dfda0b46a471c8c7b93cb8d175

Observation 49363cbe-58e8-4361-ba8a-d2d6b124f612 · outbound

This paper cites Word alignment by fine-tuning embeddings on parallel corpora,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Word alignment by fine-tuning embeddings on parallel corpora,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.145667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.954909Z digest=sha256:7c4e4fb251f2ba26780b430c815f01b280feb161765734ddf9f0aa620be58711

Observation 7de69e08-4efd-427d-a9c5-8c5d58ce77ba · outbound

This paper cites The CHiME- 8 MMCSG Challenge: Multi-modal conversations in smart glasses,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition The CHiME- 8 MMCSG Challenge: Multi-modal conversations in smart glasses,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.133245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.958868Z digest=sha256:19abd174650262a9d8f2a56cad88a7c9a3ce165d5756575c77df8564e10bd27d

Observation 4fd6157e-9f6e-46e5-9444-1df7ab2e4a8b · outbound

This paper cites CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WEB.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition CCMatrix: Mining Billions of High-Quality Parallel Sentences on the WEB

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T11:30:54.962759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:30:54.962759Z digest=sha256:a450f134475c191eb1b55a5536550c3388d5b4c4eed218ad09a6a1c763094477

Observation ec52ab41-d255-4ceb-ba7b-808fd5c9990c · outbound

This paper cites Beyond english-centric multilingual machine translation,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Beyond english-centric multilingual machine translation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.121279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.967336Z digest=sha256:12d43dc17d68f5ec71e3046424f62c8734cb9e2e43c2321bd81e338153814bac

Observation 3066038c-9a27-45aa-ac6d-121b9fd789fd · outbound

This paper cites Opensubtitles2016: Extracting large parallel corpora from movie and tv subtitles,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Opensubtitles2016: Extracting large parallel corpora from movie and tv subtitles,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.110074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.971009Z digest=sha256:0924be342d72f799cf94b826fb55aafda5c24ef60802828cfe3c079ea438e041

Observation 63b8f64c-4a4e-4437-969e-8ed9e568dda3 · outbound

This paper cites Building subject-aligned comparable corpora and mining it for truly parallel sentence pairs,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Building subject-aligned comparable corpora and mining it for truly parallel sentence pairs,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.098512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.974901Z digest=sha256:5c79690ad5d46351f7f1f595636cbcae63c5b679700b2164b126e6e9d5ceead5

Observation f4bccd77-b549-4c97-a813-75e3c1c64b25 · outbound

This paper cites Chime 8 task 3: Multi-talker word error rate,.

Transcribing and Translating, Fast and Slow: Joint Speech Translation and Recognition Chime 8 task 3: Multi-talker word error rate,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:30:55.085628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T11:30:54.980427Z digest=sha256:5aee05170d049514a22471bc877c75813d16d21750f860492d19256fa35743b4

Pith citing papers

No inbound Pith citation observations are available.