Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:36:00.786445Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 2 inbound Pith citation observations for arXiv:2506.10299.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:36:00.786445Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:35:57.016869Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-10T07:11:53.109624Z
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0c67365c-715a-4f4d-a9cb-6e9e61eb62f5 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 201aee1b-781f-40e8-ac1a-d16ca8846b8d · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Speech-to-speech translation (S2ST) End-to-end speech-to-speech translation (S2ST) systems have been actively studied, which is jointly optimized as a speech-to- speech task [2–9]
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4277612b-a8d2-4dfc-92b8-82741556f477 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Speech-to-speech translation system As shown in Figure 2, we adopt a speech-to-speech translation (S2ST) system fine-tuned from an LLM, as in AudioPaLM [7]
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e6072cdc-72ec-4206-af51-5341a59ebcb4 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 29a1df17-1875-4cf4-b0f5-c04bf5000bd2 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs The CVSS corpus is a widely used corpus for multilingual S2ST, built by speech synthe- sis from the CoV oST2 [36] speech-to-text translation corpus
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5b6d84b8-286b-42c8-8a83-10bcb2e5469e · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs We use interleaved speech–text units as the input and output of LLM, instead of the speech units, during fine-tuning LLM
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 73674bb9-7054-462b-8acb-996cc960ecfb · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs The ATR multilingual speech-to-speech translation system,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6f79be64-ec0e-480b-8638-c10156d67053 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Direct speech-to-speech translation with a sequence- to-sequence model,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation db1f0174-f87b-443f-b0ef-a8640d38851a · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Trans- latotron 2: High-quality direct speech-to-speech translation with voice preservation,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ad958d3a-8311-4ff3-bb1a-66ca0925185c · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Direct speech- to-speech translation with discrete units,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 50be0966-0b4c-4b07-8434-4417e400eb88 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs UnitY: Two- pass direct speech-to-speech translation with discrete units,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8f2a3717-4052-46c7-855c-62ccf491932c · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs SeamlessM4T: Massively mul- tilingual&multimodal machine translation,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 42335768-8cee-481f-a316-74ac65f2a976 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs AudioPaLM: A large language model that can speak and listen,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0d9a00c6-1051-476a-bb3a-d9488c89ec48 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs PolyV oice: Language models for speech to speech translation,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dc4a533c-2a14-457a-a9cf-3f9722ce4c60 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs MSLM-S2ST: A multitask speech language model for textless speech-to-speech translation with speaker style preserva- tion,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 213bbfb2-9406-4881-ae48-23092541654c · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8cf1f694-4944-4e25-a4e9-d4ddc3a0b0df · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre- training,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation bd4740d8-e80e-48d8-87c8-a2829e247d6c · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs SoundStream: An end-to-end neural audio codec,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f33f6157-a8ab-448b-bdf8-c0aa3a4ed88a · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs High fidelity neural audio compression,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3079b3e6-108e-46bf-a7cf-1d20af200453 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Language models are few-shot learners,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9882ea6b-851b-4660-b8c2-53fafa7a4a02 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Prompting large language models with speech recognition abilities,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 75ca9f52-f304-4fc7-b7e6-e728ac618b7c · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs SLM: Bridge the thin gap between speech and text foun- dation models,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 41737303-affc-4460-86a2-7eb37e006e1e · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs On decoder-only architecture for speech-to-text and large language model integration,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 63102fb5-df9e-46ad-b8f1-cfe2e5323c48 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs V oxtLM: Unified decoder-only models for consolidating speech recognition, synthesis and speech, text continuation tasks,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a6619f27-2a8f-4112-b20b-547c52b1d9e1 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Qwen-Audio: Advancing universal audio understanding via unified large-scale audio-language models,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation efd7679b-354e-4eed-b0f1-8d48a8138c50 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7e7c0c7e-ac86-4bc7-afe0-eda957170126 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs SALMONN: Towards generic hearing abilities for large language models,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2d787354-2c2d-4963-a1d9-cd376e6fdb64 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs PaLM 2 technical report,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 219889a7-ed25-4341-b034-a064555d5787 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Bridging the modality gap for speech-to-text translation,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 39ab0872-e6e9-44f0-8922-88a87624cad7 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Push- ing the limits of zero-shot end-to-end speech translation,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 522090b3-3d5d-42f8-a843-51af0e8f6b53 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs SpiRit-LM: In- terleaved spoken and written language model,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 00aafe0a-508e-4e80-a1e1-9eb84e955021 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs The llama 3 herd of models,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f2d33f73-62c6-4ba6-85e0-f8aa3f2866de · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs CVSS corpus and massively multilingual speech-to-speech translation,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0ecd7193-bc8e-4457-a050-0ce6ad047006 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Speech resynthesis from discrete disentangled self-supervised representations,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 08d467e9-3a37-48fc-a5fe-0a1611504a52 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs HiFi-GAN: Generative adversar- ial networks for efficient and high fidelity speech synthesis,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d92952c-a5af-4fbd-841a-cc4c449d59cf · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs On generative spoken language modeling from raw audio,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e970c5ac-a21b-4d8a-9a40-19c348d83ccc · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs AudioLM: A language modeling approach to audio generation,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b251883f-ff82-4bbe-bb97-bb4c961be324 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Neural codec language models are zero-shot text to speech synthesizers,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d2cfee65-97ea-4016-9dc8-b5575ccb51d0 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Speak for- eign languages with your own voice: Cross-lingual neural codec language modeling,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8e895d53-15e0-4fe5-b322-049f26681bcf · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Scaling speech-text pre-training with synthetic interleaved data,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1ab6d6e6-20eb-4abe-a7b2-8c5e65cb6062 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs CTC-segmentation of large corpora for german end-to-end speech recognition,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4f54195b-4f99-48a7-b2ba-12aed8b44562 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs CoV oST 2 and massively multilingual speech translation,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1310da3b-42f6-487b-8c1d-6c1901b54136 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs ESPnet: End-to-end speech processing toolkit,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6624793d-9b5a-4b8f-ae3e-36b306871312 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Robust speech recognition via large-scale weak su- pervision,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f1a8bf44-a61b-4888-8541-80c4cde1c945 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs UTMOS: Utokyo-sarulab system for voicemos challenge 2022,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d6b071d8-822d-40b2-9a08-1456c7190136 · outbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Investigating decoder-only large language models for speech-to-text translation,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0c67365c-715a-4f4d-a9cb-6e9e61eb62f5 · inbound
Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5532afef-ec45-4110-992d-3bf948ad7354 · inbound
NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.