Pith. sign in

Paper Citation Record · LEDGER

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs

As of 21 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 2 inbound Pith citation observations for arXiv:2506.10299.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10299 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:36:00.786445Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:35:57.016869Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T07:11:53.109624Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy41
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0c67365c-715a-4f4d-a9cb-6e9e61eb62f5 · outbound

This paper cites Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:57.016869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:57.016869Z digest=sha256:3c0458aff5d03fc32fa4154f1741b743691224f2c5ebe774ba064139013c1869

Observation 201aee1b-781f-40e8-ac1a-d16ca8846b8d · outbound

This paper cites Speech-to-speech translation (S2ST) End-to-end speech-to-speech translation (S2ST) systems have been actively studied, which is jointly optimized as a speech-to- speech task [2–9].

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Speech-to-speech translation (S2ST) End-to-end speech-to-speech translation (S2ST) systems have been actively studied, which is jointly optimized as a speech-to- speech task [2–9]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:03.223466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:57.058314Z digest=sha256:b6c9ecb58682024da27865958c0da32ae125c95c8caaf9c80982cc49f249aa28

Observation 4277612b-a8d2-4dfc-92b8-82741556f477 · outbound

This paper cites Speech-to-speech translation system As shown in Figure 2, we adopt a speech-to-speech translation (S2ST) system fine-tuned from an LLM, as in AudioPaLM [7].

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Speech-to-speech translation system As shown in Figure 2, we adopt a speech-to-speech translation (S2ST) system fine-tuned from an LLM, as in AudioPaLM [7]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:03.203017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:57.099122Z digest=sha256:fb6369dd1980e3cd3999452a224a5a2f55ca099fe332bbf4efecad1c29041818

Observation e6072cdc-72ec-4206-af51-5341a59ebcb4 · outbound

This paper cites an unresolved cited work.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:36:03.183586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:57.154368Z digest=sha256:dc6e77ba1e35d7f595162fa2563142f123be6ff117d9bd378e15added22eae05

Observation 29a1df17-1875-4cf4-b0f5-c04bf5000bd2 · outbound

This paper cites The CVSS corpus is a widely used corpus for multilingual S2ST, built by speech synthe- sis from the CoV oST2 [36] speech-to-text translation corpus.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs The CVSS corpus is a widely used corpus for multilingual S2ST, built by speech synthe- sis from the CoV oST2 [36] speech-to-text translation corpus

Reference 5

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T04:36:03.168690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:57.229801Z digest=sha256:1045f406a442d04aa9b2ddb2d92cfb175271867d257a077c4801985ffcfb3c7b

Observation 5b6d84b8-286b-42c8-8a83-10bcb2e5469e · outbound

This paper cites We use interleaved speech–text units as the input and output of LLM, instead of the speech units, during fine-tuning LLM.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs We use interleaved speech–text units as the input and output of LLM, instead of the speech units, during fine-tuning LLM

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:03.145119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:57.293763Z digest=sha256:411e524ba410af4aeb54e4306c80bcb216e7e9cd351be448274a1c32c0bb3577

Observation 73674bb9-7054-462b-8acb-996cc960ecfb · outbound

This paper cites The ATR multilingual speech-to-speech translation system,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs The ATR multilingual speech-to-speech translation system,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:03.126804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:57.392935Z digest=sha256:b52fddf80847164bc9233bad48066c77dcaca81d71bc5d1b4ac8bde9e95338e4

Observation 6f79be64-ec0e-480b-8638-c10156d67053 · outbound

This paper cites Direct speech-to-speech translation with a sequence- to-sequence model,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Direct speech-to-speech translation with a sequence- to-sequence model,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:03.106380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:57.443833Z digest=sha256:f66179c0da9145fae9d32afccc62825fe92c8eddc2cf46fe1a7a2a34ce8bce55

Observation db1f0174-f87b-443f-b0ef-a8640d38851a · outbound

This paper cites Trans- latotron 2: High-quality direct speech-to-speech translation with voice preservation,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Trans- latotron 2: High-quality direct speech-to-speech translation with voice preservation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:03.070103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:57.557627Z digest=sha256:3432156a77a7c6c03fb6fc38458fd6061f28ca6648fbdf3afaa2257735726ec9

Observation ad958d3a-8311-4ff3-bb1a-66ca0925185c · outbound

This paper cites Direct speech- to-speech translation with discrete units,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Direct speech- to-speech translation with discrete units,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:03.049986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:57.593692Z digest=sha256:9bd422ac54c2e08dcf00ad68b4a25348ac1a6dcb9e1a32b20dd3992fae132239

Observation 50be0966-0b4c-4b07-8434-4417e400eb88 · outbound

This paper cites UnitY: Two- pass direct speech-to-speech translation with discrete units,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs UnitY: Two- pass direct speech-to-speech translation with discrete units,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:03.032451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:57.691384Z digest=sha256:d6b3d60a29729c262a454193e13bfc7410eed879aba664ebdb391dcf392ce4c5

Observation 8f2a3717-4052-46c7-855c-62ccf491932c · outbound

This paper cites SeamlessM4T: Massively mul- tilingual&multimodal machine translation,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs SeamlessM4T: Massively mul- tilingual&multimodal machine translation,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:03.014533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:57.772800Z digest=sha256:5024a7c5c416001eb5d0ef9de7ed526adff861ad53695a4abb7e2f119f0a72fc

Observation 42335768-8cee-481f-a316-74ac65f2a976 · outbound

This paper cites AudioPaLM: A large language model that can speak and listen,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs AudioPaLM: A large language model that can speak and listen,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.910716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:57.882107Z digest=sha256:0d42b7b1eb061c93bdd14b9fcd244cc17b38ea86571093329dd86f74be58999d

Observation 0d9a00c6-1051-476a-bb3a-d9488c89ec48 · outbound

This paper cites PolyV oice: Language models for speech to speech translation,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs PolyV oice: Language models for speech to speech translation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.669830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:57.936105Z digest=sha256:43fc4f8905dd42d19c96bc70f4d0a06ae7e1cc1cb09b889573efeee6d5c1fbe4

Observation dc4a533c-2a14-457a-a9cf-3f9722ce4c60 · outbound

This paper cites MSLM-S2ST: A multitask speech language model for textless speech-to-speech translation with speaker style preserva- tion,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs MSLM-S2ST: A multitask speech language model for textless speech-to-speech translation with speaker style preserva- tion,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.616522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:58.051680Z digest=sha256:71722257d11e6378106e80a5d833ae61821e4f6afb92f97ff980be289991387a

Observation 213bbfb2-9406-4881-ae48-23092541654c · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.598378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:58.123290Z digest=sha256:479505d5785bc5858ae5ba98d50cefd84a2207686d06393da72a01f17d94c176

Observation 8cf1f694-4944-4e25-a4e9-d4ddc3a0b0df · outbound

This paper cites w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre- training,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs w2v-BERT: Combining contrastive learning and masked language modeling for self-supervised speech pre- training,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.580959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:58.232629Z digest=sha256:cf451cec14d627b559777a83c45ba114879748af2b643590408ea9812b4b5c31

Observation bd4740d8-e80e-48d8-87c8-a2829e247d6c · outbound

This paper cites SoundStream: An end-to-end neural audio codec,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs SoundStream: An end-to-end neural audio codec,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.558873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:58.296214Z digest=sha256:367b74a9ebbdc8c1e6a5f9369b545fd5cb921af76be15fc5c467c3798bdf2aaa

Observation f33f6157-a8ab-448b-bdf8-c0aa3a4ed88a · outbound

This paper cites High fidelity neural audio compression,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs High fidelity neural audio compression,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:58.376350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:58.376350Z digest=sha256:30c20bd1f3c5899e7b5cb36f6b6991f54f91758d9fb1d07fae27ff608579672c

Observation 3079b3e6-108e-46bf-a7cf-1d20af200453 · outbound

This paper cites Language models are few-shot learners,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Language models are few-shot learners,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.521050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:58.481581Z digest=sha256:5468c440d1b955afa0443f8382d7dc6c7bff6ae761828af08a0a84482e6a095f

Observation 9882ea6b-851b-4660-b8c2-53fafa7a4a02 · outbound

This paper cites Prompting large language models with speech recognition abilities,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Prompting large language models with speech recognition abilities,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.498520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:58.594113Z digest=sha256:f003d9fc9654a990dd3e78e94507e4b73b86150dae8aad8a7d8d938f2a26f0cd

Observation 75ca9f52-f304-4fc7-b7e6-e728ac618b7c · outbound

This paper cites SLM: Bridge the thin gap between speech and text foun- dation models,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs SLM: Bridge the thin gap between speech and text foun- dation models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.479961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:58.689948Z digest=sha256:510edbe2120b6b8bd389ba99d8a47ff8fd7cf79370d6a9cba1b989989791deb8

Observation 41737303-affc-4460-86a2-7eb37e006e1e · outbound

This paper cites On decoder-only architecture for speech-to-text and large language model integration,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs On decoder-only architecture for speech-to-text and large language model integration,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.462356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:58.762763Z digest=sha256:6d860460497c1cf5f641062eb52b53c5e58b9eef6e2e1ca51df2ff15a61b568a

Observation 63102fb5-df9e-46ad-b8f1-cfe2e5323c48 · outbound

This paper cites V oxtLM: Unified decoder-only models for consolidating speech recognition, synthesis and speech, text continuation tasks,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs V oxtLM: Unified decoder-only models for consolidating speech recognition, synthesis and speech, text continuation tasks,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.444074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:58.847398Z digest=sha256:b63de9906233e2eb339522e4fcba8f08dc2f1a34508e76c48024330b7ee133e2

Observation a6619f27-2a8f-4112-b20b-547c52b1d9e1 · outbound

This paper cites Qwen-Audio: Advancing universal audio understanding via unified large-scale audio-language models,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Qwen-Audio: Advancing universal audio understanding via unified large-scale audio-language models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.427769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:58.922240Z digest=sha256:80b43e06241539e54df88aa4885128834fef075dd9962f2e0375b4159ce6784f

Observation efd7679b-354e-4eed-b0f1-8d48a8138c50 · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.409782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:59.005127Z digest=sha256:3bf267229fdf6fa40549bf66bc0bb9a31323d9f342b7714947111b0993dc0eab

Observation 7e7c0c7e-ac86-4bc7-afe0-eda957170126 · outbound

This paper cites SALMONN: Towards generic hearing abilities for large language models,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs SALMONN: Towards generic hearing abilities for large language models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.390320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:59.070034Z digest=sha256:7c6fbdf84d083a6a987d432e1abb4cf8403e8e29152a487a30c427f5d8dd138e

Observation 2d787354-2c2d-4963-a1d9-cd376e6fdb64 · outbound

This paper cites PaLM 2 technical report,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs PaLM 2 technical report,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.373985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:59.161674Z digest=sha256:9821410a81913fe73c62b28d694d2cda3e7ba7ac679797e36c7ebc6add881621

Observation 219889a7-ed25-4341-b034-a064555d5787 · outbound

This paper cites Bridging the modality gap for speech-to-text translation,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Bridging the modality gap for speech-to-text translation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.354529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:59.240096Z digest=sha256:d8105e7b1f50bf7eec041f93978f583439d6e5d8110295a7496e1760d49f203f

Observation 39ab0872-e6e9-44f0-8922-88a87624cad7 · outbound

This paper cites Push- ing the limits of zero-shot end-to-end speech translation,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Push- ing the limits of zero-shot end-to-end speech translation,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.334713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:59.352544Z digest=sha256:9422cb249347d2684511e5de71469bb8fdcee4cf9e36330d9744e03afcb4ce03

Observation 522090b3-3d5d-42f8-a843-51af0e8f6b53 · outbound

This paper cites SpiRit-LM: In- terleaved spoken and written language model,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs SpiRit-LM: In- terleaved spoken and written language model,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.318937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:59.431515Z digest=sha256:be3727761c7245504ca5bd8f41f29cec463dada28eaedd42639c374d5ba1e63f

Observation 00aafe0a-508e-4e80-a1e1-9eb84e955021 · outbound

This paper cites The llama 3 herd of models,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs The llama 3 herd of models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.300516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:59.527407Z digest=sha256:5b48470e7d20cd0e9e6ac7448416f67cf49ad9ec81e33cda380c466dc7bb5d29

Observation f2d33f73-62c6-4ba6-85e0-f8aa3f2866de · outbound

This paper cites CVSS corpus and massively multilingual speech-to-speech translation,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs CVSS corpus and massively multilingual speech-to-speech translation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.284842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:59.635644Z digest=sha256:f396536b08e23cd9178ba4e3f5abb18e50d147200cfee3fed54c5d77884fd134

Observation 0ecd7193-bc8e-4457-a050-0ce6ad047006 · outbound

This paper cites Speech resynthesis from discrete disentangled self-supervised representations,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Speech resynthesis from discrete disentangled self-supervised representations,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.268446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:59.703342Z digest=sha256:49393ba0cfe5f325c514a4bae264516fc9e2d61c28c1969b8fce0c3c3267da3a

Observation 08d467e9-3a37-48fc-a5fe-0a1611504a52 · outbound

This paper cites HiFi-GAN: Generative adversar- ial networks for efficient and high fidelity speech synthesis,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs HiFi-GAN: Generative adversar- ial networks for efficient and high fidelity speech synthesis,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:59.801777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:59.801777Z digest=sha256:b5d22977958418bca3294087430b8d4a61fbf01dea14cd59bdb50d1f2ea50681

Observation 8d92952c-a5af-4fbd-841a-cc4c449d59cf · outbound

This paper cites On generative spoken language modeling from raw audio,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs On generative spoken language modeling from raw audio,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.238915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:35:59.907749Z digest=sha256:8df96eb86a5c5698c320b93b6063b36b2daafa08abd74c85f5fa82eff94cd0eb

Observation e970c5ac-a21b-4d8a-9a40-19c348d83ccc · outbound

This paper cites AudioLM: A language modeling approach to audio generation,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs AudioLM: A language modeling approach to audio generation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.222869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:36:00.019228Z digest=sha256:8d591e430320f21e6ac409b563bd064659ae44c359d19aab8795d1c78f1529fc

Observation b251883f-ff82-4bbe-bb97-bb4c961be324 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Neural codec language models are zero-shot text to speech synthesizers,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.207684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:36:00.112987Z digest=sha256:3606c4064c0f491e1d6ca108b9e4ba3b76d9aaee30571e2e0b06c50a189aa9ad

Observation d2cfee65-97ea-4016-9dc8-b5575ccb51d0 · outbound

This paper cites Speak for- eign languages with your own voice: Cross-lingual neural codec language modeling,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Speak for- eign languages with your own voice: Cross-lingual neural codec language modeling,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.190651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:36:00.165920Z digest=sha256:d2f931f69201fb0e7d47c6c55e3cad4c68778cf50b9f0bc4b78ff362b48f05d7

Observation 8e895d53-15e0-4fe5-b322-049f26681bcf · outbound

This paper cites Scaling speech-text pre-training with synthetic interleaved data,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Scaling speech-text pre-training with synthetic interleaved data,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.170997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:36:00.295306Z digest=sha256:ee15c6cd233be62926f131861567a52dd0826da2fc99b5e02a3a47420015253a

Observation 1ab6d6e6-20eb-4abe-a7b2-8c5e65cb6062 · outbound

This paper cites CTC-segmentation of large corpora for german end-to-end speech recognition,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs CTC-segmentation of large corpora for german end-to-end speech recognition,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.151621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:36:00.370584Z digest=sha256:6ad1fca1767ce180e32702d54674b24fe03b4814e07ab338f7285320fdcd5caf

Observation 4f54195b-4f99-48a7-b2ba-12aed8b44562 · outbound

This paper cites CoV oST 2 and massively multilingual speech translation,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs CoV oST 2 and massively multilingual speech translation,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:02.127549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:36:00.470798Z digest=sha256:b776792be7ed3279f57b9c5f0b25b36970e2b5f9e4504d06e009dd1a96945b8f

Observation 1310da3b-42f6-487b-8c1d-6c1901b54136 · outbound

This paper cites ESPnet: End-to-end speech processing toolkit,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs ESPnet: End-to-end speech processing toolkit,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:01.944452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:36:00.567759Z digest=sha256:db4060bce433799e2d7e6e0c91d628a98b74bb1b99f0d6cc7a29033a3c75236f

Observation 6624793d-9b5a-4b8f-ae3e-36b306871312 · outbound

This paper cites Robust speech recognition via large-scale weak su- pervision,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Robust speech recognition via large-scale weak su- pervision,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:01.711048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:36:00.682970Z digest=sha256:d2c00e380cc2880ebc9aa292a7ff9b0b5155f948d43fb17fbf9e06aa1ea225b1

Observation f1a8bf44-a61b-4888-8541-80c4cde1c945 · outbound

This paper cites UTMOS: Utokyo-sarulab system for voicemos challenge 2022,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs UTMOS: Utokyo-sarulab system for voicemos challenge 2022,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:01.329064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:36:00.775392Z digest=sha256:c92ad6b62aaf7497773193406e9714eacde69ff6e8aa8484147fdd2825f26b44

Observation d6b071d8-822d-40b2-9a08-1456c7190136 · outbound

This paper cites Investigating decoder-only large language models for speech-to-text translation,.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Investigating decoder-only large language models for speech-to-text translation,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:36:01.089859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:36:00.786445Z digest=sha256:fe70b35fabfd1caaac53300e03ac537cc3b69f45695069fbfb1711ea2094b1df

Pith citing papers

Observation 0c67365c-715a-4f4d-a9cb-6e9e61eb62f5 · inbound

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs cites this paper.

Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:57.016869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:57.016869Z digest=sha256:3c0458aff5d03fc32fa4154f1741b743691224f2c5ebe774ba064139013c1869

Observation 5532afef-ec45-4110-992d-3bf948ad7354 · inbound

NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages cites this paper.

NaijaS2ST: A Multi-Accent Benchmark for Speech-to-Speech Translation in Low-Resource Nigerian Languages Scheduled Interleaved Speech-Text Training for Speech-to-Speech Translation with LLMs

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:11:53.111298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T07:11:32.500279Z digest=sha256:d99aa88cd70913ef94fb213663aabf89616bcdd0c53f8aec51564b8e9dbd8c45