Pith. sign in

Paper Citation Record · LEDGER

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation

As of 4 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2606.30944.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.30944 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T01:01:24.536821Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact29
  • verified fuzzy22
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 300eb9ff-6fef-4fb1-b473-b40dd573e2f9 · outbound

This paper cites Recent advances in speech language models: A survey.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Recent advances in speech language models: A survey

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.871588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:cb137e6f38b660330a077ec6c3d13fa744aecf8e4c93c6a7db111d4cba662c41

Observation 5fed312d-0459-4563-8504-e014a0f908ed · outbound

This paper cites On The Landscape of Spoken Language Models: A Comprehensive Survey.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation On The Landscape of Spoken Language Models: A Comprehensive Survey

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.883322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:bd0fc780735398de79f1d28dc2e4f06cc0c7159924e0ebca01d3ef792a208ea0

Observation ff8859a0-4e56-439c-bd81-20434726eb29 · outbound

This paper cites Whislu: End-to-end spoken language under- standing with whisper,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Whislu: End-to-end spoken language under- standing with whisper,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.896735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:f72719bc8ac69e4c98601007df67575da390e1a51da75e9305a7908943a4e805

Observation 0ea9fb9d-5818-477d-8733-3062002070c3 · outbound

This paper cites Adapting large language model with speech for fully formatted end- to-end speech recognition,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Adapting large language model with speech for fully formatted end- to-end speech recognition,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.863174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:f5cd8114ea947015ce0397d3c8ce27dd6260e8d1cb717aa9f9a697e34ecceb0d

Observation 389a6a33-c15d-43d9-8cf5-abd64ccdf160 · outbound

This paper cites Speechgpt: Empowering large language models with intrinsic cross- modal conversational abilities.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Speechgpt: Empowering large language models with intrinsic cross- modal conversational abilities

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.874224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:12a28f4b652893967a2ea73f0797ef75c91de4c9c47b45438523dc6e3209eb16

Observation 5e81ae1d-a6f6-4baa-9178-f24eeb577039 · outbound

This paper cites Qwen2-Audio Technical Report.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Qwen2-Audio Technical Report

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.874065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:566bd5416d53ce0b63c146ff0eb2070e91cacd37e24c30623cbecb616f683a9e

Observation 0e0b2984-01f7-4fab-b5c3-ffcd1abad8e1 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.877482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:34f491f4edc5a4bdd83579f52f61fb2797ca1d2060afc83f5d2a3dd6bf1f2585

Observation 93fe1a6b-ad6a-4f55-ad08-6aff920b25bc · outbound

This paper cites Understanding the modality gap: An empirical study on the speech-text alignment mechanism of large speech language models,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Understanding the modality gap: An empirical study on the speech-text alignment mechanism of large speech language models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.868810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:15788fbcda8bacbdaf3c14b627de2b2c16632f27123fbeec63c7e187cc212d04

Observation c4d0555f-3f17-48b2-871e-541e176a93f1 · outbound

This paper cites Alignformer: Modality matching can achieve better zero-shot instruction-following speech-llm,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Alignformer: Modality matching can achieve better zero-shot instruction-following speech-llm,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.855242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:9a8d3e0b680c52da9f5612842c9fa0e41be03aef6d0e597921e682b9bd97ecb6

Observation f041ffa6-11a8-425f-af65-d53aae0d00bb · outbound

This paper cites Closing the gap between text and speech under- standing in llms.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Closing the gap between text and speech under- standing in llms

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.880630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:78bdfd88810bda0707d6c33c0323d20147396183b23c521370636d9cd82fb919

Observation 902335f4-2250-40f8-80a1-404cd4776d29 · outbound

This paper cites Speech discrete tokens or continuous features? a comparative analysis for spoken language understanding in speechllms,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Speech discrete tokens or continuous features? a comparative analysis for spoken language understanding in speechllms,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.851735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:f269a24d58cf20a1cbc3ade90db315fb67be86bdd85db524e2fc5cb8d4656c3d

Observation 2f5244f6-835b-4b1b-aa46-d3a482c1ba55 · outbound

This paper cites Closing the Modality Reasoning Gap for Speech Large Language Models.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Closing the Modality Reasoning Gap for Speech Large Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.861469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:5c2a23d7726f40122c1c42ae754eff4247dfc7898b1c1e97f110eafbe3f66509

Observation 36f65776-7fda-4dcc-8e86-e7f4e60b5a83 · outbound

This paper cites E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.866228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:092486b5b2792b2d256661dc0684373d68c3be7e7fee3bb140a2baa64b94ff49

Observation a5a1814d-23ba-44b9-b394-ba3af7ed3611 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Neural codec language models are zero-shot text to speech synthesizers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.901398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:e7ebc9a6c9e39eefac85786a3ae832f7bb2eaaeff3116b4da1ddd5ca44333473

Observation c9232717-a8ce-44d3-b864-9e668f4e3b41 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.867950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:c6a52c103682f32838a67c5fe1bba1a6656b3617341ddfa255e879fd3af4b4b1

Observation 4ba1e5bb-777c-4576-960b-85610a148913 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.854863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:4012ddf1ec982d9dc65958fced08842ce9ff79594707f7d73c0185c982e22b8f

Observation b5b13e28-7619-4746-90bf-68e753d28afd · outbound

This paper cites Pseudo-autoregressive neural codec language models for efficient zero-shot text-to-speech synthesis,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Pseudo-autoregressive neural codec language models for efficient zero-shot text-to-speech synthesis,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.889481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:6eb0791700a10b2b449d29a41bc3e348b3efcfe62b8cb5b6c1d6efcc3592edcc

Observation d599cbf4-104d-49dd-baff-0f6bad2c5390 · outbound

This paper cites Indextts2: A breakthrough in emotionally expressive and duration- controlled auto-regressive zero-shot text-to-speech,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Indextts2: A breakthrough in emotionally expressive and duration- controlled auto-regressive zero-shot text-to-speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.884335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:10f50774fa3abf799fd71c66f5c93b402cf89ec2214b16bcb50cc1a8d58e82b2

Observation 38a83d7b-9754-4a12-a95c-4ce473a84579 · outbound

This paper cites Fish audio s2 technical report.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Fish audio s2 technical report

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.844200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:acb417738ed2b51fa0a25347b1d95fd13fa59a10324ae15b5cfeda8bf7bb83b2

Observation 7fa710d7-b84e-4d7a-a5e8-d77392b18b4e · outbound

This paper cites Zipvoice-dialog: Non-autoregressive spoken dialogue generation with flow matching,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Zipvoice-dialog: Non-autoregressive spoken dialogue generation with flow matching,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.887023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:35a3a29aa609b29cecab1d0650f20983211d203bd2721bb3630acf4beae50d64

Observation bb39d317-dfdd-4d0c-91ab-35088b251255 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Moshi: a speech-text foundation model for real-time dialogue

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.840599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:bc1ad991a51fc0fd85d08c42f1c809ed3db4746b4f7cb2e3d08f34cb9387285f

Observation 3cbc1c14-2319-4db7-a69a-b13171bb03a9 · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.847968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:ffbe501c3caa75a4a503f60a3ce437124a00836e55a52bb2718dcfdef0d544d0

Observation e39f0f1d-08dc-41f0-9592-42e850563044 · outbound

This paper cites Mimo-audio: Audio language models are few-shot learners.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Mimo-audio: Audio language models are few-shot learners

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.851849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:e18d5c1cd59b0367e4eac5e92df36ed12d2aab8e81d48b8372182e8c1733a5f6

Observation 4b757680-88ed-4970-a735-e69877b02bfa · outbound

This paper cites Kimi-Audio Technical Report.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Kimi-Audio Technical Report

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.857825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:0807df21014363c20743ad9d68ef2077f6d450d5756b89b99d91ec73a4568d74

Observation f667b127-47d2-4a9e-9010-6a67b987b411 · outbound

This paper cites Slam-omni: Timbre-controllable voice interaction system with single-stage training,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Slam-omni: Timbre-controllable voice interaction system with single-stage training,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.879207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:985e3bbe37af2e617f1aa9a508f1efbcb9e292bde75e7a446b9790492334c490

Observation 6284bc4d-a100-4ddd-a54e-bcf61a6de7ff · outbound

This paper cites Fun-audio-chat technical report.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Fun-audio-chat technical report

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.864542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:9a7a41c89c3342825a2a40db139489a708ccfd08b1bdc60d15f0616cf782eeb9

Observation 72f844c2-04dc-4cec-9949-df97b3a813a5 · outbound

This paper cites MOSS-Audio Technical Report.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation MOSS-Audio Technical Report

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.831036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:3ee9fd7b711eeb890ff4aa427d47d33371d1a103fdc485307ab5bf237cba2e2b

Observation dd8dc385-a372-410b-8289-a590a05290c1 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.834436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:6371b5ca8bc636736be90f3d5dbf4f9cf100d30ee62814ae902a6702155a257b

Observation b9d06f2e-9fcc-459b-87e8-153712bfeec4 · outbound

This paper cites LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.816779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:eecb6389ba69675ab407b2229713014345ee1c30a99b3a33564762814fa0c040

Observation 1147bd70-7f4b-4557-a576-e455ba885218 · outbound

This paper cites Step-Audio 2 Technical Report.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Step-Audio 2 Technical Report

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.819741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:e16c1881b9dfa2f80bfc2b45bb7fb0e464ef685d7f4c6a6db95641914a7a5171

Observation 5a8cc085-5f67-4bc4-a36c-60adafae2633 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Qwen2.5-Omni Technical Report

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.809038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:78feb7dd85ba588f6677ddac1552b0852491294a74a651dc6aa68adcc7f90732

Observation 0f5caf8f-c1c8-4e9e-b8fe-ac27a14719bb · outbound

This paper cites Qwen3-Omni Technical Report.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Qwen3-Omni Technical Report

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.805334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:edeef35218714fc4da5168589c24c009a0fa6df55e1c4a0d74f5f6b937f1221d

Observation 32763740-303e-4d41-870f-d964311c304e · outbound

This paper cites MinMo: A Multimodal Large Language Model for Seamless Voice Interaction.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T13:05:45.810988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:5917cc7a2056739960fbc88a00a57312f210c4249412917425597fe59dd331c8

Observation 152cf258-b15f-43c8-81c8-60d6d204e8b2 · outbound

This paper cites Qwen3.5-Omni Technical Report.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Qwen3.5-Omni Technical Report

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.788459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:dd7b38fcda5d77eb5ea3f551c095c4cf346d320ef03c79b64b1f09a3a2e045ce

Observation ad9c68d7-5822-4c77-bd6e-dadc3a363a96 · outbound

This paper cites DeepSeek-V3 Technical Report.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation DeepSeek-V3 Technical Report

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.791055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:3f9c9ac57948af6948bf43dfd2da1930f8dcb0faa08703e66a1cd133efe21ea5

Observation a14c16ae-f3aa-48e1-bf3e-fc1ca63c2bde · outbound

This paper cites Better & faster large language models via multi-token prediction,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Better & faster large language models via multi-token prediction,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.881490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:f041a84c40cecdc4ef71fbdad8c318f3d84b042306767e09397e3c90d34cca25

Observation 646895cf-01a7-4a82-aa94-c4542d973006 · outbound

This paper cites Vita-audio: Fast interleaved cross-modal to- ken generation for efficient large speech-language model.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Vita-audio: Fast interleaved cross-modal to- ken generation for efficient large speech-language model

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.813181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:f6a1155994f3edd1b8e8ed0e77c8d32df1578a8d6c7cc59e5045333dd7b63e7d

Observation ffef7a5f-8d1f-4a94-8c56-2abd76e49ffb · outbound

This paper cites V ocalnet: Speech llm with multi-token prediction for faster and high- quality generation,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation V ocalnet: Speech llm with multi-token prediction for faster and high- quality generation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.894341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:9cacef3c18784f10ae7f6230372a60a4ebcf0d13e446608c81e77a7105c40fbf

Observation 284d6510-4b87-4449-8a25-cf50737a712c · outbound

This paper cites Similarity of neural network representations revisited.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Similarity of neural network representations revisited

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.904137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:dd1600054936379013e392ef88d1c400ec9755752ebdf57a2977200b038d27d9

Observation 54f3c543-76f8-4265-9022-990960a64bf8 · outbound

This paper cites Libriheavy: a 50,000 hours asr corpus with punctuation casing and context,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Libriheavy: a 50,000 hours asr corpus with punctuation casing and context,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.891831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:48df7950fd65caa3bfc57c520a7305409dee531033ace976f4833c68bdd67088

Observation 0a63bf1d-d401-45f8-b874-0cafc1f2d112 · outbound

This paper cites Covost 2 and massively multilingual speech translation,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Covost 2 and massively multilingual speech translation,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.899075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:60f7bc48916b797d8e8604b3bed6540d2dd40044afcf0a1cb99b57e133b9c725

Observation 2a62a2a5-5f38-43e4-a3cf-359a9620df83 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.823291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:a774d97013c26f8dbf14a69c88347c1bd653a4995475e895da8d1f62fa00e565

Observation f53f3c67-8aba-40e1-9f0a-f97c394ae11f · outbound

This paper cites Towards efficient speech-text jointly decoding within one speech language model,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Towards efficient speech-text jointly decoding within one speech language model,

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.827711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:38247f19a58a562e66727a8fdcb05a965727e03f202caf4790985e214896a629

Observation d4fec4bf-a7f2-43ad-9246-5f7d04420371 · outbound

This paper cites Cvss corpus and massively multilingual speech-to-speech translation,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Cvss corpus and massively multilingual speech-to-speech translation,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.876551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:3df38fc36bf9eb739ccaf92bc69a40f1c54686c5f4b31915dbb3a3b459472939

Observation 4ef932e0-924c-4dc8-a955-0ffc0f0082a8 · outbound

This paper cites SLM-S2ST: A multimodal language model for direct speech- to-speech translation,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation SLM-S2ST: A multimodal language model for direct speech- to-speech translation,

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.793848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:7952cc775caadbca2738c1eaa80a8a949e6da037a99911d729abab5c6b3ba1a8

Observation ac4e1bae-a699-4041-a37b-3006168c7580 · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.801831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:3835d2859b9861169b5a4ec8d9599ec64256968ac787ede3c2e0e46a5ed6cc97

Observation 7b7ddd3d-8416-466c-891b-aad929053b09 · outbound

This paper cites Hello gpt-4o,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Hello gpt-4o,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.906505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:90f7863efba71a0e1686d0b0e5e67a14a9d65ceec51577ba34cbc34d941bc2a4

Observation f7551b2a-e575-4830-96ee-99ca274072c9 · outbound

This paper cites Fleurs: Few-shot learning evaluation of universal representations of speech.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Fleurs: Few-shot learning evaluation of universal representations of speech

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.858208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:079a7bfe43f019dc2f1366adb7e3a2491080a739973c0bdf8a2acb1c9dd7c473

Observation 4352eb73-d13b-46e8-9f44-9d7a586a73fb · outbound

This paper cites Ultraeval-audio: A unified framework for comprehensive evaluation of audio foundation models,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Ultraeval-audio: A unified framework for comprehensive evaluation of audio foundation models,

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.797735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:6dda4856dd79228e5eea0d0ef99847296ce3e715f8aa6918c84afa69548d42f9

Observation a4ad798f-ee07-4c15-9f5d-1f73655c1a4f · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:05:45.837610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:6e08fce30ada843504c33b18d74c06e8eda4e021940a0782e29cf1f24e57a144

Observation dbdb48ea-2166-4f36-8e38-5cf7a9447902 · outbound

This paper cites Vocalbench: Benchmarking the vocal conversational abilities for speech interaction models.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Vocalbench: Benchmarking the vocal conversational abilities for speech interaction models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:05:45.870823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:4b5ca5f5cfb8b272edbac20c2c88731a86cd06e2d5dc669c3c4bad399e784599

Observation da0035ce-93ca-427e-bcaa-12191012280f · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation Robust speech recognition via large-scale weak supervi- sion,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T07:43:30.860708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-01T01:01:24.536821Z digest=sha256:56394121d3980d2c4e512b778478a9331ecbf331e5e4dcfc0c0e7b49e7565e59

Pith citing papers

No inbound Pith citation observations are available.