Pith. sign in

Paper Citation Record · LEDGER

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

As of 19 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 13 inbound Pith citation observations for arXiv:2412.15649.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15649 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:18:40.522869Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:52:07.798602Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:18:12.775752Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved50
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ddbcb98b-f99e-491f-850d-c25103875095 · outbound

This paper cites online" 'onlinestring :=.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training online" 'onlinestring :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.470848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.470848Z digest=sha256:840490b79e6da3dc15b83f563a4904c495efe1e10511b36c463bd9893ea942b1

Observation e83085ff-21c2-441f-8abd-09ae069ae99e · outbound

This paper cites write newline.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.548428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.548428Z digest=sha256:1cc1897f8c2f33d9e7e2faa5b38c0338257e53b70306bb23ff8ee8b78907d1e8

Observation 0243ff5e-c9e0-4320-a928-865ebea1d1ca · outbound

This paper cites GPT-4 Technical Report.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.591288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.591288Z digest=sha256:6e50a78a58f57eb466959115190bba84acb599f91eaa496d5ed9f1b9985629ff

Observation 6646d445-79b5-4725-b853-f2045132588b · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.629693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.629693Z digest=sha256:9c9ff4e7daa28c70cbd459a3d8bbd359717dbef6cf4fda72ee640e97ca8c7ad5

Observation e9ca6d90-c988-43e9-8243-f2763bb735e8 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.633711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.633711Z digest=sha256:d01d6e4b0fe0ab4675944fae3ba47b1700431e81c5798486339b0a1ef630e5ba

Observation 9bedda1e-26fe-4ac7-9716-f8d7f61661de · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Common Voice: A Massively-Multilingual Speech Corpus

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.639543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.639543Z digest=sha256:cdbde9011048e3a2c6bb0ea476f5b155e549fb4f30c8f742c45502dfd908f042

Observation fcb239eb-9466-422a-aa2d-3a257d5f1400 · outbound

This paper cites MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training MT-Bench-101: A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.644586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.644586Z digest=sha256:f49a075f7cc72513bb100d794673f7a6e582875c682b8c8d29615f476a5d199d

Observation 7a12acbc-2928-4026-8100-ffad69bab4e1 · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:18:41.335383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T11:18:39.753307Z digest=sha256:1cab14e7f40da6552ca197ee528a0e04c098aeb2d56057a30b75706be6e483b2

Observation da4b9510-7d37-4deb-ab38-3e54ffa94035 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.782229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.782229Z digest=sha256:e1d3bc454ba65457a738a91df5ed6a4ad7153776665dfd0b787474db9e697141

Observation 9c99d1c8-2276-4e12-bead-355ff292e314 · outbound

This paper cites VoiceBench: Benchmarking LLM-Based Voice Assistants.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training VoiceBench: Benchmarking LLM-Based Voice Assistants

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.785833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.785833Z digest=sha256:1e55aef9da2897b4cefe78739dba966c60d3efa7644c0bfd86b4672cd4a08099

Observation 8bed62e0-65e2-46e4-944f-dbcdc3f73cc3 · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.789782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.789782Z digest=sha256:3b26be696c87fd2caa6389d57b2493bfd8e594c6808a3629efe3c992516714bb

Observation b733b015-4871-4e6b-9d17-0a4ce6535436 · outbound

This paper cites High Fidelity Neural Audio Compression.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training High Fidelity Neural Audio Compression

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.794488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.794488Z digest=sha256:61a3bd6ba00868b3b34fc33ffb9a6bad3077ea7ff872857fced116da515ffdf4

Observation 32a30e75-cec8-4151-8087-a7c8f5fe1346 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Moshi: a speech-text foundation model for real-time dialogue

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.851181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.851181Z digest=sha256:e0190ba2ff557c84741daacbf1aae971f4e0a007ccc31155cdd11e8268be72d5

Observation d130f468-0397-46b1-ae1c-316c7b770213 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.902223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.902223Z digest=sha256:77918024b205688b68b6b0f56ff6f9a376d4dddbc2be60d8633d7309a553e3c2

Observation e9e7e71d-b9aa-4077-b5a0-9a430e071d00 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.907370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.907370Z digest=sha256:43d782725527a951add4c7d175b7c4e86203eb89a89f4c65c02df071fc0dcf68

Observation 68e1af96-30e5-418c-9825-7b78dd791802 · outbound

This paper cites The Llama 3 Herd of Models.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training The Llama 3 Herd of Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.911264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.911264Z digest=sha256:03e29070c3562ae316306261a0e8d964efa45385d3985983c2739a8c61fed4fc

Observation a6b9e071-c877-45b1-afdd-4bd4948362f2 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.915261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.915261Z digest=sha256:863725c9aaabfd3cc3584263eb63b48b9fdb8abedea1f3580391c7bfa9b2091f

Observation 4382f760-a711-4ee5-b5ae-8b9d71ea80c7 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.919302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.919302Z digest=sha256:61ad00a30d9d929af63ff0b4faa759decf7a01741840676b345e68c243313b6a

Observation 39366192-dac6-481f-9ab6-facd55d3f5e5 · outbound

This paper cites A Corpus for Understanding and Generating Moral Stories.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training A Corpus for Understanding and Generating Moral Stories

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.923266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.923266Z digest=sha256:127ac118c1ae94f1922c9d9f2bba6439e78c8a775b0ac18421a4cfa6d8711492

Observation dbc72e4f-adc7-45e5-a90a-d006e925a7e3 · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.927221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.927221Z digest=sha256:c43f48c3b1aefe64721c8e3b88dae9d7b8e3371e27b9e010f6551073e67036be

Observation 09f4424b-ffcd-4ce6-a9b4-1e55d6eb61dc · outbound

This paper cites LCSTS: A Large Scale Chinese Short Text Summarization Dataset.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training LCSTS: A Large Scale Chinese Short Text Summarization Dataset

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-11T11:18:40.971085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T11:18:39.930688Z digest=sha256:d400b97c9363f3c586766d619d1b8dc66dc6268be07dce595757d0299ac33ed8

Observation 85f77c63-dc55-4ddb-942b-4ab0570238d1 · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training WavChat: A Survey of Spoken Dialogue Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.935097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.935097Z digest=sha256:71e8e2d2da4a04e24572cfc2c315cdf8143ecd5fd33dd3c65cc2ec83242b5eff

Observation 82945976-e661-4e4e-9b2f-ec765e720055 · outbound

This paper cites Exploring the Impact of Instruction Data Scaling on Large Language Models: An Empirical Study on Real-World Use Cases.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Exploring the Impact of Instruction Data Scaling on Large Language Models: An Empirical Study on Real-World Use Cases

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.939255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.939255Z digest=sha256:69d67297818a3a25a05058e85d47f5c5a1c062ad55b3c9e1d549a955d0f1d347

Observation f8ebe8fd-1b2e-48bd-a50a-f31102c79e6e · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.945789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.945789Z digest=sha256:742174b376efc3ea11390249994d4be043f8d3eaee50bf08e05799b2d090101f

Observation 330b4b18-429a-430f-b5f6-5bb0986cc305 · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:18:41.279260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T11:18:39.955831Z digest=sha256:6d6a8a7ba5e942df512878da19f7bfecd9108b1682cade4a135a04e0a2f8d60c

Observation abce87c4-abe4-4d13-8ab5-a161a7012173 · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.959731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.959731Z digest=sha256:bb1403d488ede5a4498c2de9cf7e7ec7ae9db0808cd9f2b3ef09000907a8b46a

Observation 6f8872c4-908a-4ffe-bd73-8a315dda5b4b · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.012706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.012706Z digest=sha256:a2f9c9efef1e1c76ee3856d3b4c9513aabc78026b18cca6b7a0f417a49e3c67a

Observation 141e6d54-130f-4f56-bd88-f1b2aac659e1 · outbound

This paper cites Decoupled Weight Decay Regularization.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Decoupled Weight Decay Regularization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.080442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.080442Z digest=sha256:f104a7c5d718f1378823506312569910b7498155d24759d6837a0114c75b496e

Observation 7dd56083-4d5d-40e5-8eee-4edecf9f44c6 · outbound

This paper cites Language Model Can Listen While Speaking.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Language Model Can Listen While Speaking

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.165779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.165779Z digest=sha256:6bfd675f46027d0454b65799e99ae700cb381f455166863bf46cbf968c484fbf

Observation 754d09cb-3d67-460c-8aae-a341229697ad · outbound

This paper cites An Embarrassingly Simple Approach for LLM with Strong ASR Capacity.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training An Embarrassingly Simple Approach for LLM with Strong ASR Capacity

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.272578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.272578Z digest=sha256:33f43ad58be3383ae807b4bcffd555ed51e0166bc8f69382ed57d7b87d9f29e7

Observation 897e0a9c-4bec-447e-a42a-601ff5a3a95c · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.276761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.276761Z digest=sha256:3334ad6dc97b1d842b83991d4ea93975265d4afe5528e4a844253b92295195a6

Observation f56e587c-99c2-48ed-b502-d968d05b5375 · outbound

This paper cites PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training PSLM: Parallel Generation of Text and Speech with LLMs for Low-Latency Spoken Dialogue Systems

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.281167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.281167Z digest=sha256:70f4fb4ae98fed675b3d7ca69eb5b700ec355ff0a83193a0e2c8444435fd528a

Observation 3cb85342-c33a-43e9-8a76-931dcba67608 · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Spirit LM: Interleaved Spoken and Written Language Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.285303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.285303Z digest=sha256:7fd34fdff4adda24a31485c2de305c35b3ac86ac1190f3738e4aaede472390ec

Observation 39fba148-f830-467e-bfa5-0e8ddbb82c74 · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:18:41.244889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T11:18:40.289562Z digest=sha256:d7e8ad1634465f32b6b850106c09418c8e40cbec321fdfbcd0454556d23a957d

Observation 7d62b5d5-6219-4b35-a24d-f4c26ac9e0ff · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T11:18:41.230750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-11T11:18:40.292799Z digest=sha256:299c27508e879eddb8a20fdf9d288eec7cda732f080b96228b0f489bc402ee54

Observation de3a80aa-6cc6-4dfc-ba39-3a686575559b · outbound

This paper cites VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training VoiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.296375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.296375Z digest=sha256:4cb353918c4722a4a940613f03a8c18768b49c6e63db13cf3118c72640a1027f

Observation b6369a44-f92b-422e-86fe-2bc974adc059 · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.299843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.299843Z digest=sha256:8b81213cfd29336332f7c68340307e577b114c55eb53c911a2c378c522652222

Observation caac40a2-f337-4233-9396-d9a0b5b9b09a · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.303273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.303273Z digest=sha256:0f2f80f9fbf6201297cca1a99848f14cf9f939d105f6e1ca3321b6408a2c546c

Observation 6e1924e2-d598-4bf7-9d26-8bc01c011287 · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.306659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.306659Z digest=sha256:fa94dbe2de9a547e07a570d071a46d9405b24b7ea4d065b759bc5b8e6a9273b8

Observation f5a3ae68-a2de-4725-b868-0f426f862264 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.310427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.310427Z digest=sha256:497559c1d5a1daa36ba1b190a42a327190a1f250157cd9e97577e3221d111e07

Observation 0438a759-7569-4c47-b734-ffebbdc8970b · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.313983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.313983Z digest=sha256:a52c1e99e0602208440a7a8140b2289458d9e24ed1bb0100c880afb8befc2bdc

Observation 7c824507-10c2-43dc-9a6c-b270c09b4131 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.343120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.343120Z digest=sha256:7db91057d975a252ec42f3e3dc58cb8b0f9a8d1397521b8845de03528457c6cf

Observation beded3b0-8a18-4c26-abc3-c6c025291f37 · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.421279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.421279Z digest=sha256:88bd932d20a5b0f46178bcdef465b188b179256f0f7479743d95b5cbdc1dc18d

Observation 3955d7b3-f724-43ad-a180-de007bd56207 · outbound

This paper cites Qwen2 Technical Report.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Qwen2 Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.447005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.447005Z digest=sha256:e62cbc05271e90faac89909a78933c8807f96a95109d09675a69df87503d6b27

Observation 6eeafa04-2715-4596-bc18-ff1178ad1d7d · outbound

This paper cites AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training AIR-Bench: Benchmarking Large Audio-Language Models via Generative Comprehension

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.451045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.451045Z digest=sha256:f606da5189a205c576e1fa4f5c995c72c7708181c40cab7f66b74893d910ef43

Observation 57dd6b34-a653-4f7f-bfaa-bb9334055bbe · outbound

This paper cites an unresolved cited work.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.455391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.455391Z digest=sha256:e6482c31b165fc70bb3fce50335c05fa14167fc8e10379aad6b3752264dc3fa8

Observation 81f76123-dd3a-464e-91bc-16ec47d7dc7c · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.459612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.459612Z digest=sha256:31573170b94652d33f2677900bf354cecb8cfeba08037767e3ba562d07a59ff0

Observation a3f8f60c-e12e-49db-9c4a-0b24462bd9c6 · outbound

This paper cites Scaling Speech-Text Pre-training with Synthetic Interleaved Data.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Scaling Speech-Text Pre-training with Synthetic Interleaved Data

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.463008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.463008Z digest=sha256:68acba3e3f7a61b0f03e9d9f176d6f17e18209953ebe02cbe8b8debda0fd790b

Observation e9e83824-f1c5-420a-93dc-a24743da2b36 · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.467445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.467445Z digest=sha256:916841435d1cd41ca93f0eda2b5673bdf2aaf706ec13355e8a9bb3a5e8c40aff

Observation 4128e765-f20e-41c7-abfa-7ff1187d36ef · outbound

This paper cites OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training OmniFlatten: An End-to-end GPT Model for Seamless Voice Conversation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.471370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.471370Z digest=sha256:18a96df40f71efce0521aeff3600e83e513e713b62fa136aaa85d75d12349b6b

Observation 3d875b17-fcfd-4be9-be9d-05520b2506d5 · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.522869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.522869Z digest=sha256:cc0d063772d90f766eeec7e638f37caf61ab2b3d3e15b279a47f1102d2c2ec05

Pith citing papers

Observation 8a279d26-5421-4784-a928-f7085b96354b · inbound

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey cites this paper.

Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-11T19:32:24.029868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:32:24.029868Z digest=sha256:a554d130dc9a17cc4a7bb61935a0926b2e00c27be9573cb482232c38a3419746

Observation e4443bba-4982-4deb-b57b-e7ec85031f24 · inbound

Real-Time Textless Dialogue Generation cites this paper.

Real-Time Textless Dialogue Generation SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T21:27:45.525857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:27:45.525857Z digest=sha256:1674ef79475f7bf6e488d69075fa0df4485e90b8a033996e2c4b2dd9f516c8f0

Observation a214485f-e063-4ae5-80fa-d8217ce3ef7a · inbound

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her cites this paper.

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-10T13:34:25.641434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:34:25.641434Z digest=sha256:f80de536a40e9c6f5be16cbd64eed79a8bc674a068af7bc0fe7890d8aa6ea3be

Observation 370d4204-988e-4b50-a439-f59ff1b8a613 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:45:08.035913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:f3dd44692610bbdeee55d629062cfef0f0069c7b0c503d936cc1d6ff1a98e579

Observation b3c951d7-c019-4bfc-9812-a5b0b8cd35d8 · inbound

LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis cites this paper.

LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:52:07.798602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:52:07.798602Z digest=sha256:e603b207480ddb206a258d739e249f815228fb7a18941b389d3d5abd227857c2

Observation 5e189fb4-7b27-416c-932d-4bf450e13f64 · inbound

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation cites this paper.

SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:49:58.254721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:49:58.254721Z digest=sha256:245eeb3128ac89105dbd496d9940367e2b3a9cb5bc3467f8962ec1789195bc2a

Observation 20ba2d0d-0fac-499b-a766-14073c675d5b · inbound

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model cites this paper.

Stream-Omni: Simultaneous Multimodal Interactions with Large Language-Vision-Speech Model SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:33:19.356530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:33:19.356530Z digest=sha256:dda458dc8c161dd142a69442bfe3e5f7bc23d7fbeb96ba338ffcf98ba7ddbec4

Observation ce890845-c7a6-471c-a0cb-d32897a5357a · inbound

FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems cites this paper.

FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T18:06:21.278681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:06:21.278681Z digest=sha256:ef93ad8a28c1a0d96c5bb36f3e3bddf9242ba846358dae2dbb33b4d056b18852

Observation 81f83c4b-45f8-43b8-9d37-016b0153cecf · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.772401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.772401Z digest=sha256:30b689ded49bfce9742d6f4e1d05c217d9fcb50b7613b4ebb240d2e2d1a7615e

Observation 619f1271-d061-4ea1-9907-7f6fc93723ca · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:25:30.001718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:ec30f824102acb4c1a8019b724f3c6337f3a827ca43a3ef025ed12ba8e64d17d

Observation 9cb079e0-7989-4bc5-bb9d-7dc195febb24 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.237941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:8d3b17f3e1c4317331243b4d0de4a0ad79b2044bcee0c9dae082ba560bf4b3fa

Observation babbd305-f199-4614-9712-658ef4f26b97 · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:18:12.777326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:3974507fec67d94710ab39838c69f3f50e80fcdbed562762b9ac943cee5669cd

Observation 46fb2046-e84d-4c07-840c-83b750e7cd33 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:08.148931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:08.148931Z digest=sha256:1f57472fb5cbfc3cf998a66175aa16af572cfb107b336cdbd63f535e95bfeeee