Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 62 inbound Pith citation observations for arXiv:2409.06666.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-08T22:47:39.105733Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T08:29:41.322265Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 8ddcf9c2-c033-4b24-87c2-3fbb7dc8e4a2 · inbound
VoiceBench: Benchmarking LLM-Based Voice Assistants LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f6d2f7bb-0212-41b2-821f-8761c3a01feb · inbound
GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 31b3f180-c57e-4f57-a86e-2b8d1f4efb8f · inbound
VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7dc3286d-0748-4cc8-888e-601fe4d49c74 · inbound
Ola: Pushing the Frontiers of Omni-Modal Language Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5002ec4-157b-4ce9-983b-9e285c585c7e · inbound
SparQLe: Speech Queries to Text Translation Through LLMs LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d54a85f0-f0cb-403c-84be-bf3409720cca · inbound
A Preliminary Exploration with GPT-4o Voice Mode LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d3ae2fc-a0b0-4740-8024-10f7dbc5bf8e · inbound
Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 033a91c0-6928-4fc0-91f1-3a1bdd811f0a · inbound
Kimi-Audio Technical Report LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c07ba69b-84a4-4031-a186-ef95e4cb43fe · inbound
Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 507fb53f-1c52-4fb5-81b3-7edc9f0ec53b · inbound
S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 180b0dc2-a501-411a-aaf2-b27ae8112244 · inbound
ModRWKV: Transformer Multimodality in Linear Time LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc9e1a5e-b5fa-4f82-8711-d91ee80fa622 · inbound
Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d13692c-21c8-42a9-bc12-cfe3f818b5c3 · inbound
Speechless: Speech Instruction Training Without Speech for Low Resource Languages LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ca4f5b7-4a33-40d1-aa74-c22a6e2656c5 · inbound
VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6c1b79e-1f1d-42bd-9d60-646fde200bf9 · inbound
MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 132089e7-bcf9-4ff1-ada2-857a733acffe · inbound
Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d802cbdc-6ea6-4981-8b07-cc34f1544fc9 · inbound
OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3b73640-6507-408d-a905-1fce1cd35252 · inbound
Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce5592c6-903f-44ff-ae2d-861a45d3c6c1 · inbound
Chain-of-Thought Training for Open E2E Spoken Dialogue Systems LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9277ae40-93a8-402b-a2fb-6f5c118d4c39 · inbound
NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a38ced2f-e6bf-428d-9ab4-bc652ad434f4 · inbound
StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a38c437-6768-4ac8-b6b1-defd92aa8332 · inbound
SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 076e0b54-13c8-48fa-aa92-1a3b5b45342a · inbound
TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1985f09c-790f-448f-ba7c-612b0d17d029 · inbound
Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d734e001-cd5d-4eee-9cd4-638a837d3b22 · inbound
KERAG_R: Knowledge-Enhanced Retrieval-Augmented Generation for Recommendation LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60989bf5-f0e4-4c5f-ae7d-85c95a1d3c30 · inbound
Unlocking Speech Instruction Data Potential with Query Rewriting LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c396d604-debf-4f00-846d-96cf0fba6658 · inbound
AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9ecc50f-fee8-4c79-bd65-941214304f01 · inbound
Personalized Socially Assistive Robots With End-to-End Speech-Language Models For Well-Being Support LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdb8ac00-ddad-47a7-abd9-6780ddf68f6e · inbound
Step-Audio 2 Technical Report LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 275754a4-9a8c-48fc-8896-d67dea81c1d9 · inbound
BoSS: Beyond-Semantic Speech LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65f20b7f-5213-4063-be88-bbe4978d02b0 · inbound
SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60e45a49-bb4c-47c9-8812-24061e498963 · inbound
Dual Information Speech Language Models for Emotional Conversations LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee928dbd-a44f-4c45-9c7c-8826b58e7890 · inbound
Training-Free Multimodal Large Language Model Orchestration LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ddc7832d-be84-4b6a-a021-89c7de4e3344 · inbound
Training-Free Multimodal Large Language Model Orchestration LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ff509164-128d-453b-85bd-ae5fef3a68fa · inbound
LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef60b9df-dd01-4cfb-a286-ad7713ca6c18 · inbound
Enhancing Speech Large Language Models through Reinforced Behavior Alignment LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 91a0b4c1-b79f-45db-8bae-2e63819c3e44 · inbound
FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bca84238-c1e2-4017-a014-0ee3f5b7d255 · inbound
StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 55ce28e0-4813-4353-8197-81b50a8f88de · inbound
AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73d20c5c-f39d-4b87-a946-54358ef0a2a3 · inbound
Same Words, Different Judgments: How Preferences Vary Across Modalities LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 941f052b-a146-45b8-b2c5-7596fdacdd7f · inbound
Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8cf905bd-13d5-4852-88e2-5551808ab626 · inbound
Controllable Accent Normalization via Discrete Diffusion LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18bc937b-afd2-45c0-be4a-79329bb0cbe7 · inbound
MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e57d5363-c6af-4b86-8681-b2d59a1be6c0 · inbound
Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation be8c5c01-a85f-4883-87da-1846317b008e · inbound
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7cb90703-4bc7-4fcc-9550-4a20e3920691 · inbound
Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cf1b243e-3b0c-46ae-8eb7-1a1e3436288e · inbound
VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation df2398ca-422d-4bb7-b0a5-6b4af5991141 · inbound
A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 118
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation db74770c-bf98-4781-9492-e175e48af0a3 · inbound
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4838ffaa-416f-4089-8ce4-244f17c814de · inbound
DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a6d8c1b2-019d-468f-ad1a-814d5b2a6a22 · inbound
A Survey of Audio Reasoning in Multimodal Foundation Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a43b7391-d3e0-4f63-9793-c00a05313c7a · inbound
TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 65b49c1b-dac7-4476-b120-6205597d2b6d · inbound
Audio Interaction Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation eb3d44dd-5b67-4797-96c8-e90f5dc44120 · inbound
Watch, Remember, Reason: Human-View Video Understanding with MLLMs LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 084cd552-38f0-4ee4-a5de-dac922514b7d · inbound
Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a53e1ce9-d567-4f85-9ed0-d0ab6702f9c2 · inbound
Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6f7c5947-006b-4b80-9cfb-cf4596cd38eb · inbound
Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2ae65a48-4f96-4a77-8557-222645a0aa62 · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 215
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6199ec6f-8f95-4bcb-a4d8-b464e14d29db · inbound
Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a130e201-e309-4fdf-bb17-135aa45b6047 · inbound
TokAN: Accent Normalization Using Self-Supervised Speech Tokens LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 788183f2-e31e-4971-a137-16cb22e0c543 · inbound
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e6763b9-2544-45e1-82ae-e2b4c8cbe348 · inbound
Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO LLaMA-Omni: Seamless Speech Interaction with Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.