Pith. sign in

Paper Citation Record · LEDGER

LLaMA-Omni: Seamless Speech Interaction with Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 62 inbound Pith citation observations for arXiv:2409.06666.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.06666 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 62 of 62 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:47:39.105733Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:29:41.322265Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8ddcf9c2-c033-4b24-87c2-3fbb7dc8e4a2 · inbound

VoiceBench: Benchmarking LLM-Based Voice Assistants cites this paper.

VoiceBench: Benchmarking LLM-Based Voice Assistants LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:50:13.971194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T00:50:13.841689Z digest=sha256:99371f496b233c4e0e97bb33a448c8e45a1b1330364fb35ce3ffe963a1fa7982

Observation f6d2f7bb-0212-41b2-821f-8761c3a01feb · inbound

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot cites this paper.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.496972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:4a21143c5087eeca5210f7801baa8c66fb23ce2acf43e360262f799f2112d2ee

Observation 31b3f180-c57e-4f57-a86e-2b8d1f4efb8f · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.661305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:018971b21943500564d24dfd2f7e45fbb74699645ae7dcf5a087a96696efe8b8

Observation 7dc3286d-0748-4cc8-888e-601fe4d49c74 · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.105733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.105733Z digest=sha256:d32c79f819de45d3db1818d70bb434681099481003540ad49eb88188f6fef019

Observation b5002ec4-157b-4ce9-983b-9e285c585c7e · inbound

SparQLe: Speech Queries to Text Translation Through LLMs cites this paper.

SparQLe: Speech Queries to Text Translation Through LLMs LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T22:05:23.804113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:05:23.804113Z digest=sha256:f0c3daec966c58edb43586d8658d9f7cd5617e40a137722ab3b68280aa53ece3

Observation d54a85f0-f0cb-403c-84be-bf3409720cca · inbound

A Preliminary Exploration with GPT-4o Voice Mode cites this paper.

A Preliminary Exploration with GPT-4o Voice Mode LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.322585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.322585Z digest=sha256:e930be0bf98da5fe9d8702f8497efd8bf64b6e77ad93c096031899165a43758b

Observation 2d3ae2fc-a0b0-4740-8024-10f7dbc5bf8e · inbound

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction cites this paper.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.342720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:5b3653f831d9b8e6c3eb497d9f9d7594651d7f603dd50d01473173535b253f7c

Observation 033a91c0-6928-4fc0-91f1-3a1bdd811f0a · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.234568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:afc659157dce9372a2a69dc7d36941de50f25657805391bbad17d06a736433cf

Observation c07ba69b-84a4-4031-a186-ef95e4cb43fe · inbound

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach cites this paper.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:33.966621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:33.966621Z digest=sha256:ef9862b6b3f4770a7783a76d0d2fc67767578af7dd8d4e7d559b6469a2bf5b06

Observation 507fb53f-1c52-4fb5-81b3-7edc9f0ec53b · inbound

S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models cites this paper.

S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:36.938679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:36.938679Z digest=sha256:82ff6baf243a2c1672a81cce187cb0b2b0cb0c783403ed466cdafa9adaf01969

Observation 180b0dc2-a501-411a-aaf2-b27ae8112244 · inbound

ModRWKV: Transformer Multimodality in Linear Time cites this paper.

ModRWKV: Transformer Multimodality in Linear Time LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:34.644120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:34.644120Z digest=sha256:84ba5d6f3ead14e6e87dd36c415f8e948efedb5aaf671150034b73d62d8a0083

Observation dc9e1a5e-b5fa-4f82-8711-d91ee80fa622 · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:01.147173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:01.147173Z digest=sha256:e14f0cf46d11af4516229bce7ba8c1b6dfa9cb301ca79b643664a2ef1ed758e4

Observation 0d13692c-21c8-42a9-bc12-cfe3f818b5c3 · inbound

Speechless: Speech Instruction Training Without Speech for Low Resource Languages cites this paper.

Speechless: Speech Instruction Training Without Speech for Low Resource Languages LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:16.069890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:16.069890Z digest=sha256:beb9cb7624c6c4b845c7ae3adcb7c5749c1273310ae60da8e4aa55d5e1903ec0

Observation 9ca4f5b7-4a33-40d1-aa74-c22a6e2656c5 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.644063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.644063Z digest=sha256:52e2472c99d59a6866fca4539c46e9d8bf7c071fc9a8ad1f72129780edd9812a

Observation a6c1b79e-1f1d-42bd-9d60-646fde200bf9 · inbound

MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding cites this paper.

MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:33.562854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:33.562854Z digest=sha256:fff941d808065f12a38bee7e8ec10ef75c143468933a82b5d50cb27fff76bb37

Observation 132089e7-bcf9-4ff1-ada2-857a733acffe · inbound

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling cites this paper.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.598942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.598942Z digest=sha256:9cbb761e9f238034d534cb83784e34801100678760f96e0f7eb9258aabe1bbbf

Observation d802cbdc-6ea6-4981-8b07-cc34f1544fc9 · inbound

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction cites this paper.

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:30.307140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:01:30.307140Z digest=sha256:4a0c3c88855092294b31b57ba976b0aca3b0dad21f2ddfec6a4b015b3cd72469

Observation d3b73640-6507-408d-a905-1fce1cd35252 · inbound

Universal Visuo-Tactile Video Understanding for Embodied Interaction cites this paper.

Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.344367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.344367Z digest=sha256:2ed77f596e5f9b6ea3d212635626dfe709a052d3e45770e1fe7b03d7bc270f58

Observation ce5592c6-903f-44ff-ae2d-861a45d3c6c1 · inbound

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems cites this paper.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:31.966356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:31.966356Z digest=sha256:69d0391d99e8bda4af195346833c492caf8273d6d0f30fa4124282418e29c8ef

Observation 9277ae40-93a8-402b-a2fb-6f5c118d4c39 · inbound

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction cites this paper.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:35.040733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:35.040733Z digest=sha256:fd249f86170e9d8cf09920b4a6aaaf065e9990922737d95ca531586922dc26c8

Observation a38ced2f-e6bf-428d-9ab4-bc652ad434f4 · inbound

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion cites this paper.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:39.960930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:39.960930Z digest=sha256:0161cb728f340fad8281e1c5e36ed9f14e44606a94354e1cecfd80e619c24896

Observation 3a38c437-6768-4ac8-b6b1-defd92aa8332 · inbound

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant cites this paper.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.008363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.008363Z digest=sha256:8f44c552de1539268e7fcd57c8a0a31b8f0c841235b30c7dd086437cd39e1a4d

Observation 076e0b54-13c8-48fa-aa92-1a3b5b45342a · inbound

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment cites this paper.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.924179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.924179Z digest=sha256:c8f3d02497ca65181a90cb0bf7ff4a6cdac3701e350792cc9b2c34497210703e

Observation 1985f09c-790f-448f-ba7c-612b0d17d029 · inbound

Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models cites this paper.

Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.921978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:08:40.921978Z digest=sha256:c450c3b743933098b29d1743b27f7ad3e87d3eefb6861e9d6e704c5a57ee18e7

Observation d734e001-cd5d-4eee-9cd4-638a837d3b22 · inbound

KERAG_R: Knowledge-Enhanced Retrieval-Augmented Generation for Recommendation cites this paper.

KERAG_R: Knowledge-Enhanced Retrieval-Augmented Generation for Recommendation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:23:47.250006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:23:47.250006Z digest=sha256:3d58992a92f43ab756008b6d970935f476f82ddabdcc4157e25c57b2a33852ba

Observation 60989bf5-f0e4-4c5f-ae7d-85c95a1d3c30 · inbound

Unlocking Speech Instruction Data Potential with Query Rewriting cites this paper.

Unlocking Speech Instruction Data Potential with Query Rewriting LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:21:32.233590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:21:32.233590Z digest=sha256:74da8dd7f8b5e4e002feaa83f00328a20580ef8a88bb8dfae793933444a6bb47

Observation c396d604-debf-4f00-846d-96cf0fba6658 · inbound

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation cites this paper.

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:47:55.871612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:47:55.871612Z digest=sha256:9960e0fbc5163fd87d5440fe0602ccf8d0f51c09cce85f7d164bd3a195dc5f35

Observation f9ecc50f-fee8-4c79-bd65-941214304f01 · inbound

Personalized Socially Assistive Robots With End-to-End Speech-Language Models For Well-Being Support cites this paper.

Personalized Socially Assistive Robots With End-to-End Speech-Language Models For Well-Being Support LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:09:04.402999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:09:04.402999Z digest=sha256:1037976fb7298f634eea011d737d1e333e5fc5a1a4ca618db44f8fca018a073f

Observation cdb8ac00-ddad-47a7-abd9-6780ddf68f6e · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.023035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:5306479e8b2a8cf8c176801b1fb10c373f533fa076a41ca1cb37b1cf588d8a71

Observation 275754a4-9a8c-48fc-8896-d67dea81c1d9 · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.291383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.291383Z digest=sha256:1100c2b79da2f058b3a293b77484f342c104b39e0fa2ecb570f405fce94942ba

Observation 65f20b7f-5213-4063-be88-bbe4978d02b0 · inbound

SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding cites this paper.

SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:56.043822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:56.043822Z digest=sha256:9972c4bd5a87c0f8198ff04a8361da7814af208b82993e7e1c170efec382edec

Observation 60e45a49-bb4c-47c9-8812-24061e498963 · inbound

Dual Information Speech Language Models for Emotional Conversations cites this paper.

Dual Information Speech Language Models for Emotional Conversations LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T21:45:23.907075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:45:23.907075Z digest=sha256:769837001fc02ec95cd6860cf07131068b52c8323eaefe02d524fbdf40338f69

Observation ee928dbd-a44f-4c45-9c7c-8826b58e7890 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:12:54.167660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:ce24b1c0e45c94c20140635affb99a8b54d01b627b927c6fd26c165085486ef8

Observation ddc7832d-be84-4b6a-a021-89c7de4e3344 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:05:30.612129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:ef277b4ac32c97091342ed84ea0daafcc06d2abb706491987c4bddca203cbb65

Observation ff509164-128d-453b-85bd-ae5fef3a68fa · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:51.856505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:51.856505Z digest=sha256:2f34d523dfe9aea4d7b9bbbeaaa7445ec749aa29e2e756df96d9d92729cca073

Observation ef60b9df-dd01-4cfb-a286-ad7713ca6c18 · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.513736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:8aa245531d050653f4443f591dcc47584cdf8ba66830e3f98febdf739944f713

Observation 91a0b4c1-b79f-45db-8bae-2e63819c3e44 · inbound

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations cites this paper.

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:04.355944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:33:04.355944Z digest=sha256:6196d9a2839c3d294e1fba65b946a967ec5eebaa76ebd4cc51a1f89d83d0bf2d

Observation bca84238-c1e2-4017-a014-0ee3f5b7d255 · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:01:24.361876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:0679ddd1a53ed7b9e6ac22a4003e9056694ad2928c8c91f109582506370a54a7

Observation 55ce28e0-4813-4353-8197-81b50a8f88de · inbound

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues cites this paper.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:44.086529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:44.086529Z digest=sha256:2c782034e649d9f4bde762bc9e830a93dfb1c5344baf903ca4652e2f00b16875

Observation 73d20c5c-f39d-4b87-a946-54358ef0a2a3 · inbound

Same Words, Different Judgments: How Preferences Vary Across Modalities cites this paper.

Same Words, Different Judgments: How Preferences Vary Across Modalities LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:36:33.038978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T19:31:51.996480Z digest=sha256:b79ba1d5812ffc41d988154ea5b9aa15d140fb8ea9f33e82bc20885c8b94f6a0

Observation 941f052b-a146-45b8-b2c5-7596fdacdd7f · inbound

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision cites this paper.

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T18:42:05.886184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:42:05.886184Z digest=sha256:047cc1ad31991696b723f3e6d77d4c2d41c61f582226d84fda1a0f77e4dbab79

Observation 8cf905bd-13d5-4852-88e2-5551808ab626 · inbound

Controllable Accent Normalization via Discrete Diffusion cites this paper.

Controllable Accent Normalization via Discrete Diffusion LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T21:25:19.253843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:25:19.253843Z digest=sha256:27ecb2a1ce2e7acdc3a040ea630923f38488439c52aa189d31801317b9469c8a

Observation 18bc937b-afd2-45c0-be4a-79329bb0cbe7 · inbound

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model cites this paper.

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:08.488015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T15:34:50.848124Z digest=sha256:4e90899506f3ee1d255b930dc0541faa1eabfbeac10465839d46ea0b43fce2a3

Observation e57d5363-c6af-4b86-8681-b2d59a1be6c0 · inbound

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization cites this paper.

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:45:44.079790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T18:08:16.051117Z digest=sha256:8fc05f0867134aa7fc33d1b34bfdbe968a8e789b8a16c10406deb75393f97a60

Observation be8c5c01-a85f-4883-87da-1846317b008e · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:46:15.262510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T11:00:52.196039Z digest=sha256:3786440abfc324776fd59169525ad8a72d00ef622dbfb4fc7bbd09af07cf1786

Observation 7cb90703-4bc7-4fcc-9550-4a20e3920691 · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:49.611468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T00:49:26.507281Z digest=sha256:faa917bb475c99e5bd35977638c520d5773056b5cef08e9dcf1ac3a0d27b6699

Observation cf1b243e-3b0c-46ae-8eb7-1a1e3436288e · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:55.866632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:033dc9ae6a6daae1ce484b2f840e9434b15e989dbddda895cb565e732da09f4f

Observation df2398ca-422d-4bb7-b0a5-6b4af5991141 · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.969176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:591caef9175b07b2d9ea0fcf64c408228d7302fe0bc6878354e700019e65c68f

Observation db74770c-bf98-4781-9492-e175e48af0a3 · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:43:55.027484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T02:41:13.583493Z digest=sha256:41cb7c1ea905e0dd1e61c5c6daf7432952580f113c8574492cb6d99ae29bf92a

Observation 4838ffaa-416f-4089-8ce4-244f17c814de · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.361577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:87b1285c4e419f2d90c90ca36b45fe1bfdc9a430724ca370ae87d9df0aa5489b

Observation a6d8c1b2-019d-468f-ad1a-814d5b2a6a22 · inbound

A Survey of Audio Reasoning in Multimodal Foundation Models cites this paper.

A Survey of Audio Reasoning in Multimodal Foundation Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:09:24.368254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T02:08:06.976461Z digest=sha256:b8e3df4d0f2ffc55391f8f87b15065c7a5f90b634027ce262b0699a7f95b9f0c

Observation a43b7391-d3e0-4f63-9793-c00a05313c7a · inbound

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens cites this paper.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:25:59.872963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:75b275069fa1bf35622ebdbdd5a00ea64abb0988906f2c64752af45e29a78d34

Observation 65b49c1b-dac7-4476-b120-6205597d2b6d · inbound

Audio Interaction Model cites this paper.

Audio Interaction Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T10:46:52.413130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T04:57:05.062465Z digest=sha256:179213c2cf4213d0acf7ba46182f08b44eaa90432016e95ef39cb40d004bfd48

Observation eb3d44dd-5b67-4797-96c8-e90f5dc44120 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.662681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:20e123d048298e43da8d12b2809fc48423ce13fd1df3aa23ca77bed6dfd54020

Observation 084cd552-38f0-4ee4-a5de-dac922514b7d · inbound

Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models cites this paper.

Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:57:44.819997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T11:29:36.012349Z digest=sha256:faa59e5b86fe4969077de6cdbd1deb5060c2be120746c56cae70204ca3a13b8e

Observation a53e1ce9-d567-4f85-9ed0-d0ab6702f9c2 · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:12.825126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:c1b5dfd6ca52a4cc5db2a5db829ed2c2040d4713a8b8677a099d7f73ba3f9e08

Observation 6f7c5947-006b-4b80-9cfb-cf4596cd38eb · inbound

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead cites this paper.

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:29:41.323872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T11:42:44.035902Z digest=sha256:a8f8c7f95ce9f435f3ad8f36ddb1d321cdd28070c67262e070e8f83ac8095c1f

Observation 2ae65a48-4f96-4a77-8557-222645a0aa62 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 215

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:42.281653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:506a2e557226eb52ca0097075133a331e29d732e8314a7cad3bce3abdcf13e3d

Observation 6199ec6f-8f95-4bcb-a4d8-b464e14d29db · inbound

Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving cites this paper.

Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T08:09:44.957032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T08:09:44.957032Z digest=sha256:f328e08f2d38bb0c42e349cfe6dbc2a6cd7eaf5999eeab274c9570383f83c4eb

Observation a130e201-e309-4fdf-bb17-135aa45b6047 · inbound

TokAN: Accent Normalization Using Self-Supervised Speech Tokens cites this paper.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:2b9e60f6f8de0a09daebd747360a5f3e30b11d3ec3a787d6fa3b1728b42557a8

Observation 788183f2-e31e-4971-a137-16cb22e0c543 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:13.127289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:13.127289Z digest=sha256:83944c4f97ffc6c4d36c0aa1978b85b8e59d59c5723789e327b0d50d54d696d7

Observation 5e6763b9-2544-45e1-82ae-e2b4c8cbe348 · inbound

Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO cites this paper.

Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T01:56:11.344932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:56:11.344932Z digest=sha256:d58f93ffc41e9f7ca55ed5ca1d8d96e602af7bf78c9c86459fefda4904d23fec