Pith. sign in

Paper Citation Record · LEDGER

LLaMA-Omni: Seamless Speech Interaction with Large Language Models

As of 19 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 84 inbound Pith citation observations for arXiv:2409.06666.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.06666 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 84 of 84 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 84 of 84 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:33:47.104090Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:29:41.322265Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8ddcf9c2-c033-4b24-87c2-3fbb7dc8e4a2 · inbound

VoiceBench: Benchmarking LLM-Based Voice Assistants cites this paper.

VoiceBench: Benchmarking LLM-Based Voice Assistants LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:50:13.971194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-17T00:50:13.841689Z digest=sha256:7d11c0c44289e94d233f1f4767edac5a3ab221a59648e49113c9c3512e016e51

Observation c7753991-651d-4217-b129-458db8c2962a · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.240429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.240429Z digest=sha256:5055c1e546bcf97d193d2fa1f033759067027a55ecfc8a14a4021b3eec2ac054

Observation b9dd0ca6-42d6-4fdf-8f04-b1419870419b · inbound

Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models cites this paper.

Who Can Withstand Chat-Audio Attacks? An Evaluation Benchmark for Large Audio-Language Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T14:53:50.206384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T14:53:50.206384Z digest=sha256:7b6f181de10a1ac2dd275c0e57c037eca33b6741b092586868d77eafb7fb028e

Observation 6b07796b-78fe-467d-88ca-5518f05d3131 · inbound

Scaling Speech-Text Pre-training with Synthetic Interleaved Data cites this paper.

Scaling Speech-Text Pre-training with Synthetic Interleaved Data LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T12:02:40.403597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:02:40.403597Z digest=sha256:f4156884a52cf3384a2b31e0e9d68a71be558da95165390354ae830fe0511418

Observation 793b710c-ca7f-4778-9085-faa4fee8132c · inbound

SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation cites this paper.

SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T11:31:16.076477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:31:16.076477Z digest=sha256:fa12c1e651013400e75608429235d0e3783b2bcbfccf528d30a826216431d387

Observation f6d2f7bb-0212-41b2-821f-8761c3a01feb · inbound

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot cites this paper.

GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:53:47.496972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T03:53:47.396742Z digest=sha256:c735e5319d35f5312657dab526d4e3d018a0ea1c38c6ea0decbf4954db476725

Observation 5267b4e2-7a21-48c7-9c16-7b84feba1748 · inbound

Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners cites this paper.

Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T21:12:23.997188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:12:23.997188Z digest=sha256:8df89f3bcff97f684fb57a3f87e7042ad3eb514b2945b7114f07c6e320bb39f3

Observation 1740c986-f208-4950-9726-1a7c91a48106 · inbound

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls cites this paper.

YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:17:49.010166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:17:49.010166Z digest=sha256:41770ead51abfd940e3cc4d6045259e6873257cc1048503df3a422efda1ab70e

Observation 02f9e4d8-1345-404a-8365-fde8e928e8a7 · inbound

Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models cites this paper.

Typhoon 2: A Family of Open Text and Multimodal Thai Large Language Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T12:56:13.082251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:56:13.082251Z digest=sha256:bb303a2bbaec743586d22e0930d9a72bbbf8c518bcbcb8e794d6e1bb0250ddec

Observation a6b9e071-c877-45b1-afdd-4bd4948362f2 · inbound

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training cites this paper.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:39.915261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:39.915261Z digest=sha256:863725c9aaabfd3cc3584263eb63b48b9fdb8abedea1f3580391c7bfa9b2091f

Observation 1873b3af-7217-48c6-baad-3c2aa10ae3a5 · inbound

Contrastive Learning for Task-Independent SpeechLLM-Pretraining cites this paper.

Contrastive Learning for Task-Independent SpeechLLM-Pretraining LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:13:48.447978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:13:48.447978Z digest=sha256:267b4cf2dae4ab53bf3197a2af712f2e6aae6398e5b14544631b23a4715e330c

Observation 4ea0a9ce-51d8-465f-bc6c-460ce4fd60b6 · inbound

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback cites this paper.

Align Anything: Training All-Modality Models to Follow Instructions with Language Feedback LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:09:13.058336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:09:13.058336Z digest=sha256:932cc7870094c2f80869e7523d233f0d2c945f2caf9532781b9295f3572eb2c1

Observation 09b7180b-b8da-4e02-9e10-587fb06203fa · inbound

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization cites this paper.

SilVar: Speech Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T10:18:36.175024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:18:36.175024Z digest=sha256:cd542afa903818d3b98851aeb399c8389d6166814c1ad32779aebbac0276be95

Observation e1abac17-8c84-4617-9f2a-9005233c60d1 · inbound

Trading Devil RL: Backdoor attack via Stock market, Bayesian Optimization and Reinforcement Learning cites this paper.

Trading Devil RL: Backdoor attack via Stock market, Bayesian Optimization and Reinforcement Learning LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 169

Resolution
unresolved
no resolver link, observed 2026-08-11T05:12:12.117189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T05:12:12.117189Z digest=sha256:b038ed87e31d7ce0079fbbfee7f4cee27ab218238ac55703dd6a03b25ea2ef0f

Observation dbb3747d-9a84-433b-8943-faf1b8742dcc · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 119

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.635845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.635845Z digest=sha256:952c721f9e9e9fbfe45e28d0793c8f8123784afdb7d5df0a7c79b1fd65b2aafb

Observation 31b3f180-c57e-4f57-a86e-2b8d1f4efb8f · inbound

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction cites this paper.

VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:08:19.661305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-17T21:08:19.570050Z digest=sha256:d36247c93503d239958c63316cdbac68a1af2064f6f0df6713f365d1b601ac0d

Observation 5b296ea0-6cfe-4575-8c7e-66e7714aa4ac · inbound

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction cites this paper.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.674083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.674083Z digest=sha256:66dbba78e59aca13ef9a534285e271763f920a319a004e82a4f873e28d86aedf

Observation 5ad67b38-507a-4932-9a0b-a80be510e8b9 · inbound

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding cites this paper.

HumanOmni: A Large Vision-Speech Language Model for Human-Centric Video Understanding LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T14:40:55.207606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:40:55.207606Z digest=sha256:231a4ea0bcb6687a14ddb7102213d3e0548e4f4b2d96552a17dbd98506fcda5f

Observation 2b0625f8-1a5c-4afe-88c1-573aebba36b8 · inbound

Overview of the Amphion Toolkit (v0.2) cites this paper.

Overview of the Amphion Toolkit (v0.2) LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T14:21:56.749133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:21:56.749133Z digest=sha256:ebe959d0e6409e06e29ecfde457f1bba1bbbf7b93fe368d9cc45a93cb3d9bb96

Observation e45a97b2-31a9-47ef-bce1-ddf4d8165b37 · inbound

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her cites this paper.

LUCY: Linguistic Understanding and Control Yielding Early Stage of Her LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T13:34:25.666909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T13:34:25.666909Z digest=sha256:ee9a06ac520d9d434b0dc892ee120289cb12ca913f58c091807193136994359c

Observation 7dc3286d-0748-4cc8-888e-601fe4d49c74 · inbound

Ola: Pushing the Frontiers of Omni-Modal Language Model cites this paper.

Ola: Pushing the Frontiers of Omni-Modal Language Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T22:47:39.105733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T22:47:39.105733Z digest=sha256:100d6fa3cbf9c6ddd1c3d29ff43380d6fffa5a3aad3de5e3bc6f9948040e45da

Observation b5002ec4-157b-4ce9-983b-9e285c585c7e · inbound

SparQLe: Speech Queries to Text Translation Through LLMs cites this paper.

SparQLe: Speech Queries to Text Translation Through LLMs LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T22:05:23.804113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T22:05:23.804113Z digest=sha256:aec56495971eb286bd5e13f81d2482d0f4061b5343c5690ec058f922cfde8910

Observation d54a85f0-f0cb-403c-84be-bf3409720cca · inbound

A Preliminary Exploration with GPT-4o Voice Mode cites this paper.

A Preliminary Exploration with GPT-4o Voice Mode LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T20:02:51.322585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:02:51.322585Z digest=sha256:f987742958e61ecb374ea38466271f7b5a0a1d38340659806f63a9fbebb9c620

Observation 2d3ae2fc-a0b0-4740-8024-10f7dbc5bf8e · inbound

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction cites this paper.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.342720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:1965e833d7e362c9a70241b27d12f42afd2c402ee5d208adc87909744bf48c2e

Observation 400caaab-4ef9-4bcf-b1f2-ce21a04a3262 · inbound

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models cites this paper.

VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:33:47.104090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:33:47.104090Z digest=sha256:41f42e68baa5cf17efb7edba77fe767defca38e2350993ca9a6619b8652641bc

Observation c42a38e6-940d-4402-9758-56a45182388f · inbound

SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation cites this paper.

SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:30:21.844448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:30:21.844448Z digest=sha256:7a88954e5dc876f3c5d6d8d0d5af7abcf2a04675542196b171e09f0c5a421421

Observation 033a91c0-6928-4fc0-91f1-3a1bdd811f0a · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:21:27.234568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:16e16b305153f7a5235f02e7e64537159e798802058df15698ce10238b452723

Observation 2be34161-d75b-4350-bbdb-f542090ae388 · inbound

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information cites this paper.

SAKURA: On the Multi-hop Reasoning of Large Audio-Language Models Based on Speech and Audio Information LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:32.388918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:32.388918Z digest=sha256:adb1a34b9026eff5aca5af1f30f48bcb6d3472a7edc78516f143f50f928ce482

Observation c07ba69b-84a4-4031-a186-ef95e4cb43fe · inbound

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach cites this paper.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:33.966621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:33.966621Z digest=sha256:0946527cb1f7bd93f37ee27585066352170df90488db28458f1d3deec8570fba

Observation 507fb53f-1c52-4fb5-81b3-7edc9f0ec53b · inbound

S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models cites this paper.

S2SBench: A Benchmark for Quantifying Intelligence Degradation in Speech-to-Speech Large Language Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:38:36.938679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:38:36.938679Z digest=sha256:9893f282c89153e10ecf4ddd28d209de739e855b1906175c39eb838f2d022fc6

Observation 180b0dc2-a501-411a-aaf2-b27ae8112244 · inbound

ModRWKV: Transformer Multimodality in Linear Time cites this paper.

ModRWKV: Transformer Multimodality in Linear Time LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:35:34.644120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:35:34.644120Z digest=sha256:3f03511ada46f8fb5c8a4db7abec4d752bc0d9ea5cb902af674dd87982bba77a

Observation dc9e1a5e-b5fa-4f82-8711-d91ee80fa622 · inbound

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models cites this paper.

Audio Jailbreak: An Open Comprehensive Benchmark for Jailbreaking Large Audio-Language Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:01.147173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:23:01.147173Z digest=sha256:52817a401e1b107149903a5601dd6a377d0d55f74e89281c1398016e9fbbe748

Observation 0d13692c-21c8-42a9-bc12-cfe3f818b5c3 · inbound

Speechless: Speech Instruction Training Without Speech for Low Resource Languages cites this paper.

Speechless: Speech Instruction Training Without Speech for Low Resource Languages LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:51:16.069890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:51:16.069890Z digest=sha256:266347eda2b1a9d03be6ac7e96ec03e634589e8b7cc869a1a446c5e6cec6abcd

Observation 9ca4f5b7-4a33-40d1-aa74-c22a6e2656c5 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.644063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.644063Z digest=sha256:325c4db4888a4788bc8f1a9c228e4e75d63e353b714bbbb1cf5be1072258c28e

Observation a6c1b79e-1f1d-42bd-9d60-646fde200bf9 · inbound

MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding cites this paper.

MFA-KWS: Effective Keyword Spotting with Multi-head Frame-asynchronous Decoding LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:16:33.562854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:16:33.562854Z digest=sha256:b8dde2682624677253e340997d12484c52bce4e0c4a4170711ce3469d4bb45db

Observation 132089e7-bcf9-4ff1-ada2-857a733acffe · inbound

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling cites this paper.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.598942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.598942Z digest=sha256:987b71dc2ba587472ca18d0c421c34c815f394153168a019efa2c1a9f722b7f3

Observation d802cbdc-6ea6-4981-8b07-cc34f1544fc9 · inbound

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction cites this paper.

OmniCharacter: Towards Immersive Role-Playing Agents with Seamless Speech-Language Personality Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:30.307140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:01:30.307140Z digest=sha256:2094383f1987bbc841351fe18b6c81bd7be290499ea12da40ed5446808a76469

Observation d3b73640-6507-408d-a905-1fce1cd35252 · inbound

Universal Visuo-Tactile Video Understanding for Embodied Interaction cites this paper.

Universal Visuo-Tactile Video Understanding for Embodied Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:24.344367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:09:24.344367Z digest=sha256:987d8224de461f56c691a320defee34f42595539bfcb443b69a6e3a0e58f2fe8

Observation ce5592c6-903f-44ff-ae2d-861a45d3c6c1 · inbound

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems cites this paper.

Chain-of-Thought Training for Open E2E Spoken Dialogue Systems LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:04:31.966356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:04:31.966356Z digest=sha256:14396dbcc514ae44de6eb0e941b847bb5f8dbd985cf360c14144345ee333aca4

Observation 9277ae40-93a8-402b-a2fb-6f5c118d4c39 · inbound

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction cites this paper.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:35.040733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:35.040733Z digest=sha256:6fe83ff3f920bf335ceeef51179ea122bd5a3a0a9ac20a2dd5834c2cdebda53e

Observation a38ced2f-e6bf-428d-9ab4-bc652ad434f4 · inbound

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion cites this paper.

StarVC: A Unified Auto-Regressive Framework for Joint Text and Speech Generation in Voice Conversion LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:39.960930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:28:39.960930Z digest=sha256:6141c48bf28a2b45a9afb95bab331d139df8695255a5dcb5b8001d52d5596b53

Observation 3a38c437-6768-4ac8-b6b1-defd92aa8332 · inbound

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant cites this paper.

SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:23.008363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:23.008363Z digest=sha256:80e64f7081de2c35097a6da5a909259585692b16c45358367810e6e79e39ee51

Observation 076e0b54-13c8-48fa-aa92-1a3b5b45342a · inbound

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment cites this paper.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:40.924179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:40.924179Z digest=sha256:31bc442aaf894a72dc241096218eb305e713d3943979b71acb02b80fae04ed8c

Observation 1985f09c-790f-448f-ba7c-612b0d17d029 · inbound

Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models cites this paper.

Investigating Vulnerabilities and Defenses Against Audio-Visual Attacks: A Comprehensive Survey Emphasizing Multimodal Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:40.921978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:08:40.921978Z digest=sha256:4a30a19ae9fce5d755c30589c34a6a26e2f6fc2c7d4d798c00c5899d73abfe43

Observation d734e001-cd5d-4eee-9cd4-638a837d3b22 · inbound

KERAG_R: Knowledge-Enhanced Retrieval-Augmented Generation for Recommendation cites this paper.

KERAG_R: Knowledge-Enhanced Retrieval-Augmented Generation for Recommendation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:23:47.250006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:23:47.250006Z digest=sha256:b87d65a2c5415c9c77e474be5584da5c95cea9dae7e7c45dc7a046977f32c1fe

Observation 60989bf5-f0e4-4c5f-ae7d-85c95a1d3c30 · inbound

Unlocking Speech Instruction Data Potential with Query Rewriting cites this paper.

Unlocking Speech Instruction Data Potential with Query Rewriting LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:21:32.233590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:21:32.233590Z digest=sha256:1a590fdd2a1af984234c3eee47b99cc9cc2b442d287d8d39a64a3f0b9c159287

Observation c396d604-debf-4f00-846d-96cf0fba6658 · inbound

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation cites this paper.

AudioJudge: Understanding What Works in Large Audio Model Based Speech Evaluation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:47:55.871612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:47:55.871612Z digest=sha256:92aaa85ad009f82c9edca59f8fae6af5eaf1fb3ceb01c0fe5771c51a50cf87af

Observation f9ecc50f-fee8-4c79-bd65-941214304f01 · inbound

Personalized Socially Assistive Robots With End-to-End Speech-Language Models For Well-Being Support cites this paper.

Personalized Socially Assistive Robots With End-to-End Speech-Language Models For Well-Being Support LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:09:04.402999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:09:04.402999Z digest=sha256:a81d5363d93c7e280b39bb9861a43fef84a63ea621c285bf0fe593a9135f6cb8

Observation cdb8ac00-ddad-47a7-abd9-6780ddf68f6e · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:59:51.023035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:975fddb4b9c64444798d7c1377cb084f8de2b48fe2b95c68443dae35b4b141cd

Observation 275754a4-9a8c-48fc-8896-d67dea81c1d9 · inbound

BoSS: Beyond-Semantic Speech cites this paper.

BoSS: Beyond-Semantic Speech LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:50:04.291383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:50:04.291383Z digest=sha256:575cba21c3d3e71144414bac7d453442ebe5867d37e620b6841e11837e135b05

Observation e8048819-f9d5-4b44-b9d1-b147f6b4654e · inbound

Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data cites this paper.

Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T18:22:44.224384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:22:44.224384Z digest=sha256:28e4edc7ab0bcd4e773399d6866f5615753b163298b53d52cd7d2d77f7721a34

Observation 65f20b7f-5213-4063-be88-bbe4978d02b0 · inbound

SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding cites this paper.

SpecASR: Accelerating LLM-based Automatic Speech Recognition via Speculative Decoding LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:56.043822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:56.043822Z digest=sha256:6b3905afb060728edc2a0712fc4a9b8292cdf7126b6e548661959bf08f60d89b

Observation e9d9c51a-45a3-417b-88a7-66d03301e2a5 · inbound

Beyond Value Functions: Single-Loop Bilevel Optimization under Flatness Conditions cites this paper.

Beyond Value Functions: Single-Loop Bilevel Optimization under Flatness Conditions LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T17:52:28.613808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:52:28.613808Z digest=sha256:1ee01f90b8f726294687775f0d704811bb5116cbb3da1c59a26d673fed1563dc

Observation 60e45a49-bb4c-47c9-8812-24061e498963 · inbound

Dual Information Speech Language Models for Emotional Conversations cites this paper.

Dual Information Speech Language Models for Emotional Conversations LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T21:45:23.907075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:45:23.907075Z digest=sha256:1be9707fbf46dd78f37fcc1bab13a0e5dbba991a908788cfc84fa6c4b24add27

Observation ee928dbd-a44f-4c45-9c7c-8826b58e7890 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-19T00:12:54.167660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-19T00:12:39.834892Z digest=sha256:c26b7e1275177b263d0ee2561844bd2fff08193989dd3d83a8b26cf692fe7353

Observation ddc7832d-be84-4b6a-a021-89c7de4e3344 · inbound

Training-Free Multimodal Large Language Model Orchestration cites this paper.

Training-Free Multimodal Large Language Model Orchestration LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-25T08:05:30.612129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-25T08:02:15.950975Z digest=sha256:fb9c13ff26ccf6dd23602b9a6cb3239e43b06641952800c08d27637445fb147f

Observation ff509164-128d-453b-85bd-ae5fef3a68fa · inbound

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model cites this paper.

LLaSO: A Foundational Framework for Reproducible Research in Large Language and Speech Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T17:56:51.856505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:56:51.856505Z digest=sha256:8d734c57dda663904b6e8401aac9a83e72ea3f6c6feafae9a414e908512e4573

Observation ef60b9df-dd01-4cfb-a286-ad7713ca6c18 · inbound

Enhancing Speech Large Language Models through Reinforced Behavior Alignment cites this paper.

Enhancing Speech Large Language Models through Reinforced Behavior Alignment LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-21T22:24:23.513736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-21T22:23:52.392075Z digest=sha256:39814c9b1dd4594104e9493678aae88733b5b179f5b7b5b97817c99f35d93606

Observation 91a0b4c1-b79f-45db-8bae-2e63819c3e44 · inbound

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations cites this paper.

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:04.355944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:33:04.355944Z digest=sha256:470793ef67650cbf88133fe745d1141c13fb42959cc841f4704f00b92b36ac4b

Observation bca84238-c1e2-4017-a014-0ee3f5b7d255 · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:01:24.361876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:7a9cfda02af87c33b5362a208cbd6d123af659e1adcfb4fac3efb8674619ed8a

Observation 55ce28e0-4813-4353-8197-81b50a8f88de · inbound

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues cites this paper.

AV-EMO-Reasoning: Benchmarking Emotional Reasoning Capabilities in Omni-modal LLMS with Audio-visual Cues LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T11:06:44.086529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:06:44.086529Z digest=sha256:bfbb2fcaac204e31d1dea8fca8603ee35e8f695f1b8cdb9a155c18808cdfad78

Observation 73d20c5c-f39d-4b87-a946-54358ef0a2a3 · inbound

Same Words, Different Judgments: How Preferences Vary Across Modalities cites this paper.

Same Words, Different Judgments: How Preferences Vary Across Modalities LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:36:33.038978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T19:31:51.996480Z digest=sha256:e83e40b2f42c13d7a7a4316323882a4a4888d60b4f601e55743d79fffc594497

Observation 941f052b-a146-45b8-b2c5-7596fdacdd7f · inbound

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision cites this paper.

Language-Aware Distillation for Multilingual Instruction-Following Speech LLMs with ASR-Only Supervision LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T18:42:05.886184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:42:05.886184Z digest=sha256:a968e0167d005f7bfcd1a800daac9802dbd69cc82de93bf9e6166502d81b2b11

Observation 8cf905bd-13d5-4852-88e2-5551808ab626 · inbound

Controllable Accent Normalization via Discrete Diffusion cites this paper.

Controllable Accent Normalization via Discrete Diffusion LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-14T21:25:19.253843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:25:19.253843Z digest=sha256:50564c5513709842bc5df324ae65ec296fd01723674e0a15ccbae247798f43d7

Observation 18bc937b-afd2-45c0-be4a-79329bb0cbe7 · inbound

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model cites this paper.

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:08.488015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-09T15:34:50.848124Z digest=sha256:1ea1d672e63c810e9fc929cfa5bf3e3bfcef65635f16c8093cff13ac14e7d083

Observation e57d5363-c6af-4b86-8681-b2d59a1be6c0 · inbound

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization cites this paper.

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:45:44.079790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-08T18:08:16.051117Z digest=sha256:b3aa9f68670429365fa353e86a68c93f0546ab048210877e5c2a8a5bb9cf7498

Observation be8c5c01-a85f-4883-87da-1846317b008e · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:46:15.262510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-08T11:00:52.196039Z digest=sha256:f334cd5d4891a186f49c20ed690150512d677338b4830924fc755878e7a4275f

Observation 7cb90703-4bc7-4fcc-9550-4a20e3920691 · inbound

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM cites this paper.

Minimizing Modality Gap from the Input Side: Your Speech LLM Can Be a Prosody-Aware Text LLM LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:50:49.611468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T00:49:26.507281Z digest=sha256:573d290fbaaca15620eea457aa7794b863a96b51ebc4e1675ad0d467d79c2e6f

Observation cf1b243e-3b0c-46ae-8eb7-1a1e3436288e · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:55.866632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:cb5c2543577063d491bd31fcb71016879457c79b12ac9789acfc3f65ab202a41

Observation df2398ca-422d-4bb7-b0a5-6b4af5991141 · inbound

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook cites this paper.

A Survey of Large Audio Language Models: Generalization, Trustworthiness, and Outlook LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:39:48.969176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T07:38:23.099479Z digest=sha256:c0625040795913da69ac3bb9839cbcf03d9c4482b55f42a9f152e6c03dd6e4eb

Observation db74770c-bf98-4781-9492-e175e48af0a3 · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:43:55.027484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T02:41:13.583493Z digest=sha256:152ffa29c7deebd0b323455304fdd5719d7e0515fd800bf0fa2414e23678c705

Observation 4838ffaa-416f-4089-8ce4-244f17c814de · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:34:57.361577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:6ca8f3ab4fc566b16513d25333c29928119351eba1361b9a7e01b0c1418774fd

Observation a6d8c1b2-019d-468f-ad1a-814d5b2a6a22 · inbound

A Survey of Audio Reasoning in Multimodal Foundation Models cites this paper.

A Survey of Audio Reasoning in Multimodal Foundation Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T02:09:24.368254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-21T02:08:06.976461Z digest=sha256:e6679106822140be3689569f75dc502078abc7fd50996e8fba315c00c0f47481

Observation a43b7391-d3e0-4f63-9793-c00a05313c7a · inbound

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens cites this paper.

TokTalk: Expressive Real-time Facial Animation from Audio-LLM Tokens LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:25:59.872963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T22:45:39.443440Z digest=sha256:5a259a5ddcdf2c41bd9c7dfe96c0baf7530b260516fd5e49e27a74b7da1b3b1d

Observation 65b49c1b-dac7-4476-b120-6205597d2b6d · inbound

Audio Interaction Model cites this paper.

Audio Interaction Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T10:46:52.413130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T04:57:05.062465Z digest=sha256:6b496fb21956bdaec1e1cbaf37e152d7a262f92a4f02efd9f415fa3b24e2314f

Observation eb3d44dd-5b67-4797-96c8-e90f5dc44120 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.662681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:c11262b040d2792bcb177a7ea5f57fa259bf57fa5dd8ae0e6d645e3559a60e46

Observation 084cd552-38f0-4ee4-a5de-dac922514b7d · inbound

Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models cites this paper.

Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:57:44.819997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T11:29:36.012349Z digest=sha256:99bc98d394a5cf11c0fc91da337ab576cb0e3407a77963e1b05b7e10c2518134

Observation a53e1ce9-d567-4f85-9ed0-d0ab6702f9c2 · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:18:12.825126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:ee45d808413fe56ee8c55ae39ef23976d85bccf39b26b032ea627dda949ee158

Observation 6f7c5947-006b-4b80-9cfb-cf4596cd38eb · inbound

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead cites this paper.

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:29:41.323872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T11:42:44.035902Z digest=sha256:6d4e7dd0649f2ebfd369a78e2a4e2ce220a1e879a3b705641440191bdd57cfaf

Observation 2ae65a48-4f96-4a77-8557-222645a0aa62 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 215

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:42.281653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:96eef0dc51c4b4436b54cdf49536134e6de3812c15a761b172de038ba272a5a3

Observation 6199ec6f-8f95-4bcb-a4d8-b464e14d29db · inbound

Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving cites this paper.

Metronome: Bound the Cache, Keep the Beat for Real-Time Interaction Model Serving LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T08:09:44.957032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T08:09:44.957032Z digest=sha256:89747c3dc11ed82681f274319b650263a25b8606fb16e80eeaa92ebdb9ddc60c

Observation a130e201-e309-4fdf-bb17-135aa45b6047 · inbound

TokAN: Accent Normalization Using Self-Supervised Speech Tokens cites this paper.

TokAN: Accent Normalization Using Self-Supervised Speech Tokens LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-11T22:58:27.449735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T22:58:27.449735Z digest=sha256:912c71154f6c72af340b0ec8a7bff0ef1273aba829613dfed667164efb92293e

Observation 788183f2-e31e-4971-a137-16cb22e0c543 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:13.127289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:13.127289Z digest=sha256:c3542ba92c31b09ae083fe86b0a50bf2856ccd1925aece65d0abe7e6b88a4a56

Observation 5e6763b9-2544-45e1-82ae-e2b4c8cbe348 · inbound

Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO cites this paper.

Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T01:56:11.344932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:56:11.344932Z digest=sha256:99d0c7060fdab3f5d410e4bb4e5c35fb95754ea4e7ab3caaf26fff6c2b829db6