Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs

As of 6 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2607.06540.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06540 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T02:38:31.073805Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact30
  • verified fuzzy16
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8073a5d3-d128-478a-b0dd-314009358cc1 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Qwen2.5-Omni Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.619878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:fcf1ce4d6f98c82f3425ff1fc10ef0544cb994e732e0364a9c3ffc126f47abea

Observation c464032b-e06d-4e02-a8be-4caf6b8ddb17 · outbound

This paper cites Step-Audio 2 Technical Report.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Step-Audio 2 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.622641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:75f4536a420d29ff0deaffb6c9629be092e98cc4101f01109be433d9e75c722d

Observation 48bbfcc1-b184-4205-bb7a-5e9d6afab943 · outbound

This paper cites 2024 , address=.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs 2024 , address=

Reference 3

Resolution
metadata mismatch
doi, observed 2026-07-08T02:44:27.593264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:40ee6088abb65897bc7428231f82cab1892e3187683d0c475b37a8f963173ea1

Observation 2f8696ee-4a67-4815-bcff-afd98723668d · outbound

This paper cites 2024 , url=.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs 2024 , url=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.077768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:84d54bff003a719c8cc37934f3d826c532b8207adbbef6ad7263b4e62fcaa9dd

Observation 71aa6349-892f-47e6-90b7-d88dbe56116e · outbound

This paper cites Forty-second International Conference on Machine Learning,.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Forty-second International Conference on Machine Learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.100722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:017137dafba4194504c5767cbcb6f3dc87c9020bc670709a18f55e30f2beedd0

Observation c86075f5-e121-40e1-a278-6fd8a71ccb56 · outbound

This paper cites Uni-moe-2.0-omni: Scaling language- centric omnimodal large model with advanced moe, training and data.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Uni-moe-2.0-omni: Scaling language- centric omnimodal large model with advanced moe, training and data

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:44:28.048237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:a41cc008dc3770b187f9002f644a6f9635afb453275f0c1424dc8edf56125f14

Observation 9729d942-8013-4e75-8fe9-66e95fc67465 · outbound

This paper cites Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T02:44:28.050948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:da11ae7c5993713ed96381a6e1e1c2c12db472e95981f60566b7f4da0d0b6430

Observation 19139c8e-60ea-44c7-af10-fbc94cf36280 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Moshi: a speech-text foundation model for real-time dialogue

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T02:44:27.603525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:369c119ddddc8b6263271bff9c5a6caa40b1d8ebc855100a537faace5ff30f5b

Observation fd1aa9dd-4430-4e02-9d67-b0bf7431296d · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.588649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:19c79527b267fa49e1de6ebaa145c795d58e2a053a8a19774a825d2a1ebfd5c9

Observation 1cda6be4-a74d-4dc3-a973-d78313f6864a · outbound

This paper cites Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models , booktitle =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models , booktitle =

Reference 10

Resolution
verified exact
doi, observed 2026-07-08T02:44:27.558902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:02b40448f9a1adc95c7e8dc15ea5741eb6972cd78a43da5f8cc97d04c859c7fd

Observation 3de0d2c6-4aa5-4f65-b43e-e196dba0f7de · outbound

This paper cites OmniFlatten: An End-to-end.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs OmniFlatten: An End-to-end

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.089646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:fed4b5b83fbe99b1c380870ad7f0c60144e1d6a4b2e60614fdb96628f5499493

Observation 8dcc24d8-05fc-445c-a6d6-318d25da1e6c · outbound

This paper cites Peloquin and Bokai Yu and Hongyu Gong and Shyamnath Gollakota , editor =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Peloquin and Bokai Yu and Hongyu Gong and Shyamnath Gollakota , editor =

Reference 12

Resolution
verified exact
doi, observed 2026-07-08T02:44:27.611459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:25b2668f7558829d9264016472c0853e31321d24b93fcba64a918352bbb03cab

Observation f3b0cc80-33c5-45ed-bc8e-d462660701fa · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Lost in the Middle: How Language Models Use Long Contexts

Reference 13

Resolution
malformed identifier
doi_truncated, observed 2026-07-08T02:44:27.607375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:855c4d5cbf9d654d0f2b60e51a191dad0232e94df28a22508fdb6eaf698ee14a

Observation 7692298f-d65c-4698-8233-d1eee4103afb · outbound

This paper cites Flm-audio: Natural monologues improves native full-duplex chatbots via dual training.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Flm-audio: Natural monologues improves native full-duplex chatbots via dual training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:44:27.594063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:624efed14966599ec8da01fff4c6e060926135ef02ba0e8e0deea49404de9680

Observation a16fb13d-a686-47bf-abd6-aa0345df9fa1 · outbound

This paper cites FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.639858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:a5c5655bc90d034f26b37f26e4e6f6e5483660958a12515e0385d89e9b148314

Observation 6b85cadf-c3c0-4c47-83fd-d67304d4995a · outbound

This paper cites FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.581178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:0edd5a0dc2c4f96fa2e51e4196e232a538cb034cce61be298283f87a97ea806e

Observation f0ad6e29-4569-4f31-b8f9-9d6330632f51 · outbound

This paper cites Easy turn: Integrating acoustic and linguistic modalities for robust turn-taking in full-duplex spoken dialogue systems.CoRR, abs/2509.23938.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Easy turn: Integrating acoustic and linguistic modalities for robust turn-taking in full-duplex spoken dialogue systems.CoRR, abs/2509.23938

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:44:27.625563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:7a92bfa99bf627af4f4bd70c913e470283325e4c48877f468ce5c1a1a633535a

Observation bb850aba-42a8-432d-9624-3ba7af116bdb · outbound

This paper cites Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities

Reference 18

Resolution
verified exact
doi, observed 2026-07-08T02:44:27.609250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:5b310e02deaed5c019f1d07ad1b472813665ace2789221cdba3ca2bfcf1ba612

Observation 4ec6d432-c6c4-4a2f-8371-029317a04ab3 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs The Thirteenth International Conference on Learning Representations,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.084664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:571a4f340cc86253817049a96272bf2825dc99bc35cf569de040ce29ee710163

Observation 98702d34-2181-4dfb-9741-b879160a4d84 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs The Thirteenth International Conference on Learning Representations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.100942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:076d5e033821e6fc9811c34edb69053e517cd0bcd0a5ff570eb0a40180a500a4

Observation df695be2-4cf6-4ea1-8e23-298e240f1ed5 · outbound

This paper cites 2025 , eprint=.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs 2025 , eprint=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.094402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:2997d2d1528a42981ba9c07976740fd7c81284204f54d7771bda5c6603088848

Observation 2c7b9d4c-a8a2-4f96-839c-0ae026565842 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.628391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:cb4d8d2edbe61b04af91d321f8ecc1d3ab7e755288eaa5c6f6d0ecfa78f7f00e

Observation 6be500fc-2793-402f-975c-3daafed552c8 · outbound

This paper cites 7th International Conference on Learning Representations.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs 7th International Conference on Learning Representations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.096341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:1cc5875e80fd96f33211731a7873b827184a50c7a3cb286538bb097a1b94e66d

Observation 6470b60b-ec4d-4ee5-b57f-6e176d9dbf5a · outbound

This paper cites Findings of the Association for Computational Linguistics:.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Findings of the Association for Computational Linguistics:

Reference 24

Resolution
verified exact
doi, observed 2026-07-08T02:44:27.572967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:2435afc40107c7d6e90accca55282a6a82d4ee642b41f920a29fc54e18e50b7c

Observation a9882f2e-28c4-475b-82ab-4e6b4265cb3a · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.587484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:3d6737d56563a48d1bc05330265c62ee7734e1c127ee220917a0fa10e1641d0f

Observation 3d380d32-e552-4c0e-9c67-8297676ce432 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision , booktitle =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Robust Speech Recognition via Large-Scale Weak Supervision , booktitle =

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.093626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:a7435fb6202bbd2250008e3f5076e70a7256019afaaff4f3d19f75643fd1c3fc

Observation 091daa65-d903-4158-8aac-c1421ad65188 · outbound

This paper cites FD-Bench:.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs FD-Bench:

Reference 27

Resolution
verified exact
doi, observed 2026-07-08T02:44:27.598548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:bf976fc933719f87f679e9ac4a6cf9dc2251e9b6916fea9dd9f334bac7e7ed71

Observation 4c58e7d2-efcf-4901-a686-426ef16586fb · outbound

This paper cites Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.586165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:0004b89c600b7dfd40970622109d9068e7ffd6b7f1afd8bf739b682e5a79630e

Observation 197231e1-4ef1-416b-b0e8-2ef3841b8fee · outbound

This paper cites In: Interspeech 2022.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs In: Interspeech 2022

Reference 29

Resolution
metadata mismatch
doi, observed 2026-07-08T02:44:27.613541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:6508c9a09074de2dc569d9e89338b8956da0d35dd19b23c3516670120737ed1e

Observation 72153fe9-cf18-4895-a9fb-96c33f9c8cef · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.616822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:bf95d37570a316067c3561ac728dce2bd9213840fd1fd6144476b083cbbc78e9

Observation e27ccfc7-4414-4250-83e8-d30fb0de5281 · outbound

This paper cites ArXiv , year=.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs ArXiv , year=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.089467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:60b89011d364db79c5ac1872c7e5ee6bc56de19f76fc23f0aaeb63a50f3d42e8

Observation f3da0420-87d6-43a9-b592-8926a5bed2a0 · outbound

This paper cites Moe adapter for large audio language models: Sparsity, disentanglement, and gradient-conflict-free.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Moe adapter for large audio language models: Sparsity, disentanglement, and gradient-conflict-free

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:44:27.604524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:aa04382739ef6e036ab6db525936ff78e9bf4ee6caeb1335baf15c96ad176e3f

Observation b65231ce-691f-49ab-bddc-1b16309664a0 · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs WavChat: A Survey of Spoken Dialogue Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.611257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:af0a5405f46cae1fb71955de26953df02f525887c4d8d381436ba3eb702115f4

Observation 5eecad0f-c386-488b-a254-31dad33f4ff2 · outbound

This paper cites On The Landscape of Spoken Language Models:.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs On The Landscape of Spoken Language Models:

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.098690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:e7c03390a4361dea9e5c9971888db89cf030f856208457c1040d9eebfe297613

Observation 742c147c-3090-44e2-a8ca-453d497e52b9 · outbound

This paper cites Duplex Conversation: Towards Human-like Interaction in Spoken Dialogue Systems , booktitle =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Duplex Conversation: Towards Human-like Interaction in Spoken Dialogue Systems , booktitle =

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:44:27.584536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:c5ddbeb419083afcf04cd52253ab0aa003befda34fad1ca8e136d2d1fc3c68d9

Observation 6c40db74-d2fa-4534-82f6-d7499846db9c · outbound

This paper cites Efficient and direct duplex 14 DuplexSLA modeling for speech-to-speech language model.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Efficient and direct duplex 14 DuplexSLA modeling for speech-to-speech language model

Reference 36

Resolution
verified exact
doi, observed 2026-07-08T02:44:27.595665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:a0f5a3ad619c8d7e743f9e23d77aa0828a7da6aa1ee463933d48aea835e45017

Observation b8384b0e-3f14-40fb-b734-08764e7fb9a1 · outbound

This paper cites LLM integration in extended reality: A comprehensive review of current trends, challenges, and future perspectives.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs LLM integration in extended reality: A comprehensive review of current trends, challenges, and future perspectives

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:44:27.636629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:a5402b970edbe4cfe22e19b151ce4898a9ce42bfe38b3dc3a95918f715b3f0eb

Observation 21f6f1e5-7025-4737-8c84-c0a4aed49abd · outbound

This paper cites A Preliminary Exploration with GPT-4o Voice Mode.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs A Preliminary Exploration with GPT-4o Voice Mode

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.576608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:6f4a33f1f9211af370ecc249783e145c51d9a1181898d6e71910b9a2923666e2

Observation 2138829a-7b48-4d70-b364-2890d40b492a · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Advances in Neural Information Processing Systems , year=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.122930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:9989f47e4059bdc259a4017875e763c2df32f1a1e7ed818be0bbca27ddf70bbd

Observation 45215517-5e2e-4a33-8210-ccebfc14e0f3 · outbound

This paper cites Schegloff and Gail Jefferson , journal =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Schegloff and Gail Jefferson , journal =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.096738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:4f2b1f6b0d28ba26ddb8e15478e025aff5669bf6a9f670a4e134ce5bdd647b38

Observation f3c58321-cb36-4240-af14-d8bdfa193b68 · outbound

This paper cites Scaling Laws for Neural Language Models.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Scaling Laws for Neural Language Models

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T02:44:28.045394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:f0ad45e707bd881583e72de11a2082bfd195f35f4672538312a9bb91d70a6a79

Observation 9e87f6db-149f-461b-9912-f78fde4bd2a2 · outbound

This paper cites and Sifre, Laurent , title =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs and Sifre, Laurent , title =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.085448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:4ae346ba3cea395d5ee24a2b8f28e3377ce808f6490a3c3d03ffadb7d1612926

Observation a2b6b703-6c2c-41d4-8848-295f1e74aada · outbound

This paper cites An Overview of Multi-Task Learning in Deep Neural Networks.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs An Overview of Multi-Task Learning in Deep Neural Networks

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:28.042808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:fd8145287e87e14fd52d6cf332161ae417563b865a9a7cfb42b17779adffc06b

Observation 675b3f63-32f4-4cb4-8d80-7e9f420941f0 · outbound

This paper cites Gradient Surgery for Multi-Task Learning , url =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Gradient Surgery for Multi-Task Learning , url =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.098540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:668c5367813f5c5df4f659ae60c8971fb730354071202478b3f2db6b568ff645

Observation e69da6f3-542d-4795-851c-c07896143261 · outbound

This paper cites Multi-Task Learning as Multi-Objective Optimization , url =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Multi-Task Learning as Multi-Objective Optimization , url =

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.083081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:c4265bbe2e3f85c377be88eaddcba9df58db61e1a42a5592fa9cea07ea925115

Observation 6ee3558d-3725-4947-9bda-7b5a391c0bed · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing doi:10.1109/TASLP.2023.3288409.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs IEEE/ACM Transactions on Audio, Speech, and Language Processing doi:10.1109/TASLP.2023.3288409

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:44:27.561523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:a795e3c4885bacf806d56a64bd5f0b8d004209c675a41990c910f8ce6dda2c11

Observation 26063020-15b3-47ca-b4b9-faff0bc7cab5 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers , year=.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers , year=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.076868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:3214929a9df950c100523a6fea8c90b597498b371afa662ed894cef82175b6ae

Observation c95ab5c7-faf9-4d64-b2b1-ffb71d148c30 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Soundstream: An end-to-end neural audio codec

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:44:27.550416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:089c518a9d10b6fce016ec1a9a65bf81e80bfe3b4037d5fab394dfb43b8e74fa

Observation 28f62edc-ae10-41d3-aa1e-bfadcb7062b9 · outbound

This paper cites Multimodal learning with transformers: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(10):12113–12132.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Multimodal learning with transformers: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(10):12113–12132

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:44:27.556602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:98f0a4213ccc9412c67b5d451e37e2b26e7ecd20fd79da34b32bf7655f9ff365

Observation b79d7231-c488-45b4-9fe9-827beb811674 · outbound

This paper cites Damien Ernst, Pierre Geurts, and Louis Wehenkel.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Damien Ernst, Pierre Geurts, and Louis Wehenkel

Reference 50

Resolution
metadata mismatch
doi, observed 2026-07-08T02:44:27.519071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:20c8aeeb98ce83fad1245b77315e57eb2e2f5fc611cfe6780ddec5a6ccedbfde

Observation 2d06a823-eb47-47d0-b9b3-263e1e0e86b9 · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.583375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:a60b5dbed2c8568fdf319b3e46a2a178d6f2c22415ee958a12d49a74c5729692

Observation 59717bf9-801b-4e78-981b-3511d357c519 · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.591192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:437d963528c643ce09ff17095d9b8916b5d0064d2363dc9e79c3706c8f28c2fd

Observation 64ebfddc-59c9-4709-91aa-d6ae7d164469 · outbound

This paper cites LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.526294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:1fe27356fc8181eb7913c4f259814eacdd3bd1a6d83913ce5193af22716f236d

Observation c2226ca2-91bf-4957-b01f-fb74290ba53e · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.544716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:16e268ebe47b527e27d9ff724f1479365e183234784174c266d91a93b04422d7

Observation 692fd470-4a4a-4250-a939-0a77615a0aac · outbound

This paper cites LUCY: Linguistic Understanding and Control Yielding Early Stage of Her.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs LUCY: Linguistic Understanding and Control Yielding Early Stage of Her

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.522952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:1f375168b46317f66a18a709dfcbbe12faaf1739765819e9ec50b4ab1f26a174

Pith citing papers

No inbound Pith citation observations are available.