Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs

As of 22 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2607.06540.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06540 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T02:38:31.073805Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact30
  • verified fuzzy16
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch8

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8073a5d3-d128-478a-b0dd-314009358cc1 · outbound

This paper cites Qwen2.5-Omni Technical Report.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Qwen2.5-Omni Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.619878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:5d051bdd94944c26d8376b200eaed1068ce8d0c77de5d1493d9bbebab38b6509

Observation c464032b-e06d-4e02-a8be-4caf6b8ddb17 · outbound

This paper cites Step-Audio 2 Technical Report.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Step-Audio 2 Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.622641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:bfd6e4051c649e7843ddda374218f4e83fa697f0d4b2fcb6458fe5d2c860eadb

Observation 48bbfcc1-b184-4205-bb7a-5e9d6afab943 · outbound

This paper cites 2024 , address=.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs 2024 , address=

Reference 3

Resolution
metadata mismatch
doi, observed 2026-07-08T02:44:27.593264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:b53313a61772eb1153d820f225cb2349435e60e10709cc43529d1d65362a14f8

Observation 2f8696ee-4a67-4815-bcff-afd98723668d · outbound

This paper cites 2024 , url=.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs 2024 , url=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.077768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:a59ef4794de2d86e7bcca0d9aa6c81f1c74271981d97098c6aa7f0649afd374d

Observation 71aa6349-892f-47e6-90b7-d88dbe56116e · outbound

This paper cites Forty-second International Conference on Machine Learning,.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Forty-second International Conference on Machine Learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.100722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:e9e693b8c9cbb35c101083dac39a97bb24e9494880108eaa6e8b88f7522954d3

Observation c86075f5-e121-40e1-a278-6fd8a71ccb56 · outbound

This paper cites Uni-moe-2.0-omni: Scaling language- centric omnimodal large model with advanced moe, training and data.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Uni-moe-2.0-omni: Scaling language- centric omnimodal large model with advanced moe, training and data

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:44:28.048237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:edd651f295024676e31cf3e277633f9de07309e31694a5e73e4b29a0052c404a

Observation 9729d942-8013-4e75-8fe9-66e95fc67465 · outbound

This paper cites Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Perception, Reason, Think, and Plan: A Survey on Large Multimodal Reasoning Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T02:44:28.050948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:f4315f67164b728562e50b72bc9102c496c6f84e9c049cab259c06756976d214

Observation 19139c8e-60ea-44c7-af10-fbc94cf36280 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Moshi: a speech-text foundation model for real-time dialogue

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T02:44:27.603525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:8801b961f5f5ef92722419a18e411ef0ff809b2061b5e675c55f52631e123911

Observation fd1aa9dd-4430-4e02-9d67-b0bf7431296d · outbound

This paper cites VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.588649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:7bbf87c5a9dae7e55bc0bd15b74d539eb66defabc2d77979f2c994f16da43e95

Observation 1cda6be4-a74d-4dc3-a973-d78313f6864a · outbound

This paper cites Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models , booktitle =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Beyond the Turn-Based Game: Enabling Real-Time Conversations with Duplex Models , booktitle =

Reference 10

Resolution
verified exact
doi, observed 2026-07-08T02:44:27.558902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:9e127927a319a8ac53e51e8ca275c4665de8caf3c0cce00e863fe40fb70fafba

Observation 3de0d2c6-4aa5-4f65-b43e-e196dba0f7de · outbound

This paper cites OmniFlatten: An End-to-end.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs OmniFlatten: An End-to-end

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.089646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:89d2fd4fd5b7c10ba7a3eb73ab7e573f0c9a0b5e897c41f72c32845fd2df20fe

Observation 8dcc24d8-05fc-445c-a6d6-318d25da1e6c · outbound

This paper cites Peloquin and Bokai Yu and Hongyu Gong and Shyamnath Gollakota , editor =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Peloquin and Bokai Yu and Hongyu Gong and Shyamnath Gollakota , editor =

Reference 12

Resolution
verified exact
doi, observed 2026-07-08T02:44:27.611459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:8229252ab8e9fa88c0e6abceda603d6b83693c6d068c99ea6adeab7f375032a9

Observation f3b0cc80-33c5-45ed-bc8e-d462660701fa · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Lost in the Middle: How Language Models Use Long Contexts

Reference 13

Resolution
malformed identifier
doi_truncated, observed 2026-07-08T02:44:27.607375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:7e73ba7b15ed10b8e8a43306301dfa0a4996597fd566859df55c0c87e70b50be

Observation 7692298f-d65c-4698-8233-d1eee4103afb · outbound

This paper cites Flm-audio: Natural monologues improves native full-duplex chatbots via dual training.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Flm-audio: Natural monologues improves native full-duplex chatbots via dual training

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:44:27.594063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:d212cf5c5683d69cd3a574ac5063df2a69b1b733763d40a5bfcb1d393f04dc16

Observation a16fb13d-a686-47bf-abd6-aa0345df9fa1 · outbound

This paper cites FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.639858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:756a801fc0982ac04783874561f57c34fc7e4f5a3a079c7ef4bc13c1868b3fe4

Observation 6b85cadf-c3c0-4c47-83fd-d67304d4995a · outbound

This paper cites FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs FlexDuo: A Pluggable System for Enabling Full-Duplex Capabilities in Speech Dialogue Systems

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.581178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:2e9eba71cf738a13239af87b07892a141e91e6911cf2b1f5e848d2236f199548

Observation f0ad6e29-4569-4f31-b8f9-9d6330632f51 · outbound

This paper cites Easy turn: Integrating acoustic and linguistic modalities for robust turn-taking in full-duplex spoken dialogue systems.CoRR, abs/2509.23938.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Easy turn: Integrating acoustic and linguistic modalities for robust turn-taking in full-duplex spoken dialogue systems.CoRR, abs/2509.23938

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:44:27.625563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:cf292f460d4fef2e6d134af79f28ea636613654b785d326822e6c8eac8c5daf2

Observation bb850aba-42a8-432d-9624-3ba7af116bdb · outbound

This paper cites Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Speechgpt: Empowering large language models with intrinsic cross-modal conversational abilities

Reference 18

Resolution
verified exact
doi, observed 2026-07-08T02:44:27.609250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:ed9120212f28855eacd3eb744ca2a0695f8e44d182cb0b08b7f431ecc46c7eba

Observation 4ec6d432-c6c4-4a2f-8371-029317a04ab3 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs The Thirteenth International Conference on Learning Representations,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.084664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:38167ed6b259b52dd0127420e59e995ae25bfac9a3b37fc50b3d975083004acd

Observation 98702d34-2181-4dfb-9741-b879160a4d84 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs The Thirteenth International Conference on Learning Representations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.100942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:815ccd8b54bb67ec0bb335139f538c422fbf4e73e6f1abca34a0042b8c433bb1

Observation df695be2-4cf6-4ea1-8e23-298e240f1ed5 · outbound

This paper cites 2025 , eprint=.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs 2025 , eprint=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.094402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:1132d477e3d885e324e9d48a288f27112ec66f9104d6466dd24bbd0c40c1f1ab

Observation 2c7b9d4c-a8a2-4f96-839c-0ae026565842 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.628391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:d4994445e91fdb9f52b65def19c44c9b709c97730d89793d40ac54dbe09e982b

Observation 6be500fc-2793-402f-975c-3daafed552c8 · outbound

This paper cites 7th International Conference on Learning Representations.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs 7th International Conference on Learning Representations

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.096341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:26dba82ec65d85e079517a15ddc5f85ba8b18bd715421e0b5559c71ab80ee63f

Observation 6470b60b-ec4d-4ee5-b57f-6e176d9dbf5a · outbound

This paper cites Findings of the Association for Computational Linguistics:.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Findings of the Association for Computational Linguistics:

Reference 24

Resolution
verified exact
doi, observed 2026-07-08T02:44:27.572967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:e20ce56752af1493f91171b065a13d80cf66444710a45b845cf840f17fb511b1

Observation a9882f2e-28c4-475b-82ab-4e6b4265cb3a · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.587484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:31cba7cafae0615fe418ec139f9c2f6d6f09435db272eb52d07399929d3fc012

Observation 3d380d32-e552-4c0e-9c67-8297676ce432 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision , booktitle =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Robust Speech Recognition via Large-Scale Weak Supervision , booktitle =

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.093626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:394d1b17f69f40d48defa22b3b4cf8af39272494603ac82597b2b4f12b4dab22

Observation 091daa65-d903-4158-8aac-c1421ad65188 · outbound

This paper cites FD-Bench:.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs FD-Bench:

Reference 27

Resolution
verified exact
doi, observed 2026-07-08T02:44:27.598548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:ecabec3613ddee8fe2ff4f3132a1a5532b6f9170d737fcff86453d2fab1f47e5

Observation 4c58e7d2-efcf-4901-a686-426ef16586fb · outbound

This paper cites Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Full-Duplex-Bench: A Benchmark to Evaluate Full-duplex Spoken Dialogue Models on Turn-taking Capabilities

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.586165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:bffd2e38e64e2c3e9dfc8e0ce0c53b85f8c36918725b0c5db78b9d4563c1f6b5

Observation 197231e1-4ef1-416b-b0e8-2ef3841b8fee · outbound

This paper cites In: Interspeech 2022.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs In: Interspeech 2022

Reference 29

Resolution
metadata mismatch
doi, observed 2026-07-08T02:44:27.613541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:0262ac92f02070ddf0c12fa8dede70ccc37c4dc4da8ba2e823e7fa9aa0dbf64e

Observation 72153fe9-cf18-4895-a9fb-96c33f9c8cef · outbound

This paper cites FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.616822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:415565bf4113695805b950004f7115afb6d5884f0aa28585e009c9cc9cf2a125

Observation e27ccfc7-4414-4250-83e8-d30fb0de5281 · outbound

This paper cites ArXiv , year=.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs ArXiv , year=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.089467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:ed32f382c80acccc8496d9b4f60fb9074ae9588eea77e71b85ce343ef37cf988

Observation f3da0420-87d6-43a9-b592-8926a5bed2a0 · outbound

This paper cites Moe adapter for large audio language models: Sparsity, disentanglement, and gradient-conflict-free.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Moe adapter for large audio language models: Sparsity, disentanglement, and gradient-conflict-free

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:44:27.604524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:91c189a9b690886f7cbc09a8e6212d5749a3752d0ad3c22cc34d243167e042f7

Observation b65231ce-691f-49ab-bddc-1b16309664a0 · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs WavChat: A Survey of Spoken Dialogue Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.611257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:8dc4efe789af660b26804d6a3d659a3caf45d4a651904eaa6fffed9955de7068

Observation 5eecad0f-c386-488b-a254-31dad33f4ff2 · outbound

This paper cites On The Landscape of Spoken Language Models:.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs On The Landscape of Spoken Language Models:

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.098690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:545414be4d933c7103151b8858a8e73d91e97599768827f75b808d04cb922344

Observation 742c147c-3090-44e2-a8ca-453d497e52b9 · outbound

This paper cites Duplex Conversation: Towards Human-like Interaction in Spoken Dialogue Systems , booktitle =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Duplex Conversation: Towards Human-like Interaction in Spoken Dialogue Systems , booktitle =

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:44:27.584536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:8163a7cb8801a63dbc96fb41b5cb8564f0c5743fe318a96bbd72843b3b01955e

Observation 6c40db74-d2fa-4534-82f6-d7499846db9c · outbound

This paper cites Efficient and direct duplex 14 DuplexSLA modeling for speech-to-speech language model.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Efficient and direct duplex 14 DuplexSLA modeling for speech-to-speech language model

Reference 36

Resolution
verified exact
doi, observed 2026-07-08T02:44:27.595665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:a78924b50bd1431e55ce12e0b3f1a1028a1d26111336a053cd8152e770276915

Observation b8384b0e-3f14-40fb-b734-08764e7fb9a1 · outbound

This paper cites LLM integration in extended reality: A comprehensive review of current trends, challenges, and future perspectives.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs LLM integration in extended reality: A comprehensive review of current trends, challenges, and future perspectives

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T02:44:27.636629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:840cec2bf6f9ebefdfba616656065acf10946d1db864188311cb09c2b3c2e8d9

Observation 21f6f1e5-7025-4737-8c84-c0a4aed49abd · outbound

This paper cites A Preliminary Exploration with GPT-4o Voice Mode.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs A Preliminary Exploration with GPT-4o Voice Mode

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.576608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:18689817a998f96b7d396a864d6c6c4bb0370ed79498ca569f5468bda62a05af

Observation 2138829a-7b48-4d70-b364-2890d40b492a · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Advances in Neural Information Processing Systems , year=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.122930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:d0ef6c14c7bf82c74a6a58171536e42e7ead6d90dec1f87451a835700bf1e149

Observation 45215517-5e2e-4a33-8210-ccebfc14e0f3 · outbound

This paper cites Schegloff and Gail Jefferson , journal =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Schegloff and Gail Jefferson , journal =

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.096738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:51dd1a2da96c21cd81620be947fe711fc62c6c78ea8e0375cf0104a8efd0bdf8

Observation f3c58321-cb36-4240-af14-d8bdfa193b68 · outbound

This paper cites Scaling Laws for Neural Language Models.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Scaling Laws for Neural Language Models

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T02:44:28.045394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:9b2f4a70f864ab24acad519524e521dd217efc471e34d667963cf571d01598e3

Observation 9e87f6db-149f-461b-9912-f78fde4bd2a2 · outbound

This paper cites and Sifre, Laurent , title =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs and Sifre, Laurent , title =

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.085448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:67e1570deb6ef930877310d3425aff580572e91598a27c8a89607647896f2000

Observation a2b6b703-6c2c-41d4-8848-295f1e74aada · outbound

This paper cites An Overview of Multi-Task Learning in Deep Neural Networks.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs An Overview of Multi-Task Learning in Deep Neural Networks

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:28.042808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:eef83b0b85810ef885b97e1f6e03af0d9c09144c531928aff4cd719e5a4203d7

Observation 675b3f63-32f4-4cb4-8d80-7e9f420941f0 · outbound

This paper cites Gradient Surgery for Multi-Task Learning , url =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Gradient Surgery for Multi-Task Learning , url =

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.098540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:e508dc624013a1271821140db2bdd1a3d7d8ba74a584fd15e2e1bb5a32740537

Observation e69da6f3-542d-4795-851c-c07896143261 · outbound

This paper cites Multi-Task Learning as Multi-Objective Optimization , url =.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Multi-Task Learning as Multi-Objective Optimization , url =

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.083081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:34008ec5e3c311d5174b9c621a37cc84a5a891a98efa1273a2ea95da339e13e4

Observation 6ee3558d-3725-4947-9bda-7b5a391c0bed · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing doi:10.1109/TASLP.2023.3288409.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs IEEE/ACM Transactions on Audio, Speech, and Language Processing doi:10.1109/TASLP.2023.3288409

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:44:27.561523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:0c7c5fbabf9aff98c08afaca815291757e9df5ea034df076d1abb379a7f6d90d

Observation 26063020-15b3-47ca-b4b9-faff0bc7cab5 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers , year=.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers , year=

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T02:54:28.076868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:332e0b5315989fa12a7bfa5b0997baa6af48b5b35bd920434d9672d8f0349fc3

Observation c95ab5c7-faf9-4d64-b2b1-ffb71d148c30 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Soundstream: An end-to-end neural audio codec

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:44:27.550416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:f65b7d820ec8c5deea28da28b16909cb69c727b47b1d182b6a0f5e9622b4fe36

Observation 28f62edc-ae10-41d3-aa1e-bfadcb7062b9 · outbound

This paper cites Multimodal learning with transformers: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(10):12113–12132.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Multimodal learning with transformers: A survey.IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(10):12113–12132

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:44:27.556602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:19fbbcdb4c5220aac73f449ce019d0188616ee7a274a375396a0788ccba76745

Observation b79d7231-c488-45b4-9fe9-827beb811674 · outbound

This paper cites Damien Ernst, Pierre Geurts, and Louis Wehenkel.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Damien Ernst, Pierre Geurts, and Louis Wehenkel

Reference 50

Resolution
metadata mismatch
doi, observed 2026-07-08T02:44:27.519071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:dedc80d204e8af237f39646f4516364ac1f67c9c8e5f3fe53148f39e09bae81c

Observation 2d06a823-eb47-47d0-b9b3-263e1e0e86b9 · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.583375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:58fae0f20aef83d12673b316c6af001a8832e1ce4d5dbf710462c58cc2ad2118

Observation 59717bf9-801b-4e78-981b-3511d357c519 · outbound

This paper cites Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.591192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:89018b8414f05c87a16869e370e71c2f73f7450fb6667de78092b0dc4a9debb4

Observation 64ebfddc-59c9-4709-91aa-d6ae7d164469 · outbound

This paper cites LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.526294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:b0f6c60d835b13bff5dda0a23a79ea928cef95c92d41c4f7508045a1e3f0e55a

Observation c2226ca2-91bf-4957-b01f-fb74290ba53e · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.544716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:448699d018f965116f78e733390d974918fcc81f0e7989fa6481f4e52fa6af8f

Observation 692fd470-4a4a-4250-a939-0a77615a0aac · outbound

This paper cites LUCY: Linguistic Understanding and Control Yielding Early Stage of Her.

Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs LUCY: Linguistic Understanding and Control Yielding Early Stage of Her

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:44:27.522952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-08T02:38:31.073805Z digest=sha256:60dd1b929046fe4178686249ddc74849d49a837401cf0c50b84adfe3341be1c1

Pith citing papers

No inbound Pith citation observations are available.