Pith. sign in

Paper Citation Record · LEDGER

Language Model Can Listen While Speaking

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2408.02622.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.02622 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:30:21.941410Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T20:41:13.541672Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2737a321-83c1-4b02-be69-0b21d04dbe5e · inbound

WavChat: A Survey of Spoken Dialogue Models cites this paper.

WavChat: A Survey of Spoken Dialogue Models Language Model Can Listen While Speaking

Reference 143

Resolution
unresolved
no resolver link, observed 2026-08-12T20:13:57.746800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:13:57.746800Z digest=sha256:47059b2cfe8888a8ebd9d6e2c77033fc51c7e0b2b2414193cb411aa207df0538

Observation 6036e836-4175-4943-9540-61e1053cce2a · inbound

SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation cites this paper.

SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation Language Model Can Listen While Speaking

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T11:31:16.143025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:31:16.143025Z digest=sha256:5ce2e2a79aa1918f3b35fd8eea829263634fab987de5ab83d81c52645cf04100

Observation c8a474e0-d0b3-410d-8da3-75f94ae18b0f · inbound

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions cites this paper.

InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions Language Model Can Listen While Speaking

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T16:56:06.302551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:56:06.302551Z digest=sha256:c794bf5f807e0eee73a49c20bfafc990785d87477581b51ec9b2fbe6b890d4e5

Observation 7dd56083-4d5d-40e5-8eee-4edecf9f44c6 · inbound

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training cites this paper.

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training Language Model Can Listen While Speaking

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:40.165779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:40.165779Z digest=sha256:6bfd675f46027d0454b65799e99ae700cb381f455166863bf46cbf968c484fbf

Observation cbeae967-33b4-41c2-b95d-50f8d569f215 · inbound

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction cites this paper.

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction Language Model Can Listen While Speaking

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:56.744635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:10:56.744635Z digest=sha256:4d07b6a6d1c7b08b3ebbff0e5ce66af329e8514af52c98f7363b8c215036f777

Observation a8bf02f8-4065-425f-a275-5d422a254839 · inbound

Beyond Turn-taking: Introducing Text-based Overlap into Human-LLM Interactions cites this paper.

Beyond Turn-taking: Introducing Text-based Overlap into Human-LLM Interactions Language Model Can Listen While Speaking

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T00:42:24.786146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T00:42:24.786146Z digest=sha256:e7c76a033646f398f9f85c9f4221d455d5ba8c7792c26f88447449bb57837c16

Observation 2cc726c3-a956-4c4b-a178-b6f894a9ff56 · inbound

SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation cites this paper.

SimulS2S-LLM: Unlocking Simultaneous Inference of Speech LLMs for Speech-to-Speech Translation Language Model Can Listen While Speaking

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T11:30:21.941410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:30:21.941410Z digest=sha256:fd9d71f6573ecfc4053497bfcb7edf4d52aeec922521fc9c6b2f8c7223f0e1ac

Observation be552e6e-2ee5-45fa-9bde-29e1f95b63e5 · inbound

LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis cites this paper.

LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis Language Model Can Listen While Speaking

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:52:07.887315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:52:07.887315Z digest=sha256:871637634208913d870598c063f33771ddfefd427914975300f40df780845363

Observation 6d69e0e4-cd4a-4954-91cd-920b36bfd365 · inbound

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model cites this paper.

SALM-Duplex: Efficient and Direct Duplex Modeling for Speech-to-Speech Language Model Language Model Can Listen While Speaking

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:41.553345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:41.553345Z digest=sha256:f76113f8d5f2377193683532aca45da3f40f0e49b19cd5db33cd36a2057b561b

Observation b210dfb8-1829-466d-9661-3d084b5f0f4f · inbound

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs cites this paper.

LiSTEN: Learning Soft Token Embeddings for Neural Audio LLMs Language Model Can Listen While Speaking

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:33:46.411149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:33:46.411149Z digest=sha256:55c1ddc62d2cbf72def648c60999b0f863702bc7de87e0eb8262bb8195652830

Observation 91447976-ac9f-4a61-9ca5-c61009dce3b5 · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Language Model Can Listen While Speaking

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.581577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.581577Z digest=sha256:a73b58e79b5e0744846990ccb4a10cafd0e6f3258c4a932774aac33fd3dda058

Observation d6e50ae6-0bd2-4ffe-89d5-8c9e59471355 · inbound

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction cites this paper.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Language Model Can Listen While Speaking

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:37.678428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:37.678428Z digest=sha256:3064e80a724911ea6004397b4006665a7dce5d300b18a510b27d064389b2bd99

Observation 55676e28-45e8-4cf7-b2d0-689fce2729b4 · inbound

Towards a Japanese Full-duplex Spoken Dialogue System cites this paper.

Towards a Japanese Full-duplex Spoken Dialogue System Language Model Can Listen While Speaking

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:14:21.543096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:14:21.543096Z digest=sha256:2cff4b8c5dcb56192336b3354a3069d1915fc1cc73ccba43ef92699fd125f7a4

Observation 5e649e6d-6f63-41f8-b5f9-763d810e26b8 · inbound

Towards Emotion Co-regulation with LLM-powered Socially Assistive Robots: Integrating LLM Prompts and Robotic Behaviors to Support Parent-Neurodivergent Child Dyads cites this paper.

Towards Emotion Co-regulation with LLM-powered Socially Assistive Robots: Integrating LLM Prompts and Robotic Behaviors to Support Parent-Neurodivergent Child Dyads Language Model Can Listen While Speaking

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:36:23.645457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:36:23.645457Z digest=sha256:14e0903bf921eb0dff49fb5a8b01e569dbc40b0dc104f353d2e24d66b60e31bb

Observation 842c9ff1-2696-4def-863a-ca925c5bbb3d · inbound

FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems cites this paper.

FD-Bench: A Full-Duplex Benchmarking Pipeline Designed for Full Duplex Spoken Dialogue Systems Language Model Can Listen While Speaking

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:06:21.263232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:06:21.263232Z digest=sha256:ccace39beb28ac6a24ae37837890711f5c8f72712bcbcd533a9e12542e7519fc

Observation 983eb816-01c4-4e2b-8da2-113902adb740 · inbound

Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis cites this paper.

Open-Source Full-Duplex Conversational Datasets for Natural and Interactive Speech Synthesis Language Model Can Listen While Speaking

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T10:25:14.926547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:25:14.926547Z digest=sha256:bd84f39e18a3225641d20d3a0d1db3766ba701374be96ba2e57198ad0e59ec72

Observation f2ccd446-7b5b-4b38-9e07-97e6d96459d7 · inbound

Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations cites this paper.

Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations Language Model Can Listen While Speaking

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:41:13.547085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-08T08:16:47.522489Z digest=sha256:d50dfb82e61cd70c2d6926be4b6dec20d35a59742528c9beaae7282f32882950