Pith. sign in

Paper Citation Record · LEDGER

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents

As of 7 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2608.01119.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01119 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:30:12.153813Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6ef395f-e37d-4481-a367-8fab8f7e00e2 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Moshi: a speech-text foundation model for real-time dialogue

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T00:30:10.752328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:30:10.752328Z digest=sha256:23e2bab77b79f19e95fcb5311e6291e669bb1eb20e71ed16d6f76b3eb18bb978

Observation 52398a3e-4123-4ff3-866e-7a4699707149 · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T00:30:10.821180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:30:10.821180Z digest=sha256:563d2fc59b33e510c41ff4173be5ac4488aa82f1e9b6050bf1ea62f23bc6410f

Observation ab138103-f61e-474a-9fdd-fbe47ffc35e4 · outbound

This paper cites Qwen3-Omni Technical Report.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Qwen3-Omni Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T00:30:10.872210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:30:10.872210Z digest=sha256:184002dd30bdd419e43500d7b8114b98dafc987ed1395b1ea8d34d9605cf5e78

Observation d140047f-5374-4a98-8ebc-f37b8d0f1bd0 · outbound

This paper cites Full-duplex-bench v1.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Full-duplex-bench v1

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:30:14.664834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:30:10.912943Z digest=sha256:4b198fa3dba748b153efc3190ad52f01bd56573540424efae441631b3eb9f546

Observation 9d73f340-dfdf-4e12-92a4-117842d1e7be · outbound

This paper cites Joyai-llm flash: Advancing mid-scale llms with token efficiency, 2026.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Joyai-llm flash: Advancing mid-scale llms with token efficiency, 2026

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:30:14.493409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:30:10.966840Z digest=sha256:03a4ea44a489eff80ac39193e923e0034492a6859f89910df79f6037b23454f2

Observation 541ade7c-7b87-4e23-9450-07008f4c1238 · outbound

This paper cites Deepseek-v3 technical report, 2025.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Deepseek-v3 technical report, 2025

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:30:14.284823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:30:11.076411Z digest=sha256:5c10f465ac547b127d430bdfcb127908270bb59a2e3e7eb1f8efe9ee3695d535

Observation a908f1e1-28da-474a-9c9a-1f39c1edf920 · outbound

This paper cites Charles, Cheng Chen, Guanduo Chen, Haiting Chen, Huarong Chen, Jiahao Chen, et al.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Charles, Cheng Chen, Guanduo Chen, Haiting Chen, Huarong Chen, Jiahao Chen, et al

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:30:14.061828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:30:11.169630Z digest=sha256:e154c6cfcd6f864459c130177ecf7ac3c1f9f311cc8136243e034e79c7ef296a

Observation 61cfc91b-56a6-4d4c-9563-6084e2b3e3df · outbound

This paper cites Joyvoice: Long-context conditioning for anthropomorphic multi-speaker conversational synthesis.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Joyvoice: Long-context conditioning for anthropomorphic multi-speaker conversational synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T00:30:11.238953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:30:11.238953Z digest=sha256:e3c6bc434029e2423b870e7a69837150cc153c4fc93c3c2fa7a0abb55df9ef22

Observation 628f88d7-b279-4863-a476-f19a392e2d9c · outbound

This paper cites Interaction models: A scalable approach to human-ai collaboration.Thinking Machines Lab: Connectionism, May 2026.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Interaction models: A scalable approach to human-ai collaboration.Thinking Machines Lab: Connectionism, May 2026

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:30:13.800415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:30:11.311537Z digest=sha256:ab1be7d9d31eb39eff1e47fc35b9940d73cde75979921b8bd7dc3977f24d2962

Observation 8c02733f-b0e2-4e28-a769-5559d90237ed · outbound

This paper cites an unresolved cited work.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:30:13.581387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:30:11.406010Z digest=sha256:7306c46c11380ebd1f5e46e1f9f46506e6deabec6e95fbb6bbad5f6401a30bae

Observation 35e76cc7-99cc-4623-a452-fe026c25a7ed · outbound

This paper cites an unresolved cited work.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:30:13.373000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:30:11.486175Z digest=sha256:40d351f80fd153f4ac224e4942f6941794f25c79a0d2451263d758449b6af114

Observation dfae0f97-7e32-42ba-98c4-38bc3434682b · outbound

This paper cites Uro-bench: Towards comprehensive evaluation for end-to-end spoken dialogue models.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Uro-bench: Towards comprehensive evaluation for end-to-end spoken dialogue models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:30:13.183293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:30:11.575149Z digest=sha256:b7641af77837665cfc1b105a9877afbfc41cfb3d2482e52c15a5c814e53dee60

Observation 84a57bcc-1d3b-4aff-a185-7803ea456070 · outbound

This paper cites Air-bench: Benchmarking large audio-language models via generative comprehension.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Air-bench: Benchmarking large audio-language models via generative comprehension

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:30:12.972047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:30:11.677624Z digest=sha256:faf8e59e3dce08804a03636706d06c27f06b086cc9a833ccee4f3586ba5229a3

Observation fe411018-4dad-4230-b1e6-f711817657f7 · outbound

This paper cites Schuller, and Jianhua Tao.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Schuller, and Jianhua Tao

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:30:12.766278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T00:30:11.766423Z digest=sha256:5c1b3b69fd84656e5caff502416a7aa2e4eb6c0d883be47b2f6cece3c11d6451

Observation 2ba383fd-2bbc-47a4-a865-eabd4edfa3e4 · outbound

This paper cites Echomind: An interrelated multi-level benchmark for evaluating empathetic speech language models.CoRR, abs/2510.22758, 2025.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Echomind: An interrelated multi-level benchmark for evaluating empathetic speech language models.CoRR, abs/2510.22758, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:30:11.852258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:30:11.852258Z digest=sha256:85276af19a0866e36141a91ca380e3590a328ff09a407d8830f5f299b5e08b6d

Observation 4f8514ff-761c-414d-bbf3-1b6782f4248e · outbound

This paper cites Easy turn: Integrating acoustic and linguistic modalities for robust turn-taking in full-duplex spoken dialogue systems.CoRR, abs/2509.23938, 2025.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Easy turn: Integrating acoustic and linguistic modalities for robust turn-taking in full-duplex spoken dialogue systems.CoRR, abs/2509.23938, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T00:30:11.954747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:30:11.954747Z digest=sha256:e76f3c86d9e7f3d7602063e908e4722d2a0948c2e59757efd4c5eb133e3ca4be

Observation 229d3229-7ef9-48c5-8938-fa6f53bfc3e2 · outbound

This paper cites Soulx-duplug: Plug-and-play streaming state prediction module for realtime full-duplex speech conversation.arXiv preprint arXiv:2603.14877, 2026.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Soulx-duplug: Plug-and-play streaming state prediction module for realtime full-duplex speech conversation.arXiv preprint arXiv:2603.14877, 2026

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T00:30:12.006697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:30:12.006697Z digest=sha256:9ef4f71c131d161443de6ed7f96668f2194bd98a89a982b044ebad4346254d8c

Observation af842580-0665-431a-9382-1f366bd5c6ce · outbound

This paper cites FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents FastTurn: Unifying Acoustic and Streaming Semantic Cues for Low-Latency and Robust Turn Detection

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T00:30:12.085175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:30:12.085175Z digest=sha256:fd1322adb10c36d823f6ef3ba40e034fcbd86c0e8a6f5477782b4c4466ad6233

Observation 77d40865-6b3f-4ab8-9f6e-45ff83bf3fcc · outbound

This paper cites Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition.

JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents Unified Streaming and Non-streaming Two-pass End-to-end Model for Speech Recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T00:30:12.153813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:30:12.153813Z digest=sha256:304d3324fd1e2da0732cd3b3564236074521f2bf7b59617fda814247b9417a2c

Pith citing papers

No inbound Pith citation observations are available.