Pith. sign in

Paper Citation Record · LEDGER

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction

As of 13 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2606.09186.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.09186 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T15:22:02.107863Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:37:24.214630Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact8
  • verified fuzzy0
  • unresolved4
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch16

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7de6918a-8e30-45ad-bd8b-ba362a98e847 · outbound

This paper cites FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.977710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:aef94c1207bb00a1b6c096e3feb0f575935b18a7c653878843a32164d2ab045f

Observation cd1e3c9b-e5eb-41aa-9191-8d12d5f9c043 · outbound

This paper cites Qwen2-Audio Technical Report.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Qwen2-Audio Technical Report

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.967180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:4832de3551f6d4588478432ae14afb8780caf43b830c974490f1a70c8b9e937a

Observation 6c6575e6-57dd-432a-bd55-12a5ff1fc009 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Moshi: a speech-text foundation model for real-time dialogue

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.974880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:2be4bd06d50a6f5ace1ccf17d486c269195606a5c4f27f1c505d288603ba810f

Observation 616e258a-1439-478f-b0b9-83aa4e3fe9a0 · outbound

This paper cites Enhancing Chat Language Models by Scaling High-quality Instructional Conversations.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Enhancing Chat Language Models by Scaling High-quality Instructional Conversations

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.990361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:57ed51c31d58da65a27829b7b1318f5c12308575cfeeb58fda4e7923b5028fb5

Observation 98fb4e5e-9dad-49dd-88fa-f8daedc1774a · outbound

This paper cites an unresolved cited work.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T15:22:02.107863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:e49f74cc499d726077f35314f74240134e80b3b4083bd22c462abc1bce74103e

Observation 47e44b60-ac6b-4a55-9448-b864553eae8e · outbound

This paper cites Baichuan-omni-1.5 technical report.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Baichuan-omni-1.5 technical report

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.948374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:f5c5a1514cb1c10cba7c0899c3f0b87518d04a7b7b0c0866d02ca79b574a704a

Observation f3ebd363-802d-46ed-b0c1-22820ed014bf · outbound

This paper cites Towards Better Instruction Following Language Models for Chinese: Investigating the Impact of Training Data and Evaluation.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Towards Better Instruction Following Language Models for Chinese: Investigating the Impact of Training Data and Evaluation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.996815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:7b06861b728bf69a766efbe92349f329a0b131a4a15a62a22ff52041e2a9f02d

Observation 520d023e-46ba-4a4d-9dc3-fb488f1c84ce · outbound

This paper cites Kimi-Audio Technical Report.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Kimi-Audio Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:34.934034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:ff9f41292b522cbd1b6499424e48e3ac37b803e0d9d308e02b10ebe19ddc54d2

Observation a3aada79-a49b-4af6-b4b1-98a1a95fe510 · outbound

This paper cites OpenAssistant Conversations -- Democratizing Large Language Model Alignment.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction OpenAssistant Conversations -- Democratizing Large Language Model Alignment

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.964427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:97d6a303baf3cafe26864e1b8ccacd2558ceef97f359f148ca633fd1fad77320

Observation 95185f7e-3b6e-4b9c-8acf-470044852c07 · outbound

This paper cites InIEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025, Honolulu, HI, USA, December 6-10, 2025, pages 1–8.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction InIEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025, Honolulu, HI, USA, December 6-10, 2025, pages 1–8

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T15:22:02.107863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:bd548960a24efaade932b628ec21d0cbe5cc0054449ddac413731c11839c1c90

Observation da5bd0c2-fcbc-43ca-b841-79eaff1e1b73 · outbound

This paper cites CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction CleanS2S: Single-file Framework for Proactive Speech-to-Speech Interaction

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.994077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:737d7f59851969c09de3e832fc891305ab895688a504f8f9e953e407df571785

Observation 719b2b1d-0a84-4af9-8df2-24f3c64f0019 · outbound

This paper cites an unresolved cited work.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-27T15:22:02.107863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:9d68730183298f4edfa674ea8d9aa985be371d7945a039631b3929d50ae2cbee

Observation 1b0a8841-f964-46db-991b-6d8ad15baee3 · outbound

This paper cites Generative Spoken Dialogue Language Modeling.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Generative Spoken Dialogue Language Modeling

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.967134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:3a5a929390bb66df4c29c54a0919fd99cf9a2739514f76fef7fab20629d44265

Observation 1424cab8-256d-4987-901d-2e21e602c5b8 · outbound

This paper cites GPT-4o System Card.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction GPT-4o System Card

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.925449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:4498da461f6398e421fa73fdaa6f1da4d2167cc73225c0bc79c17efd9f814632

Observation f486d2a1-263a-4b1a-b8bd-ad83515765f5 · outbound

This paper cites In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2015, South Brisbane, Queensland, Australia, April 19-24, 2015, pages 5206–5210.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction In 2015 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP 2015, South Brisbane, Queensland, Australia, April 19-24, 2015, pages 5206–5210

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T15:22:02.107863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:f1808953ebca5701a677e05a461c05f791ba4892605ab06fa19e902aebb544e9

Observation 1540ff3a-9419-411e-b956-0665f08be33c · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.951605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:a08de9a6531065fc33f69777b2d9d9258f8641ac1b0f773635fc62e475a72ba9

Observation 4c428786-99d2-4a8a-a407-84414cd2cada · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.984043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:021910a1e18b76689efd46cf4dea3db16da7167e00b882e621d6a1613a4975e5

Observation 45001c07-ce4d-46b9-982f-cc7b3b081eaa · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.933850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:ecfe19f837bd9aada2906a3ae2ff9d707c1b4579030e3fa67b089555bf4ee90b

Observation b1ba49d5-88d7-4bb1-a437-86645d3b9ac5 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Gemini: A Family of Highly Capable Multimodal Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:34.964777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:ac4573087d1929c3c990ba4ffb0478d59d0f3e1804b8b238fbcd77dee46bb9ac

Observation 6075f2df-37ec-4ec1-9a62-71f21f480513 · outbound

This paper cites Fun-audio-chat technical report.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Fun-audio-chat technical report

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.982969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:89de18a4897eea1063a7587c67f5a2d52045180849006e25bc7409146329fa97

Observation 4fbddfca-2ba8-4d45-9932-5fc094535bec · outbound

This paper cites A Full-duplex Speech Dialogue Scheme Based On Large Language Models.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction A Full-duplex Speech Dialogue Scheme Based On Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.976859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:158f7eb24f85c086656d4a8ec58c8369ba6e69311f4904e0177bd035ed9b2b7e

Observation 677c4174-76b2-4d78-a095-13dd3c96d46b · outbound

This paper cites Mimo-audio: Audio language models are few-shot learners.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Mimo-audio: Audio language models are few-shot learners

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.987343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:88e98a671b9b1a850f90c5ad58deaed81d35b33577b2d8e556568635997178c5

Observation 6a7e3871-3f95-4a75-81f1-ed871e2981fe · outbound

This paper cites Mini-omni-reasoner: Token-level thinking-in-speaking in large speech models.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Mini-omni-reasoner: Token-level thinking-in-speaking in large speech models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.980125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:ad185fca4c2155c7ed6b66b64813f124d86434a0ecd6b191b751d24afa35ddb4

Observation 1c71c09e-45b4-4f30-ad3c-be4af4142bce · outbound

This paper cites Qwen2.5-Omni Technical Report.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Qwen2.5-Omni Technical Report

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.942903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:dc1a5e2f713bd39a60ec5f73b9b8e6c53d49c559e37059abbbcf0a1588ba81ed

Observation 5a806667-89b4-40b5-ab16-1e6a32e8592b · outbound

This paper cites Duplexcascade: Full- duplex speech-to-speech dialogue with vad-free cascaded asr- llm-tts pipeline and micro-turn optimization,.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Duplexcascade: Full- duplex speech-to-speech dialogue with vad-free cascaded asr- llm-tts pipeline and micro-turn optimization,

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.980639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:a11f46cc33cd2000738ea3128042b7d500fc31295933e5a330799800e7791858

Observation 830949d4-e30d-4e7b-b065-6f7a62fae384 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 26

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T03:27:34.973705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:b93590d556a3ac01c225aeb8bc113ca018224582105e4b7a3fc7cdd844d0bc9d

Observation 030fe35c-2674-446b-b5c0-b3fc2ffab60a · outbound

This paper cites SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction SALMONN-omni: A Standalone Speech LLM without Codec Injection for Full-duplex Conversation

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.936721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:7fe6422796f711d8bdf935f43fcf367087f23e9d1196b1413cb1021fbcc16fad

Observation e301429a-ae85-490a-b8ac-a531e9cb089e · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:34.931397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:a59517ef018dfe8cf7f711647cc900f8599fd32d40307433a0875baf85776ec4

Observation 569fe23b-d8b0-46b2-9cab-4d122910491d · outbound

This paper cites Mlvu: Benchmarking multi-task long video understanding.

DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction Mlvu: Benchmarking multi-task long video understanding

Reference 29

Resolution
malformed identifier
arxiv_id, observed 2026-07-03T03:27:34.970446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T15:22:02.107863Z digest=sha256:778558509371f51e5c0de24635ca7a02852e0ec502849f476539cd61f1c679f5

Pith citing papers

Observation 89b5d4f3-5ebc-40ac-b54e-9fcb33ce1ef0 · inbound

Voice Memory for Agentic Speech Recognition cites this paper.

Voice Memory for Agentic Speech Recognition DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T16:37:24.214630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:37:24.214630Z digest=sha256:becc692e42c3132e7956875f3a4619d3e5f522f455895753b4a66269f7e1faa3