Pith. sign in

Paper Citation Record · LEDGER

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training

As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2506.07081.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07081 v2

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:47:14.172217Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T05:37:17.684884Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T16:28:39.127928Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 44080555-c607-49e4-a760-a70df34c9de0 · outbound

This paper cites WavChat: A Survey of Spoken Dialogue Models.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training WavChat: A Survey of Spoken Dialogue Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:11.320572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:11.320572Z digest=sha256:0827e12e6a86f3a943e2c9c8a6f203ac9c39bc8139a12700d1b29339556aa97f

Observation 1ca75ee5-98f1-49b8-9bce-9eed04c1a35a · outbound

This paper cites A review of subjective scales measuring the user experience of voice assistants,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training A review of subjective scales measuring the user experience of voice assistants,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.964681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:11.409467Z digest=sha256:d4f544da19be01d5b632ba104215bf59d9cf1fa42b2b3d7e07bdfcdf06a7e27d

Observation 54c688bf-1a8f-4aa7-8334-91adfada3c3c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Gemini: A Family of Highly Capable Multimodal Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:11.481844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:11.481844Z digest=sha256:3e7b6baf33e78cfec9fe00b43b963e35afac680cd3df6a305e70daedd12284c0

Observation 0f0da420-5713-49ff-9605-d978e32a8c85 · outbound

This paper cites OpenAI-gpt-4o,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training OpenAI-gpt-4o,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.794138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:11.526889Z digest=sha256:89778a8ecd0b14a1e9fc348c0da73d3b310351b683f78ae864087731ae17159a

Observation 261a5a25-60ab-49de-b987-497f08b38a76 · outbound

This paper cites Improved End- of-Query Detection for Streaming Speech Recognition,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Improved End- of-Query Detection for Streaming Speech Recognition,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.682409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:11.656767Z digest=sha256:d3381f793f7390fa37904606281f175ebb5b72e6549cee74edecc02a7371b43f

Observation be2a5081-a9d9-4e68-9d62-720bfb6fda20 · outbound

This paper cites V oice activity projection: Self-supervised learning of turn-taking events,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training V oice activity projection: Self-supervised learning of turn-taking events,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.539409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:11.711258Z digest=sha256:c92cdb3d8660b81acf06783e594d3b789ab706a3ade78aa2e791591032a1f2a4

Observation 651c22f7-0f2f-4f6f-842d-169ef6f7cf64 · outbound

This paper cites Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics ,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Talking Turns: Benchmarking Audio Foundation Models on Turn-Taking Dynamics ,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.396827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:11.779580Z digest=sha256:9de286103367fda2aa29516bd2a61565fdf481713950921af7e8e53fd4058b00

Observation bf9ee3b0-9255-4331-84b8-4392ca4b0ce8 · outbound

This paper cites Root causes of lost time and user stress in a simple dialog system,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Root causes of lost time and user stress in a simple dialog system,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.274468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:11.855592Z digest=sha256:c5c8da299ccb5fa9c75e251bdd75f9f6e93d5761d3807ff89295278060173c36

Observation 96915e7a-187b-4b5b-b6f9-52620342af30 · outbound

This paper cites A statistical model-based voice activity detection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training A statistical model-based voice activity detection,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:19.096855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:11.946892Z digest=sha256:80ee5cf40bd0cc24e789885b3d4eb45ff43ccd403a53a4526bc4c95484ed6ab1

Observation 1f4f2787-25e3-4410-b0d1-4d46fba5cc3d · outbound

This paper cites A Convolutional Neural Network Smartphone App for Real-Time V oice Activity Detection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training A Convolutional Neural Network Smartphone App for Real-Time V oice Activity Detection,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.965610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:12.001979Z digest=sha256:471f52a771405ca9bac122801d880f033a3f83987f0863c421bdd8d0b595a368

Observation b96e219c-d423-4ddf-9bd5-d3b65f4c9073 · outbound

This paper cites Temporal modeling using dilated convolution and gating for voice-activity-detection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Temporal modeling using dilated convolution and gating for voice-activity-detection,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.753204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:12.092752Z digest=sha256:2086cdc44d4629fabf57e2e24121fa363abb2c0cea1b719792312766e9e3cb2d

Observation 7ad4af9c-0153-444c-8f67-863e23fd20bd · outbound

This paper cites Robust end-of-utterance detection for real-time speech recognition applications,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Robust end-of-utterance detection for real-time speech recognition applications,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.629364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:12.157555Z digest=sha256:7ef518b74c2a4188471f4a2c8893fee7e8337ff26ba245185102d4ae6b3f7101

Observation f50f5b3e-03d1-4f6b-8d10-1ec217bbbdf2 · outbound

This paper cites Combining acoustic embeddings and decoding features for end-of-utterance detection in real-time far-field speech recognition systems,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Combining acoustic embeddings and decoding features for end-of-utterance detection in real-time far-field speech recognition systems,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.503583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:12.248057Z digest=sha256:e580c4fd28f5277543bd98f96e3feb7aaa77d78cd857c8d88c3a76bab2294d73

Observation 5dc9aed4-5c79-4292-830b-050f5772222d · outbound

This paper cites Dynamic speech endpoint detection with regression targets,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Dynamic speech endpoint detection with regression targets,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.308518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:12.341985Z digest=sha256:2f4511aaa440afe6853ef5932423cec5a545a48d34c716e92e67704473e97db2

Observation b86fa938-e980-45c9-a34f-227472dbd7e5 · outbound

This paper cites SoundStream: An End-to-End Neural Audio Codec,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training SoundStream: An End-to-End Neural Audio Codec,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:18.141871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:12.410195Z digest=sha256:046b96a7d967d18c69e22e048383096062ab94b83815ed32d5217c0931ae8bd4

Observation 34bd2dc7-2de0-4b9d-a23b-3ffbca6bbfe4 · outbound

This paper cites High Fidelity Neural Audio Compression,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training High Fidelity Neural Audio Compression,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.978334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:12.520620Z digest=sha256:3b85ca6f00d3c1ab85b1ea854746f9082f1724dfed9360a871be316f5fc6a5e9

Observation 17e1c7d0-bc0f-4cfc-81f8-2d25bb286c4e · outbound

This paper cites Audiodec: An Open-Source Streaming High-Fidelity Neural Audio Codec,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Audiodec: An Open-Source Streaming High-Fidelity Neural Audio Codec,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.830311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:12.585964Z digest=sha256:bd1be78e19a238d15b79b33872aed7bdaa43fa3f8d3cb4f335cc0f5f0c3dfb12

Observation 9d41f697-b6d7-4954-95f6-dc506c121022 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Moshi: a speech-text foundation model for real-time dialogue

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:12.652698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:12.652698Z digest=sha256:2073480b2081cfaf62d0a48a137d2166d814c5061cbbfd6395f5dd4485e2e608

Observation a363b232-9cb9-456b-9fd4-f460d74c6255 · outbound

This paper cites Codec-SUPERB: An In-Depth Analysis of Sound Codec Models,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Codec-SUPERB: An In-Depth Analysis of Sound Codec Models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.616111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:12.704887Z digest=sha256:48f84241d537af4817cb29de005656a806ed0b28738934fbc00f9f3c284b2db1

Observation 6595c748-2f80-4c2f-b947-401edb609c36 · outbound

This paper cites ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training ESPnet-Codec: Comprehensive Training and Evaluation of Neural Codecs for Audio, Music, and Speech

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:47:14.407787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:12.794875Z digest=sha256:61edbfd6e9c3038d754816a9b11dd8f37b37d6b359904cadec645da1c57a3004

Observation eae1316f-4479-4d94-9707-99e0bbf49b08 · outbound

This paper cites Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Codec-ASR: Training Performant Automatic Speech Recognition Systems with Discrete Speech Representations,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.453168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:12.857921Z digest=sha256:4b168642b841bf4aa6d2ef2226990f58e7cf0d7f4bb69548e4cfbdc4480f5c83

Observation 9893ded1-943e-47cd-b054-faab046ad98c · outbound

This paper cites High-Fidelity Simultaneous Speech-To-Speech Translation.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training High-Fidelity Simultaneous Speech-To-Speech Translation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:12.941618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:12.941618Z digest=sha256:93dd663f91e66116bfd92eaf1b82107c010cc8eee886bcd80c9cf42262980c0d

Observation 63e01f50-c2ee-408e-a658-eb4bf3d9617c · outbound

This paper cites Self-supervised speech representation learning: A review,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Self-supervised speech representation learning: A review,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:13.003105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:13.003105Z digest=sha256:6cef9aac1e4fdf45d8ae96bf0f34c6cd6ddf1fe333462f4b48d661fdf1cf96df

Observation 6a621974-0784-420e-8bea-cabd6897ef01 · outbound

This paper cites HILCodec: High-Fidelity and Lightweight Neural Audio Codec.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training HILCodec: High-Fidelity and Lightweight Neural Audio Codec

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:13.130003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:13.130003Z digest=sha256:70d6f66621d039f2961fd5512fae752b20e1ffc1fdca1012eae2972d241cd884

Observation abbbbc0c-103e-4429-a95f-dbb4fbdbbd95 · outbound

This paper cites End-to-end speech endpoint detection utilizing acoustic and language modeling knowledge for online low- latency speech recognition,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training End-to-end speech endpoint detection utilizing acoustic and language modeling knowledge for online low- latency speech recognition,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.289254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.214580Z digest=sha256:77543784c95017a5cb84c23826b179e37ea2ab2b4ea0798720298566121814eb

Observation d98b2206-b1c5-4486-9f80-464af49ccf38 · outbound

This paper cites Joint endpointing and decoding with end-to-end models,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Joint endpointing and decoding with end-to-end models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:17.150228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.254678Z digest=sha256:74d724d57abe19e544ca62464a1441ea6193827e38b2c5dc4dd413f87beef297

Observation 4344bebb-23e6-4ca9-83c1-9f4ab23c83a8 · outbound

This paper cites Turn-Taking Prediction for Natural Conversational Speech,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Turn-Taking Prediction for Natural Conversational Speech,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.992507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.295077Z digest=sha256:4a8160f2bccee8e1f30f7e5ca41948ce2365a349fb9af497194d691f8ab25df1

Observation 1aad363c-8c0b-479f-872f-b23e539ee1a3 · outbound

This paper cites Towards fast and accurate streaming end-to-end ASR,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Towards fast and accurate streaming end-to-end ASR,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.879748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.378298Z digest=sha256:6fbbebbd12e7cefe74b118a6a5895ff079870f60eab70f98defc1217bec100c6

Observation 30dea1f6-734e-4b63-899a-fcc8d19ed347 · outbound

This paper cites Streaming Automatic Speech Recognition with Re-blocking Processing Based on Integrated V oice Activity Detection.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Streaming Automatic Speech Recognition with Re-blocking Processing Based on Integrated V oice Activity Detection

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.705567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.441465Z digest=sha256:0391939e59e568a841dba62df017341d9872144e2b676b8a2007d23dfe1a653a

Observation b5881084-3c11-4493-bfc4-749fdef1cbd7 · outbound

This paper cites Two-Pass Endpoint Detection for Speech Recognition,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Two-Pass Endpoint Detection for Speech Recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.631628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.504091Z digest=sha256:822e06387c984b6f4fd96af781fa0f6bd8dd1ac02af84f85a101c23697307487

Observation f50655af-451a-4b4e-9481-990a3529c45f · outbound

This paper cites Towards Accurate and Real-Time End-of-Speech Estimation,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Towards Accurate and Real-Time End-of-Speech Estimation,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.506542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.535657Z digest=sha256:08ae1c80a1c8e21e2d9ea5b0e0aae115524b52ed8332737e7b3a5fa90b4efc40

Observation 473084b1-adfe-47ea-87a7-353b44ab619d · outbound

This paper cites Text Injec- tion for Capitalization and Turn-Taking Prediction in Speech Models,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Text Injec- tion for Capitalization and Turn-Taking Prediction in Speech Models,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.264732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.576137Z digest=sha256:7a899cd8824d8628c63e3bcba63f37ae2458dd30621a526aa6f985885d0ae1dc

Observation 7e73b4fd-9911-40ce-8881-1abafd2b4420 · outbound

This paper cites Multilingual turn-taking prediction using voice activity projection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Multilingual turn-taking prediction using voice activity projection,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:16.112289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.643938Z digest=sha256:5163ad0f33d582adb55cf9f31a5d29e5d118c80259b37b089fe768894d087db5

Observation 89e63db1-7533-4c20-be69-3ed56f734810 · outbound

This paper cites Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Yeah, un, oh: Continuous and real-time backchannel prediction with fine-tuning of voice activity projection,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.914320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.681171Z digest=sha256:22049f7bf6a0bab5b84e63380c9b9395a3aea955b11ad71cdee20fe8445778c8

Observation 45cbefc9-f0af-4ab9-b512-200a9d8bfa12 · outbound

This paper cites Turn-Taking and Backchannel Prediction with Acoustic and Large Language Model Fusion,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Turn-Taking and Backchannel Prediction with Acoustic and Large Language Model Fusion,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.745811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.728539Z digest=sha256:50b243e30f7b2956ba2d4da6619b2dc2b66395d7e79fa252471b473d31aaf739

Observation 9c4f5350-1927-4723-af50-68fd90cd8876 · outbound

This paper cites End-to-end automatic speech recognition integrated with CTC-based voice activity detection,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training End-to-end automatic speech recognition integrated with CTC-based voice activity detection,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.537482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.756574Z digest=sha256:8fdbac07d821f4e3610adbb4e57ed0bbbf4bdac79f489b43b2e7b8f15575dad4

Observation 4debdea7-6324-49e2-aece-de6ddf6788b2 · outbound

This paper cites Text- free prosody-aware generative spoken language modeling,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Text- free prosody-aware generative spoken language modeling,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.345486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.809131Z digest=sha256:7cf4c7b89b836e732401a6bf1bd01743fa064ab8936c65fe582382b483d328a8

Observation dc82655a-cacc-48b4-b0f3-ee6fd9074f98 · outbound

This paper cites Simple and controllable music generation,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Simple and controllable music generation,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.129513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.835391Z digest=sha256:4eab21855d0fc466c7f78b8a72d705e1d8ba336c44cc829dc69b67d355e07612

Observation 43085501-452b-4b50-97a8-6d7ec417e90f · outbound

This paper cites SpokenWOZ: a large-scale speech-text benchmark for spoken task-oriented dialogue agents,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training SpokenWOZ: a large-scale speech-text benchmark for spoken task-oriented dialogue agents,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:15.062955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.880359Z digest=sha256:8cf754d7153face6084a0dabf6fa39953e73b6b068d7d4dd1ca35b4a6b9314ab

Observation 60401cb3-9b9e-4899-8663-ca7f1a9d073f · outbound

This paper cites Silero V AD: pre-trained enterprise-grade V oice Activity Detector (V AD), Number Detector and Language Classifier,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Silero V AD: pre-trained enterprise-grade V oice Activity Detector (V AD), Number Detector and Language Classifier,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:13.921334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:13.921334Z digest=sha256:221fe01328a1dc64a291930c260311e720dd46020ce9959fb20a244b3058cef6

Observation c425acdf-4ba8-4ec4-9a02-d5a7cb25ca4e · outbound

This paper cites PyTorch: an imperative style, high- performance deep learning library,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training PyTorch: an imperative style, high- performance deep learning library,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:14.816798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:13.960499Z digest=sha256:8d223440fca43b690fdfc83564df54164b0d6093e03dfd9b30c70763c618e042

Observation fc38f094-28c5-4cad-b465-4167ba175604 · outbound

This paper cites Transformers: State-of-the-art natural language processing,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Transformers: State-of-the-art natural language processing,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:14.679549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:14.010240Z digest=sha256:cd92a694c0a99ef15875ce3a1eef2e94cc929afc42f8c9c5e3664f019face419

Observation c25e00e3-6b18-4177-a50b-80c08dcf091c · outbound

This paper cites Modeling turn-taking in human-to-human spoken dialogue datasets us- ing self-supervised features,.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training Modeling turn-taking in human-to-human spoken dialogue datasets us- ing self-supervised features,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:47:14.564865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:47:14.075639Z digest=sha256:c86bce23f3b7cc7a9983cd395c5f24bb951eb05faef0fe879f3028803901ea9b

Observation 2768c705-0e2b-4bd9-9b30-d63fb9d189e3 · outbound

This paper cites LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training LLM-Enhanced Dialogue Management for Full-Duplex Spoken Dialogue Systems

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:14.122383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:14.122383Z digest=sha256:7d7bd0d690d6025d3442356e6a26859082c1afa1b63b6fe8da04a442dca7a8fb

Observation 0c79f311-4e44-4233-a9bf-7396fc98ea25 · outbound

This paper cites MinMo: A Multimodal Large Language Model for Seamless Voice Interaction.

Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:47:14.172217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:47:14.172217Z digest=sha256:3d68a9a57c756b1231089821a11952ce45a5ca4ffdbd67c65d5b0024e28315a2

Pith citing papers

Observation 1f8ee3e4-80fe-4a9d-bb4c-5c66b4c1ee93 · inbound

Endpoint Anticipation for Low-Latency Spoken Dialogue cites this paper.

Endpoint Anticipation for Low-Latency Spoken Dialogue Streaming Endpointer for Spoken Dialogue using Neural Audio Codecs and Label-Delayed Training

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:28:39.129488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T05:37:17.684884Z digest=sha256:e55fb13809568e09522bac8ba2da60130cb2cc2736b2cd0f801ed002e154c685