Pith. sign in

Paper Citation Record · LEDGER

Improving endpoint detection in end-to-end streaming ASR for conversational speech

As of 19 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2505.17070.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17070 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:22:18.584812Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:22:18.441360Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T20:22:18.688537Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy19
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d098c91-0aea-43bc-bafb-ae6b816383b0 · outbound

This paper cites Improving endpoint detection in end-to-end streaming ASR for conversational speech.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Improving endpoint detection in end-to-end streaming ASR for conversational speech

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:22:18.694131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.441360Z digest=sha256:563286911be717d6cd552cfc7d04475e078409bafe3640718f5c467925713d45

Observation 9ca48c20-592c-4dff-b0b6-664ca26346a2 · outbound

This paper cites occasional.

Improving endpoint detection in end-to-end streaming ASR for conversational speech occasional

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:19.097652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.447793Z digest=sha256:e22d901a17a5afe78f0b9c6827b7f5d9566c7134b330c1276845e5f4a1c9d26a

Observation 93ba72f9-3b61-4d80-8087-0a4d9449e9e4 · outbound

This paper cites an unresolved cited work.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:22:19.080975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.453071Z digest=sha256:26ca42d4eb9b0c0eec8c2111219644bc5abe02aaf2193479082d0f5768a6124e

Observation 0e80cce5-189b-4a48-a701-c7d81033d4a0 · outbound

This paper cites A separate speech detector network operates in parallel with the ASR decoder to determine speech/non-speech at the frame-level.

Improving endpoint detection in end-to-end streaming ASR for conversational speech A separate speech detector network operates in parallel with the ASR decoder to determine speech/non-speech at the frame-level

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:19.063531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.458365Z digest=sha256:949ea5c64b77350cd8b44b7546612a448ba2ebd9b3a00f70c2f07d5b4a6cef07

Observation 31e2f763-80fe-4979-b645-63a334e0dd78 · outbound

This paper cites Towards end-to-end speech recognition with recurrent neural networks,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Towards end-to-end speech recognition with recurrent neural networks,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:19.045760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.463719Z digest=sha256:b83da3490052453c30cde5cc354024387d00d9845440d31e5ff1a00c441dbeaf

Observation d2d59993-5d49-4a65-a1cd-4a4847d02ea9 · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Sequence Transduction with Recurrent Neural Networks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:22:18.469004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:22:18.469004Z digest=sha256:656739f2c9e481ce7f2db82ebb8461099ea60570c365e3c2a90295584c490328

Observation abc6b5c8-a755-4754-996d-cf85a35cf669 · outbound

This paper cites End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results.

Improving endpoint detection in end-to-end streaming ASR for conversational speech End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:22:18.474739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:22:18.474739Z digest=sha256:a3b83085298a50b61d4747df777a154953b35f782887224c2d56c6dd3c4dea61

Observation b4f7863f-ca48-4bf7-a2ba-8861ed6198a1 · outbound

This paper cites Online and linear-time attention by enforcing monotonic alignments,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Online and linear-time attention by enforcing monotonic alignments,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:22:18.480105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:22:18.480105Z digest=sha256:724fadae0843f05c484d58552ed5f3d3871e71aae4a4c20d7a758f26fd526535

Observation 5ece3553-1c55-4190-9353-eea097327c0b · outbound

This paper cites Joint ctc-attention based end-to-end speech recognition using multi-task learning,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Joint ctc-attention based end-to-end speech recognition using multi-task learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:22:18.485624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:22:18.485624Z digest=sha256:1f95c765103938d3083a21135f2d84332aba7faff646a2bd6c600d91768dcc63

Observation 0e63d94b-f795-472d-9510-ab4f13d77206 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Conformer: Convolution-augmented Transformer for Speech Recognition,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:19.002999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.491193Z digest=sha256:607ab930e482b35bdfa66eb8e3eaaa9d63c92075710d83354b0cac5eaf768256

Observation 3b3c1adb-57ec-4ae9-bfb2-b2b42c11bf89 · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Zipformer: A faster and better encoder for automatic speech recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.985886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.495977Z digest=sha256:75903c5d069ef726791b2d3a9c2eda9c07f7817da7ce46ea203a0cfda805a696

Observation 86e1d173-1fdc-498e-a31e-4c4853784067 · outbound

This paper cites A com- parison of streaming models and data augmentation methods for robust speech recognition,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech A com- parison of streaming models and data augmentation methods for robust speech recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.968882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.501158Z digest=sha256:390c8f4df4a3f25f730bd9e4c553cbcca5fed308e73a37551a29824d5ca7c9c0

Observation fb2d1710-5b08-495a-8068-1dee596eaa2b · outbound

This paper cites Towards fast and accurate streaming end-to-end asr,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Towards fast and accurate streaming end-to-end asr,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.952724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.506505Z digest=sha256:6b98585eb4773927c847b918c902413c1ad406df6628965d3e86f79fd1595ef3

Observation cbe5b4fd-5699-47b8-98f3-66488ce8ba74 · outbound

This paper cites Fastemit: Low- latency streaming asr with sequence-level emission regulariza- tion,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Fastemit: Low- latency streaming asr with sequence-level emission regulariza- tion,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.936700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.511465Z digest=sha256:d824f5ceeb55f7f185cbff7d7bdd6875ce010907026c918897131528b1280f80

Observation 2688c00f-c026-4ab7-9104-bb2a1e313f6a · outbound

This paper cites Delay-penalized transducer for low- latency streaming asr,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Delay-penalized transducer for low- latency streaming asr,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.919701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.516283Z digest=sha256:c84506dc9b1a63f77d4a786b310a53d0560b8e4c6c723f64c920f75671b28a65

Observation e16103be-7be4-4946-83a2-4d81ec5b19f2 · outbound

This paper cites Alignment restricted stream- ing recurrent neural network transducer,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Alignment restricted stream- ing recurrent neural network transducer,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.902297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.520871Z digest=sha256:02ac94ff0023cb3f994772b61dd5f9f9560d02ebb5b47df433cf8019e4e4b642

Observation 92e7fefd-1b17-4617-bedb-0d7abf092add · outbound

This paper cites Reducing Streaming ASR Model Delay with Self Alignment,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Reducing Streaming ASR Model Delay with Self Alignment,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.884025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.525456Z digest=sha256:fdf30aa8e8782be7991fa4f4fc33dfb3c5729e1a73bfdb0563259cc6e8e6b622

Observation fa69a3f4-e68c-445c-b836-7d30e7e59682 · outbound

This paper cites Minimum latency training of se- quence transducers for streaming end-to-end speech recognition,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Minimum latency training of se- quence transducers for streaming end-to-end speech recognition,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.867526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.530478Z digest=sha256:d6a8acbd3cbb6a51d497786a80486cffba17a3bb1aeb96ccec6fe0c254952d96

Observation d89f23b3-7f93-4b3b-8ba0-6b8e3c768eb8 · outbound

This paper cites Endpoint detection for streaming end- to-end multi-talker ASR,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Endpoint detection for streaming end- to-end multi-talker ASR,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.848576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.535402Z digest=sha256:9339fb22400f9f220e7b0c6936e9b890c66c5d85ddd4f583f6af46c4d79dc659

Observation 844a9ba7-818b-4c01-ae43-2c484e110579 · outbound

This paper cites Towards accurate and real-time end-of-speech estimation,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Towards accurate and real-time end-of-speech estimation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.830818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.540400Z digest=sha256:1618bb1489f3553748173ae6b9cb4a55dbb6f11eea9302cd0b03f854f6cb9c60

Observation 45492703-8431-40cf-8600-5e972c8d3786 · outbound

This paper cites Unified end-to-end speech recognition and endpointing for fast and efficient speech systems,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Unified end-to-end speech recognition and endpointing for fast and efficient speech systems,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.814063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.545562Z digest=sha256:0cb92a94ebefdddde8d4f63eabe0970b71c994d9d455e717b783f7dd030f324d

Observation 9103b586-6564-40f6-b523-c03eac0cac57 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Lib- rispeech: an asr corpus based on public domain audio books,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:22:18.550553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:22:18.550553Z digest=sha256:c1e32b247ae6fe7c3f17ba52901dcb05dc22013ae3302f8e08ed7475e56ea0b0

Observation 0f5a4116-3d3c-4455-b92e-97865e07ea09 · outbound

This paper cites Switchboard: Telephone speech corpus for research and development,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Switchboard: Telephone speech corpus for research and development,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.785063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.555368Z digest=sha256:225125f4268a30eed0c43d118c82461afc1bcac7d9595d677d8d91303ff13304

Observation 06b9a936-1060-442b-a263-e36887451514 · outbound

This paper cites Is the speaker done yet? Faster and more accurate end-of-utterance detection using prosody,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Is the speaker done yet? Faster and more accurate end-of-utterance detection using prosody,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.768195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.560359Z digest=sha256:c7276605765b2540b9af4f78ba7d24c9d285f57302f2e95d4e002a0c2d1bd3e8

Observation 365e75d5-2a22-4418-ba43-7a00be30e260 · outbound

This paper cites Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:22:18.632786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.565258Z digest=sha256:71ad5ad483f8c77c8526f288d7d7e308b31accf9a04a8d8786e1b806260ac176

Observation f35c6df4-6f89-4af8-b87d-e10f28cb3822 · outbound

This paper cites Dissecting User-Perceived Latency of On-Device E2E Speech Recognition,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Dissecting User-Perceived Latency of On-Device E2E Speech Recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.751414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.570629Z digest=sha256:7c58f9067cdfda8a5abce245fdd3c4503fbaa4dbf2897ad1f3f35a52e78fa555

Observation 777baf3b-7a0b-4fa2-854d-62c7daa4eb1b · outbound

This paper cites Dialogue act modeling for automatic tagging and recognition of conversational speech,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Dialogue act modeling for automatic tagging and recognition of conversational speech,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:22:18.575241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:22:18.575241Z digest=sha256:6c7a77dfa7802608b3ab5da8aa4f774eb0b171c83a7915a9b5dfb5a86564f6b8

Observation d2b41aee-662a-45d4-96a8-f259e2818f33 · outbound

This paper cites Metrics for polyphonic sound event detection,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Metrics for polyphonic sound event detection,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:22:18.580035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:22:18.580035Z digest=sha256:d29c30e70b20c8b93643925e0c6f4ad7338eece596327554891c113ad18bf859

Observation 4cfcf93e-1fca-4971-a299-06f6deaa6940 · outbound

This paper cites The third DIHARD di- arization challenge,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech The third DIHARD di- arization challenge,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.711766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.584812Z digest=sha256:03726d73753b663aa2b1463d8d93bce5d1d268ee181f08d8b0f22072b5d2809b

Pith citing papers

Observation 4d098c91-0aea-43bc-bafb-ae6b816383b0 · inbound

Improving endpoint detection in end-to-end streaming ASR for conversational speech cites this paper.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Improving endpoint detection in end-to-end streaming ASR for conversational speech

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:22:18.694131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T20:22:18.441360Z digest=sha256:563286911be717d6cd552cfc7d04475e078409bafe3640718f5c467925713d45