Pith. sign in

Paper Citation Record · LEDGER

Improving endpoint detection in end-to-end streaming ASR for conversational speech

As of 23 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2505.17070.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17070 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:22:18.584812Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:22:18.441360Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T20:22:18.688537Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy19
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d098c91-0aea-43bc-bafb-ae6b816383b0 · outbound

This paper cites Improving endpoint detection in end-to-end streaming ASR for conversational speech.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Improving endpoint detection in end-to-end streaming ASR for conversational speech

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:22:18.694131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.441360Z digest=sha256:83a1c9c19ba28135e1b19e1e36d328b2b183900de5bf2e8a60ed4691640c3f05

Observation 9ca48c20-592c-4dff-b0b6-664ca26346a2 · outbound

This paper cites occasional.

Improving endpoint detection in end-to-end streaming ASR for conversational speech occasional

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:19.097652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.447793Z digest=sha256:2313be30e848b21af5071e4b51393038fd51995c93cbce8a217d3d63d63787be

Observation 93ba72f9-3b61-4d80-8087-0a4d9449e9e4 · outbound

This paper cites an unresolved cited work.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-15T20:22:19.080975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.453071Z digest=sha256:a1d0d202b820156f69830e731f339d232691087619360220fa0bb0cc63d4756e

Observation 0e80cce5-189b-4a48-a701-c7d81033d4a0 · outbound

This paper cites A separate speech detector network operates in parallel with the ASR decoder to determine speech/non-speech at the frame-level.

Improving endpoint detection in end-to-end streaming ASR for conversational speech A separate speech detector network operates in parallel with the ASR decoder to determine speech/non-speech at the frame-level

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:19.063531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.458365Z digest=sha256:b27d735267dc5b043c5ba48731cdb5346819d0617659ecb9eac962a604038c85

Observation 31e2f763-80fe-4979-b645-63a334e0dd78 · outbound

This paper cites Towards end-to-end speech recognition with recurrent neural networks,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Towards end-to-end speech recognition with recurrent neural networks,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:19.045760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.463719Z digest=sha256:bd464125107d6bd07089eac4b3c5740213079c39b0eb4712d2c5c9cca042f739

Observation d2d59993-5d49-4a65-a1cd-4a4847d02ea9 · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Sequence Transduction with Recurrent Neural Networks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:22:18.469004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:22:18.469004Z digest=sha256:656739f2c9e481ce7f2db82ebb8461099ea60570c365e3c2a90295584c490328

Observation abc6b5c8-a755-4754-996d-cf85a35cf669 · outbound

This paper cites End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results.

Improving endpoint detection in end-to-end streaming ASR for conversational speech End-to-end Continuous Speech Recognition using Attention-based Recurrent NN: First Results

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:22:18.474739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:22:18.474739Z digest=sha256:a3b83085298a50b61d4747df777a154953b35f782887224c2d56c6dd3c4dea61

Observation b4f7863f-ca48-4bf7-a2ba-8861ed6198a1 · outbound

This paper cites Online and linear-time attention by enforcing monotonic alignments,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Online and linear-time attention by enforcing monotonic alignments,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:22:18.480105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:22:18.480105Z digest=sha256:724fadae0843f05c484d58552ed5f3d3871e71aae4a4c20d7a758f26fd526535

Observation 5ece3553-1c55-4190-9353-eea097327c0b · outbound

This paper cites Joint ctc-attention based end-to-end speech recognition using multi-task learning,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Joint ctc-attention based end-to-end speech recognition using multi-task learning,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:22:18.485624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:22:18.485624Z digest=sha256:1f95c765103938d3083a21135f2d84332aba7faff646a2bd6c600d91768dcc63

Observation 0e63d94b-f795-472d-9510-ab4f13d77206 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Conformer: Convolution-augmented Transformer for Speech Recognition,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:19.002999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.491193Z digest=sha256:d80f3a27e9a643b27d3e0ca065b6470fdec4f0bf5ee9efa9386afb888c3745d4

Observation 3b3c1adb-57ec-4ae9-bfb2-b2b42c11bf89 · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Zipformer: A faster and better encoder for automatic speech recognition,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.985886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.495977Z digest=sha256:68c7c8509e9f4c34146a27276a6daed47a2c765427386e9c00fe98c8bfb54966

Observation 86e1d173-1fdc-498e-a31e-4c4853784067 · outbound

This paper cites A com- parison of streaming models and data augmentation methods for robust speech recognition,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech A com- parison of streaming models and data augmentation methods for robust speech recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.968882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.501158Z digest=sha256:13d11bb682b1f8741bd13235f4554d786b58b62eff188ed2cab09f4eb24770d6

Observation fb2d1710-5b08-495a-8068-1dee596eaa2b · outbound

This paper cites Towards fast and accurate streaming end-to-end asr,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Towards fast and accurate streaming end-to-end asr,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.952724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.506505Z digest=sha256:60b8f6e4b90e18fd1850f1e8c098e11edaa66e5ec5725e71a5b9590e992e957a

Observation cbe5b4fd-5699-47b8-98f3-66488ce8ba74 · outbound

This paper cites Fastemit: Low- latency streaming asr with sequence-level emission regulariza- tion,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Fastemit: Low- latency streaming asr with sequence-level emission regulariza- tion,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.936700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.511465Z digest=sha256:dd95c129835d8e6c31466edceeee0d57e5d0cb397a88988465d62b83e64135ab

Observation 2688c00f-c026-4ab7-9104-bb2a1e313f6a · outbound

This paper cites Delay-penalized transducer for low- latency streaming asr,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Delay-penalized transducer for low- latency streaming asr,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.919701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.516283Z digest=sha256:31935f2d0a726798f03ce180d9ee2de8a4cb136628b9373528733d6fea31da40

Observation e16103be-7be4-4946-83a2-4d81ec5b19f2 · outbound

This paper cites Alignment restricted stream- ing recurrent neural network transducer,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Alignment restricted stream- ing recurrent neural network transducer,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.902297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.520871Z digest=sha256:24f6df3f84c23b986306d7d07ef65294cb7abef2e04b93234c73b6f45011678e

Observation 92e7fefd-1b17-4617-bedb-0d7abf092add · outbound

This paper cites Reducing Streaming ASR Model Delay with Self Alignment,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Reducing Streaming ASR Model Delay with Self Alignment,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.884025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.525456Z digest=sha256:ad6a1777ec2e212c579d547bba926b73378ac5cc6410780bed05afdb141d71a6

Observation fa69a3f4-e68c-445c-b836-7d30e7e59682 · outbound

This paper cites Minimum latency training of se- quence transducers for streaming end-to-end speech recognition,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Minimum latency training of se- quence transducers for streaming end-to-end speech recognition,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.867526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.530478Z digest=sha256:861d01de731fd255b23753163c73d1da5b70dbcb0e6c2efc1a209e73ab30494a

Observation d89f23b3-7f93-4b3b-8ba0-6b8e3c768eb8 · outbound

This paper cites Endpoint detection for streaming end- to-end multi-talker ASR,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Endpoint detection for streaming end- to-end multi-talker ASR,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.848576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.535402Z digest=sha256:bc343910be21df077b3d6abe43eb3b4c1cbba5db7af09c79e3342bacf5bb6f46

Observation 844a9ba7-818b-4c01-ae43-2c484e110579 · outbound

This paper cites Towards accurate and real-time end-of-speech estimation,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Towards accurate and real-time end-of-speech estimation,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.830818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.540400Z digest=sha256:1651eba558e0c65060f9e462a767d461c51a3b321846f5ff75831c988de6ab9e

Observation 45492703-8431-40cf-8600-5e972c8d3786 · outbound

This paper cites Unified end-to-end speech recognition and endpointing for fast and efficient speech systems,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Unified end-to-end speech recognition and endpointing for fast and efficient speech systems,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.814063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.545562Z digest=sha256:2a0f48ed8c6804a619368e2d830513b774dc1b7434efd055013b128eaf30ed3e

Observation 9103b586-6564-40f6-b523-c03eac0cac57 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Lib- rispeech: an asr corpus based on public domain audio books,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:22:18.550553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:22:18.550553Z digest=sha256:c1e32b247ae6fe7c3f17ba52901dcb05dc22013ae3302f8e08ed7475e56ea0b0

Observation 0f5a4116-3d3c-4455-b92e-97865e07ea09 · outbound

This paper cites Switchboard: Telephone speech corpus for research and development,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Switchboard: Telephone speech corpus for research and development,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.785063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.555368Z digest=sha256:53a1e24948b6a675f992fa8c67e805b6cce2f2b43bcf0c28bf7ab4d5c30db523

Observation 06b9a936-1060-442b-a263-e36887451514 · outbound

This paper cites Is the speaker done yet? Faster and more accurate end-of-utterance detection using prosody,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Is the speaker done yet? Faster and more accurate end-of-utterance detection using prosody,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.768195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.560359Z digest=sha256:d677ab366cdc8192eed8ebadf005aa3d819f5ee9f04a39345711ac2d4f0f4e95

Observation 365e75d5-2a22-4418-ba43-7a00be30e260 · outbound

This paper cites Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Turn-taking and Backchannel Prediction with Acoustic and Large Language Model Fusion

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:22:18.632786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.565258Z digest=sha256:2e36d41f2521a89d077a3015371111910495118fde6a3e330742b2c7f2cb6aaa

Observation f35c6df4-6f89-4af8-b87d-e10f28cb3822 · outbound

This paper cites Dissecting User-Perceived Latency of On-Device E2E Speech Recognition,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Dissecting User-Perceived Latency of On-Device E2E Speech Recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.751414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.570629Z digest=sha256:1c593711773a02e7e77b17b5dbcaddd76528789e7fff71e5e9612b6f9393f8cb

Observation 777baf3b-7a0b-4fa2-854d-62c7daa4eb1b · outbound

This paper cites Dialogue act modeling for automatic tagging and recognition of conversational speech,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Dialogue act modeling for automatic tagging and recognition of conversational speech,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:22:18.575241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:22:18.575241Z digest=sha256:6c7a77dfa7802608b3ab5da8aa4f774eb0b171c83a7915a9b5dfb5a86564f6b8

Observation d2b41aee-662a-45d4-96a8-f259e2818f33 · outbound

This paper cites Metrics for polyphonic sound event detection,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Metrics for polyphonic sound event detection,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:22:18.580035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:22:18.580035Z digest=sha256:d29c30e70b20c8b93643925e0c6f4ad7338eece596327554891c113ad18bf859

Observation 4cfcf93e-1fca-4971-a299-06f6deaa6940 · outbound

This paper cites The third DIHARD di- arization challenge,.

Improving endpoint detection in end-to-end streaming ASR for conversational speech The third DIHARD di- arization challenge,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:22:18.711766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.584812Z digest=sha256:43de9b6a9ba3fa70f630682b1c91305009041154ca35b0c273ed933a4b1a9f59

Pith citing papers

Observation 4d098c91-0aea-43bc-bafb-ae6b816383b0 · inbound

Improving endpoint detection in end-to-end streaming ASR for conversational speech cites this paper.

Improving endpoint detection in end-to-end streaming ASR for conversational speech Improving endpoint detection in end-to-end streaming ASR for conversational speech

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-15T20:22:18.694131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-15T20:22:18.441360Z digest=sha256:83a1c9c19ba28135e1b19e1e36d328b2b183900de5bf2e8a60ed4691640c3f05