Pith. sign in

Paper Citation Record · LEDGER

Word Level Timestamp Generation for Automatic Speech Recognition and Translation

As of 5 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2505.15646.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.15646 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:29:22.077126Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-07T14:53:55.835313Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c7367329-b074-4690-8396-c8d89d527ccf · inbound

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions cites this paper.

In-Sync: Adaptation of Speech Aware Large Language Models for ASR with Word Level Timestamp Predictions Word Level Timestamp Generation for Automatic Speech Recognition and Translation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.723435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:25:50.524448Z digest=sha256:c83456aa3f1ddc3db96c37f0a60dae5c69e6fe80af348d89269fc16b5f9bc950

Observation 47effc7e-584b-43b7-bf48-3762a318aa6a · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue Word Level Timestamp Generation for Automatic Speech Recognition and Translation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.101650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:3406c8f7ee044e8e638a89e285a6ef05c640bb3c6bfc10d02fe0becedf5daa92

Observation e2cad170-3d16-4002-9f89-620863b6aef0 · inbound

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing cites this paper.

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing Word Level Timestamp Generation for Automatic Speech Recognition and Translation

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-07T14:53:55.837853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-07T14:52:35.127533Z digest=sha256:291fb1bbf8b536a3d4801ec2f43c9a9d5f9a1506e6beb189888cb1cc5f33dee4

Observation f45fb52e-531b-4ab7-bdae-ec25e6d94d36 · inbound

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing cites this paper.

REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing Word Level Timestamp Generation for Automatic Speech Recognition and Translation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T08:34:02.696146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:34:02.696146Z digest=sha256:398a69e54325ef977d96593e2158d6bd050a546c835463ffb903d681af94c484

Observation 0f40e806-b1e9-419f-b475-bf27c4ce6c10 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Word Level Timestamp Generation for Automatic Speech Recognition and Translation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:22.077126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:22.077126Z digest=sha256:f8cabf7f9cd7f4520761870a93e68668cf7f4eea6e34f974067f685a814a7e36