Pith. sign in

Paper Citation Record · LEDGER

SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2110.07205.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2110.07205 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:58:41.673699Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T00:56:25.108341Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 1a9a4096-7be8-487e-bd11-85d4f50bebe4 · inbound

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models cites this paper.

Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:57:28.714811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T18:57:28.666194Z digest=sha256:4ec99957814b54eb4e0f0af16f8cd570433f97ff49a744d97f389720b3c61250

Observation 8f19f970-ef08-4992-94ed-0e9b227a247c · inbound

Qwen2-Audio Technical Report cites this paper.

Qwen2-Audio Technical Report SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T02:14:45.400835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T02:14:45.371564Z digest=sha256:75fb7a26f22168f40308f4e83b9cdba4e3003cf79bdb0b269765406868cf8863

Observation 4e0f11e2-22a8-4db6-94d5-26769db66329 · inbound

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment cites this paper.

TESU-LLM: Training Speech-LLMs Without Speech via Unified Encoder Alignment SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:41.673699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:58:41.673699Z digest=sha256:198814dd67edceff0775639f71c13dca28994443051c67e3bc8043d6f90b41bc

Observation 08c2ee89-05c8-4112-bbd2-b3cc7cab53c4 · inbound

Your Spending Needs Attention: Modeling Financial Habits with Transformers cites this paper.

Your Spending Needs Attention: Modeling Financial Habits with Transformers SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T10:59:01.261072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:59:01.261072Z digest=sha256:6fc5ad262c33399df4edd8f5c970c16467f8996617f1fb7a041487364ec299ac

Observation 4cc58584-f19d-4b4d-8286-78ba2c0b84b0 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:29.783850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:29.783850Z digest=sha256:f7c800a35cc6cf2bf0037c624a7bae786ab4aa62706822dcd7119d6e3c719291

Observation d9642255-09a9-46d6-807c-b1e781165aed · inbound

Evaluating Generalization and Robustness in Russian Anti-Spoofing: The RuASD Initiative cites this paper.

Evaluating Generalization and Robustness in Russian Anti-Spoofing: The RuASD Initiative SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:46:12.434663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T02:27:05.776290Z digest=sha256:ea0837f4046311c5bb84a4cbd82f37c10f37e4ca72b4b73a53039addd031f316

Observation 731d1eee-03d8-462a-b9b5-bc1eae320375 · inbound

MOSS-Audio Technical Report cites this paper.

MOSS-Audio Technical Report SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:56:25.110474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:ee32c7820008bc114d5f1f31745dba4bf2f70a16b9f1b11ccbb8340fe7ff8548