Pith. sign in

Paper Citation Record · LEDGER

Speech Model Pre-training for End-to-End Spoken Language Understanding

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:1904.03670.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1904.03670 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:21:35.934460Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T15:47:23.273387Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4baa1ecd-4333-4fac-91df-0a4ca85a36e4 · inbound

An Overview and Discussion of the Suitability of Existing Speech Datasets to Train Machine Learning Models for Collective Problem Solving cites this paper.

An Overview and Discussion of the Suitability of Existing Speech Datasets to Train Machine Learning Models for Collective Problem Solving Speech Model Pre-training for End-to-End Spoken Language Understanding

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-11T04:44:40.867977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:44:40.867977Z digest=sha256:6b1cba289edf69415ef258a5c06f2fc487c6ed4040f682951842fc19c6b6eed1

Observation 08bbda33-2a61-404b-9116-d7ce1ba5c64b · inbound

The ICME 2025 Audio Encoder Capability Challenge cites this paper.

The ICME 2025 Audio Encoder Capability Challenge Speech Model Pre-training for End-to-End Spoken Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T14:27:09.106687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:27:09.106687Z digest=sha256:bc8eac31f406027a8dabe54579f011d7e56028361a2e948fd562e49ad657cfbf

Observation bee27332-0ace-4efd-81be-dd65522065c7 · inbound

QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding cites this paper.

QUADS: QUAntized Distillation Framework for Efficient Speech Language Understanding Speech Model Pre-training for End-to-End Spoken Language Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T20:21:35.934460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:21:35.934460Z digest=sha256:fe9a65441bd02c0cce6f5b2e649faf78f51a8103700b9a175fdc6d090137de88

Observation d5e70fa5-877b-462d-8c31-0932f4dfd7eb · inbound

X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance cites this paper.

X-ARES: A Comprehensive Framework for Assessing Audio Encoder Performance Speech Model Pre-training for End-to-End Spoken Language Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:05:14.197126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:05:14.197126Z digest=sha256:8837a91cb4da50c3cfb466262c16f6cee7dd77517ce0ad92995ed0678971e9e8

Observation 5cc71a69-6342-4500-8b2d-4106f84f54fa · inbound

Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving cites this paper.

Enhancing Generalization of Speech Large Language Models with Multi-Task Behavior Imitation and Speech-Text Interleaving Speech Model Pre-training for End-to-End Spoken Language Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:30:33.562288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:30:33.562288Z digest=sha256:02a034db7ef364089bb11cd018ec0d2c36d71163212eabe15b33cda0bb5e7975

Observation c2240fbd-62e5-4a08-9b63-b53fdd165702 · inbound

Continual Speech Learning with Fused Speech Features cites this paper.

Continual Speech Learning with Fused Speech Features Speech Model Pre-training for End-to-End Spoken Language Understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T11:46:59.294977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:46:59.294977Z digest=sha256:1288e6ed1f5ed64ff51979a8dc3ed05bd24c3d3b2b811085cb7d23da08a9805d

Observation 9107ecd9-b34d-4846-b3be-055b6f96949a · inbound

findsylls: A Language-Agnostic Toolkit for Syllable-Level Speech Tokenization and Embedding cites this paper.

findsylls: A Language-Agnostic Toolkit for Syllable-Level Speech Tokenization and Embedding Speech Model Pre-training for End-to-End Spoken Language Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T17:37:12.659819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T17:37:12.659819Z digest=sha256:6880b35dd5d5dfa0387e05b6986dffbe0c0f6a8f6d6bcf43f60cfc46ac0e3e1e

Observation b66e635a-b0a0-4956-a1ae-7a1e5a142bb2 · inbound

WQ-Fusion: Dynamic Gated Attention for Cross-Domain Audio Representation cites this paper.

WQ-Fusion: Dynamic Gated Attention for Cross-Domain Audio Representation Speech Model Pre-training for End-to-End Spoken Language Understanding

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-04T14:39:57.212554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T03:21:39.842959Z digest=sha256:e1384983a9298e10cb9e268c819f55ff1c7bb962b5d6687ca27866f43dfe6826

Observation 741fbe33-76d3-4287-8649-b56b6537444b · inbound

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks cites this paper.

Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks Speech Model Pre-training for End-to-End Spoken Language Understanding

Reference 249

Resolution
metadata mismatch
local_arxiv, observed 2026-07-10T15:47:23.274542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-07-10T15:38:58.361411Z digest=sha256:2d04702ed12172679a31fc2f32006019c5f24e4a9f45ecc8cb8872319f8e181c

Observation c69cc7ce-597c-4dab-8348-63faafe4144f · inbound

VAD to the Bone: Ultra-Tiny Speech Activity Detection for Edge Deployment cites this paper.

VAD to the Bone: Ultra-Tiny Speech Activity Detection for Edge Deployment Speech Model Pre-training for End-to-End Spoken Language Understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T01:19:07.656877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:19:07.656877Z digest=sha256:76523bcac296ff5a6f196d95c8cf34920fee52d5d010bcdc214b43575743022f

Observation f248c76e-2494-4794-a3f7-03a9e0cb401f · inbound

Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning cites this paper.

Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning Speech Model Pre-training for End-to-End Spoken Language Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T08:50:37.709583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:50:37.709583Z digest=sha256:7e434d1eca992c5e012d6f0111eae223e243882e35cb9c0b9086bb2e99a4cc64