Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T06:02:07.932691Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 3 of 3 outbound references and 1 inbound Pith citation observation for arXiv:2605.29209.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-29T06:02:07.932691Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T10:13:57.787394Z
A source-named dated measurement, never combined with another source.
Source: cited_works
3 of 3 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5f8f2a8e-d722-4b33-879c-17de4a9412ca · outbound
The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 540b4a8b-dd47-466f-b4b3-ec4f989f556b · outbound
The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models Step-audio-r1 technical report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dc72c124-18db-419a-80a7-5dd153362632 · outbound
The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models Tokenizer Backbone.Both the encoder and decoder utilize a robust Transformer architecture: • Encoder:32 layers, 20 attention heads, 1280 hidden dimension, and 5120 linear units
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92418081-62ad-4080-9924-bd5911203b0b · inbound
ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition The WER Trap: Shattering the Illusion of Unified Tokens in Speech Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.