Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-25T02:52:22.610397Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 2 inbound Pith citation observations for arXiv:2605.23463.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-25T02:52:22.610397Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T10:13:56.110991Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T17:18:44.075741Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 086e75aa-a6ef-417f-b944-038c36ad785f · outbound
StepAudio 2.5 Technical Report Connectionist temporal classification
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 14ae607c-95e5-44d1-bdf8-4b5aacb3f70c · outbound
StepAudio 2.5 Technical Report Sequence Transduction with Recurrent Neural Networks
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6eab5b32-aab8-4e2a-a3fe-7d9fdbfc6b2a · outbound
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5eeac4c3-be4e-46d5-b52c-f412f6d551a2 · outbound
StepAudio 2.5 Technical Report Robust speech recognition via large-scale weak supervision
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation df029add-8e4b-4238-ab59-6d9334d01e84 · outbound
StepAudio 2.5 Technical Report VIBEVOICE-ASR technical report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5e20d531-5ce5-478e-afeb-3a15edf3fa37 · outbound
StepAudio 2.5 Technical Report Fun-ASR technical report
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 91001d88-a7f2-4930-8541-dc76486a95f5 · outbound
StepAudio 2.5 Technical Report Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9fd13268-45a3-43f5-a739-bfc06b0f025e · outbound
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ae35f11f-b45b-4570-a385-d37c6d821ce6 · outbound
StepAudio 2.5 Technical Report Step-Audio 2 Technical Report
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 713f12d6-bd4a-407b-8965-ba391750f8e0 · outbound
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0953c54f-6449-4679-b7a8-a5382add346c · outbound
StepAudio 2.5 Technical Report Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fc8064bb-97ff-403c-9dfd-eedabf2c2453 · outbound
StepAudio 2.5 Technical Report Salmonn: Towards generic hearing abilities for large language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 57515c3d-9c92-4e8e-9189-36d3eb900b8a · outbound
StepAudio 2.5 Technical Report Audiolm: a language modeling approach to audio generation.IEEE/ACM transactions on audio, speech, and language processing, 31:2523–2533
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 350987e5-fa75-45d3-8571-0bcdb7a12996 · outbound
StepAudio 2.5 Technical Report Recent advances in speech language models: A survey
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c59d0276-31d1-4f4f-b648-9517b90875b5 · outbound
StepAudio 2.5 Technical Report Paralinguistics-aware speech-empowered large language models for natural conversation.Advances in Neural Information Processing Systems, 37:131072–131103
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0caf796e-e190-4ba0-b488-fdd8162c462c · outbound
StepAudio 2.5 Technical Report Freeze-omni: A smart and low latency speech-to-speech dialogue model with frozen LLM
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2c452943-18a0-4a86-b196-e4b4c9f47119 · outbound
StepAudio 2.5 Technical Report Depflow: Disentangled speech generation to mitigate semantic bias in depression detection
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9b34497f-a2a0-43e7-a002-f32ab6dbe7ee · outbound
StepAudio 2.5 Technical Report A new approach to extract fetal electrocardiogram using affine combination of adaptive filters
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e2be10c6-91b2-40e5-abf2-c6973cc89d8b · outbound
StepAudio 2.5 Technical Report Multi-bench: A multi-turn interactive benchmark for assessing emotional intelligence ability of spoken dialogue models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation be94f0cd-1a89-49c5-9cc2-8a8b361e35e7 · outbound
StepAudio 2.5 Technical Report Gemini: A Family of Highly Capable Multimodal Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a0b65575-9b51-44bd-9706-5a1487dd1153 · outbound
StepAudio 2.5 Technical Report Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3f84023c-a4ee-4d5f-bc42-95785eff768a · outbound
StepAudio 2.5 Technical Report Chronological thinking in full-duplex spoken dialogue language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 63082297-3c5c-4db9-aca2-327d8a2db705 · outbound
StepAudio 2.5 Technical Report Duplexsla: A full-duplex spoken language model with synchronized speech, language, and action
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 20c44ea9-682f-4e47-ab7b-0581c94a5b9b · outbound
StepAudio 2.5 Technical Report Mamba in speech: Towards an alternative to self-attention.IEEE Transactions on Audio, Speech and Language Processing
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation df5d89ca-d3a6-4fe0-b631-566a1f8999a9 · outbound
StepAudio 2.5 Technical Report Code-switching speech recognition under the lens: Model-and data-centric perspectives.IEEE Transactions on Audio, Speech and Language Processing
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 413f7c81-0419-41ee-9e47-258ff55a93e6 · outbound
StepAudio 2.5 Technical Report Step-audio-r1 technical report
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 61a5a62a-d7e4-43fe-b7da-eb8aaabcd02b · outbound
StepAudio 2.5 Technical Report Step-Audio-R1.5 Technical Report
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6bef68ea-45ae-41f8-adba-1faf1a74e168 · outbound
StepAudio 2.5 Technical Report Park, William Chan, Yu Zhang, et al
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 1a6e80ea-7904-4bf1-ba00-a51938a47ef3 · outbound
StepAudio 2.5 Technical Report Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 697f37fc-63c2-408f-8eb4-c9cef2fdab7e · outbound
StepAudio 2.5 Technical Report AIShell-1: An open-source mandarin speech corpus and a speech recognition baseline
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 01cf541e-269d-45f9-9451-c73612401bd9 · outbound
StepAudio 2.5 Technical Report AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8d2d211f-feb6-4434-83a8-468b733a68fc · outbound
StepAudio 2.5 Technical Report WenetSpeech: A 10000+ hours multi-domain mandarin corpus for speech recognition
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 49a49c25-9c27-466e-ae3e-3b83d2481bb9 · outbound
StepAudio 2.5 Technical Report FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 01041449-4711-47c4-8cb1-e0f0585030bc · outbound
StepAudio 2.5 Technical Report LibriSpeech: An ASR corpus based on public domain audio books
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8f1435fe-f4a1-4c63-b748-925370e9a75e · outbound
StepAudio 2.5 Technical Report Common voice: A massively-multilingual speech corpus
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6c171711-18e3-4de3-9885-1f7b8eac9738 · outbound
StepAudio 2.5 Technical Report V oxpopuli-cleaned-aa: Cleaned ground truth transcripts for voxpopuli english test set
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9d1bfb95-8e08-47a5-b14c-f7e005da005c · outbound
StepAudio 2.5 Technical Report Earnings22-cleaned-aa: Cleaned ground truth transcripts for earnings22 english test set
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4534e43b-4358-4eec-be62-10f5a597de92 · outbound
StepAudio 2.5 Technical Report Step-audio-editx technical report
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation ae24bc43-fe61-4d31-a7a5-2443a782c1a9 · outbound
StepAudio 2.5 Technical Report Proximal Policy Optimization Algorithms
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 789faeb4-f885-46e4-a64f-704563c7b88e · inbound
Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech StepAudio 2.5 Technical Report
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a05dba3a-3f4d-45fe-ad13-fa4c033e29f2 · inbound
ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition StepAudio 2.5 Technical Report
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.