Pith. sign in

Paper Citation Record · LEDGER

StepAudio 2.5 Technical Report

As of 6 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 2 inbound Pith citation observations for arXiv:2605.23463.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.23463 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-25T02:52:22.610397Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T10:13:56.110991Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T17:18:44.075741Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact20
  • verified fuzzy18
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 086e75aa-a6ef-417f-b944-038c36ad785f · outbound

This paper cites Connectionist temporal classification.

StepAudio 2.5 Technical Report Connectionist temporal classification

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.361557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:d9d89c48591bf59d201c5672e2213523dab51865ba8c03ef979916c08abe4a20

Observation 14ae607c-95e5-44d1-bdf8-4b5aacb3f70c · outbound

This paper cites Sequence Transduction with Recurrent Neural Networks.

StepAudio 2.5 Technical Report Sequence Transduction with Recurrent Neural Networks

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.521251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:37a9a9f12ac72403b977e7dca4e70ecc098a4f9fa4da4fb4e0c55552f561b654

Observation 6eab5b32-aab8-4e2a-a3fe-7d9fdbfc6b2a · outbound

This paper cites Listen, Attend and Spell.

StepAudio 2.5 Technical Report Listen, Attend and Spell

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.509973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:68629928e5510872ef3a17001701826a49744589011269fb1fda47cd01a8c50e

Observation 5eeac4c3-be4e-46d5-b52c-f412f6d551a2 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

StepAudio 2.5 Technical Report Robust speech recognition via large-scale weak supervision

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.353701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:952c368b9420a7b406f93b85b4eca707ed33a46fe983907478be26cc01191f2e

Observation df029add-8e4b-4238-ab59-6d9334d01e84 · outbound

This paper cites VIBEVOICE-ASR technical report.

StepAudio 2.5 Technical Report VIBEVOICE-ASR technical report

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.526933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:661002dff81d0d53bcdce463ccbd158c81e67419e2f65269d4e2628d6e56e2f8

Observation 5e20d531-5ce5-478e-afeb-3a15edf3fa37 · outbound

This paper cites Fun-ASR technical report.

StepAudio 2.5 Technical Report Fun-ASR technical report

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.532217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:d62aa202d3f42d1c5f23e16eebd1d3e4bac138586988e8bb07263f9709003172

Observation 91001d88-a7f2-4930-8541-dc76486a95f5 · outbound

This paper cites Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition.

StepAudio 2.5 Technical Report Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.515698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:0d7139c4d5aec47cd02e17bcd75092841822b125aae917a8c52ca483c764a620

Observation 9fd13268-45a3-43f5-a739-bfc06b0f025e · outbound

This paper cites Qwen3-ASR Technical Report.

StepAudio 2.5 Technical Report Qwen3-ASR Technical Report

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.483222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:d38fadb54043430dcd2f2f2c14026a244750959a8576cfc9c8380c56d0ef935b

Observation ae35f11f-b45b-4570-a385-d37c6d821ce6 · outbound

This paper cites Step-Audio 2 Technical Report.

StepAudio 2.5 Technical Report Step-Audio 2 Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.487852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:90acad101f777216641f6d4aa21981653bc06241c81654924549023a1f86e049

Observation 713f12d6-bd4a-407b-8965-ba391750f8e0 · outbound

This paper cites Qwen3-Omni Technical Report.

StepAudio 2.5 Technical Report Qwen3-Omni Technical Report

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.441821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:eeee58a1e8c0d7149cf6b2752f5d9d70d0b5921b81687d2ae8d6794901e3a8cd

Observation 0953c54f-6449-4679-b7a8-a5382add346c · outbound

This paper cites Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation.

StepAudio 2.5 Technical Report Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.477459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:35f5e62a37a33096fdc85bf31590de4510ee57d23518edaeb6b6a0965709ca0d

Observation fc8064bb-97ff-403c-9dfd-eedabf2c2453 · outbound

This paper cites Salmonn: Towards generic hearing abilities for large language models.

StepAudio 2.5 Technical Report Salmonn: Towards generic hearing abilities for large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.371149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:f553bdd234997d862edca0625514c53e8bf2795c2190febb51a427aa0292b7ad

Observation 57515c3d-9c92-4e8e-9189-36d3eb900b8a · outbound

This paper cites Audiolm: a language modeling approach to audio generation.IEEE/ACM transactions on audio, speech, and language processing, 31:2523–2533.

StepAudio 2.5 Technical Report Audiolm: a language modeling approach to audio generation.IEEE/ACM transactions on audio, speech, and language processing, 31:2523–2533

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.374738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:b390edc310cbb359b625600db253febe7605a8d41d1ee1ea17a068b236125739

Observation 350987e5-fa75-45d3-8571-0bcdb7a12996 · outbound

This paper cites Recent advances in speech language models: A survey.

StepAudio 2.5 Technical Report Recent advances in speech language models: A survey

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.379842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:698af2c071d71777d632cb98437953860abc0be9c952586482bd8bd20b88910b

Observation c59d0276-31d1-4f4f-b648-9517b90875b5 · outbound

This paper cites Paralinguistics-aware speech-empowered large language models for natural conversation.Advances in Neural Information Processing Systems, 37:131072–131103.

StepAudio 2.5 Technical Report Paralinguistics-aware speech-empowered large language models for natural conversation.Advances in Neural Information Processing Systems, 37:131072–131103

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.383483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:0f72a595569e3ff3d2779fea85e5b6ce2c1ac9c0ca890313d6c7392de5a6b2a3

Observation 0caf796e-e190-4ba0-b488-fdd8162c462c · outbound

This paper cites Freeze-omni: A smart and low latency speech-to-speech dialogue model with frozen LLM.

StepAudio 2.5 Technical Report Freeze-omni: A smart and low latency speech-to-speech dialogue model with frozen LLM

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.389815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:e37d21524ef8a932e8e7caa0c8eb637ae53e43f8880b867a98fe62f305982bb3

Observation 2c452943-18a0-4a86-b196-e4b4c9f47119 · outbound

This paper cites Depflow: Disentangled speech generation to mitigate semantic bias in depression detection.

StepAudio 2.5 Technical Report Depflow: Disentangled speech generation to mitigate semantic bias in depression detection

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.431713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:7a06c6f775089ca939e076d6fe2354b18a8868d99011bf20c94f288c7f7ecb0c

Observation 9b34497f-a2a0-43e7-a002-f32ab6dbe7ee · outbound

This paper cites A new approach to extract fetal electrocardiogram using affine combination of adaptive filters.

StepAudio 2.5 Technical Report A new approach to extract fetal electrocardiogram using affine combination of adaptive filters

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.367549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:bb8b526c38850ea011188aa5e6fc1a88238b6f7f1ad6870ffc44a21178672be9

Observation e2be10c6-91b2-40e5-abf2-c6973cc89d8b · outbound

This paper cites Multi-bench: A multi-turn interactive benchmark for assessing emotional intelligence ability of spoken dialogue models.

StepAudio 2.5 Technical Report Multi-bench: A multi-turn interactive benchmark for assessing emotional intelligence ability of spoken dialogue models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.436892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:3c511c6de6fe995c04e8d948d0889ae0bb42ead54883b7aecce6d4bdf3665442

Observation be94f0cd-1a89-49c5-9cc2-8a8b361e35e7 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

StepAudio 2.5 Technical Report Gemini: A Family of Highly Capable Multimodal Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.463143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:cdcd8bc2db6e6f39fb4ac3c2dc9405078dd2a5d392e716cb24b590820178ceb9

Observation a0b65575-9b51-44bd-9706-5a1487dd1153 · outbound

This paper cites Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models.

StepAudio 2.5 Technical Report Mind-Paced Speaking: A Dual-Brain Approach to Real-Time Reasoning in Spoken Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.472393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:4f2b4d1bf52649fda749d7fe38cc42f0eec3ecdf00cba7cd9a8d53d596c824cc

Observation 3f84023c-a4ee-4d5f-bc42-95785eff768a · outbound

This paper cites Chronological thinking in full-duplex spoken dialogue language models.

StepAudio 2.5 Technical Report Chronological thinking in full-duplex spoken dialogue language models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.505381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:66fd683bb9b3a42b8925b04cd506d88adda1baa5fe1b41e1670977ed9c86fa3f

Observation 63082297-3c5c-4db9-aca2-327d8a2db705 · outbound

This paper cites Duplexsla: A full-duplex spoken language model with synchronized speech, language, and action.

StepAudio 2.5 Technical Report Duplexsla: A full-duplex spoken language model with synchronized speech, language, and action

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.386641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:49cc0550dc5ad7d33c3b3371d11b616f1bf9860903aa3f3ec85e9bc749c02fa9

Observation 20c44ea9-682f-4e47-ab7b-0581c94a5b9b · outbound

This paper cites Mamba in speech: Towards an alternative to self-attention.IEEE Transactions on Audio, Speech and Language Processing.

StepAudio 2.5 Technical Report Mamba in speech: Towards an alternative to self-attention.IEEE Transactions on Audio, Speech and Language Processing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.392860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:30476b1a023924433f2354723b6dbbb7bb3a79d8c905aa1f7e9a55acd9de0d78

Observation df5d89ca-d3a6-4fe0-b631-566a1f8999a9 · outbound

This paper cites Code-switching speech recognition under the lens: Model-and data-centric perspectives.IEEE Transactions on Audio, Speech and Language Processing.

StepAudio 2.5 Technical Report Code-switching speech recognition under the lens: Model-and data-centric perspectives.IEEE Transactions on Audio, Speech and Language Processing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.396848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:d58b322e4f0ede215bbeb0a4906199086954bc78d769709f8242c1593f01aeb6

Observation 413f7c81-0419-41ee-9e47-258ff55a93e6 · outbound

This paper cites Step-audio-r1 technical report.

StepAudio 2.5 Technical Report Step-audio-r1 technical report

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.500260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:8263ed65b4d6e7012501dc2fbacec2325a9a590ae8e3a35b2c9170636fafab80

Observation 61a5a62a-d7e4-43fe-b7da-eb8aaabcd02b · outbound

This paper cites Step-Audio-R1.5 Technical Report.

StepAudio 2.5 Technical Report Step-Audio-R1.5 Technical Report

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.494617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:26fe5466774d8c44ba534e7c80434d0f39dfa5ed065ab873f8d48b0c139f5da8

Observation 6bef68ea-45ae-41f8-adba-1faf1a74e168 · outbound

This paper cites Park, William Chan, Yu Zhang, et al.

StepAudio 2.5 Technical Report Park, William Chan, Yu Zhang, et al

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.402758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:10f18c17b10c60c1789aa3e6b70b4bc876c60a9f5be62e6a5b0bbdffceed6066

Observation 1a6e80ea-7904-4bf1-ba00-a51938a47ef3 · outbound

This paper cites an unresolved cited work.

StepAudio 2.5 Technical Report Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-05-25T02:56:35.406529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:8d004653391c3e69c7cbe758b7e5fb2f2bed1864cb5a05fe771949f1d672d271

Observation 697f37fc-63c2-408f-8eb4-c9cef2fdab7e · outbound

This paper cites AIShell-1: An open-source mandarin speech corpus and a speech recognition baseline.

StepAudio 2.5 Technical Report AIShell-1: An open-source mandarin speech corpus and a speech recognition baseline

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.409728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:3c939bb47b307a1ff6dc7cbc14343c4b4bfe24623cc3b9f519f79874582d5b19

Observation 01cf541e-269d-45f9-9451-c73612401bd9 · outbound

This paper cites AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale.

StepAudio 2.5 Technical Report AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.451910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:bf7c0552fb03a4db4ffb48b0351713bf61e24153eb402dcc24948786137e825d

Observation 8d2d211f-feb6-4434-83a8-468b733a68fc · outbound

This paper cites WenetSpeech: A 10000+ hours multi-domain mandarin corpus for speech recognition.

StepAudio 2.5 Technical Report WenetSpeech: A 10000+ hours multi-domain mandarin corpus for speech recognition

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.422611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:ff7233401431b4a25a1af86908833848854b42f3fd74a21d2ba62c59363c9dd8

Observation 49a49c25-9c27-466e-ae3e-3b83d2481bb9 · outbound

This paper cites FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech.

StepAudio 2.5 Technical Report FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.447086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:98c679476edb3e4739c5409ff07f8dd855a8baf374fdd8ab6825846d1fd39a49

Observation 01041449-4711-47c4-8cb1-e0f0585030bc · outbound

This paper cites LibriSpeech: An ASR corpus based on public domain audio books.

StepAudio 2.5 Technical Report LibriSpeech: An ASR corpus based on public domain audio books

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.425885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:dc4a78ca696ff22a8212bf1f76469323521247ab5f1aa488c159851c49c665d5

Observation 8f1435fe-f4a1-4c63-b748-925370e9a75e · outbound

This paper cites Common voice: A massively-multilingual speech corpus.

StepAudio 2.5 Technical Report Common voice: A massively-multilingual speech corpus

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.419215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:c9e616f6351127509136366886ce58af3ead10309af33a3b1e3d0128297094e6

Observation 6c171711-18e3-4de3-9885-1f7b8eac9738 · outbound

This paper cites V oxpopuli-cleaned-aa: Cleaned ground truth transcripts for voxpopuli english test set.

StepAudio 2.5 Technical Report V oxpopuli-cleaned-aa: Cleaned ground truth transcripts for voxpopuli english test set

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.415860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:2349d4fce244dc9b019fedca67526396476c9adf54de97088dcc40e94acd1063

Observation 9d1bfb95-8e08-47a5-b14c-f7e005da005c · outbound

This paper cites Earnings22-cleaned-aa: Cleaned ground truth transcripts for earnings22 english test set.

StepAudio 2.5 Technical Report Earnings22-cleaned-aa: Cleaned ground truth transcripts for earnings22 english test set

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T02:56:35.412796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:fe7eae167149fa617a177a79eb707e01fcb6c0e664cdbc0bd3be455cdf182826

Observation 4534e43b-4358-4eec-be62-10f5a597de92 · outbound

This paper cites Step-audio-editx technical report.

StepAudio 2.5 Technical Report Step-audio-editx technical report

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-25T02:55:16.458036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:10f6e1310176f399917c94542dbf7b2c5c54fdf00d70c191e8ef180971979fad

Observation ae24bc43-fe61-4d31-a7a5-2443a782c1a9 · outbound

This paper cites Proximal Policy Optimization Algorithms.

StepAudio 2.5 Technical Report Proximal Policy Optimization Algorithms

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:55:16.467568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:52:22.610397Z digest=sha256:955b6b385ce02741183d556bd6fb89371e7b7ab4d5bff7f5e31471e9f2ec65c8

Pith citing papers

Observation 789faeb4-f885-46e4-a64f-704563c7b88e · inbound

Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech cites this paper.

Improving Code-Switching ASR with Code-Mixing Guided Synthetic Speech StepAudio 2.5 Technical Report

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:18:44.077038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T04:19:13.689689Z digest=sha256:ee3026dd5b50f0c35cdafdea6976739fdac3d97d60a470a2d1f88a4d5d298be1

Observation a05dba3a-3f4d-45fe-ad13-fa4c033e29f2 · inbound

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition cites this paper.

ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition StepAudio 2.5 Technical Report

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T10:13:56.110991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T10:13:56.110991Z digest=sha256:682b0cb92719fb5720372acea7a819f95609509892dfcde34c9b523904716350