Pith. sign in

Paper Citation Record · LEDGER

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

As of 2 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 35 inbound Pith citation observations for arXiv:2502.11946.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.11946 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-18T13:39:48.225482Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:43:34.049064Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T02:49:25.012133Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact21
  • verified fuzzy17
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 140d72a0-7b80-44a1-a3cd-11115e1b2599 · outbound

This paper cites The Method of Paired Comparisons , author=.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction The Method of Paired Comparisons , author=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.451611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:b17a328e79f3507ead80432eb6f8564f1b5cd24c105167f05c8974f404b2886e

Observation 59504f3a-02fa-4189-8681-69058398f4db · outbound

This paper cites International conference on machine learning , pages=.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction International conference on machine learning , pages=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.418058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:7cb89a4f6b5d355cb0dec254a91292c756386bd1aba63eed9167917232d1e919

Observation db7516ef-eed4-463e-9032-585405001c8b · outbound

This paper cites ICASSP 23 , year=.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction ICASSP 23 , year=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.421909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:2a2cb416811bdda6404b35df233d89e1a313562c383d3820cbd88479e2a9e844

Observation 7999e265-947b-48c0-8b19-8ee96d91f9a3 · outbound

This paper cites 2017 20th conference of the oriental chapter of the international coordinating committee on speech databases and speech I/O systems and assessment (O-COCOSDA) , pages=.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction 2017 20th conference of the oriental chapter of the international coordinating committee on speech databases and speech I/O systems and assessment (O-COCOSDA) , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.426351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:f82f06452afa8d05eb9574f48eea01f1105e711816ab1b6c59008f08cc2c2292

Observation caefa079-983b-414a-b880-a257b96ff91e · outbound

This paper cites ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) , pages=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.430746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:9f2da733a529b9e69540436537ad55a67b8f84fbcc081418e569ecfb72f05237

Observation 09d7c2d1-71a6-44f1-ba0d-78e57b6748b4 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.434662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:9d26e91b1edf73a3db09dc17508b4e2731215448eaba70577ee3e7da389f1407

Observation e6957376-f545-4975-a7a1-1d0188dc2d06 · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.438169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:45d9647492783822071a8388ec5a2c27ebd0ec9467f3985bb6f698365c5bbff6

Observation 4c593f38-50e3-4bd9-b64d-ea8488ed5415 · outbound

This paper cites 2024 , howpublished =.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction 2024 , howpublished =

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.441489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:40c53a519752b4b57358d850ceb7d3e3a945edb61dd2edfb20f4b18f8d438900

Observation 11452c23-5220-4e88-a6f6-ef9c92a2b030 · outbound

This paper cites 2024 , howpublished =.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction 2024 , howpublished =

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.444793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:b5136eed2f3382c3e42a187211a61ae165556d6a9fba9e1e6a1f2e9f867931b3

Observation 5af1c320-77b5-4d62-9907-d3a655eb9244 · outbound

This paper cites 2024 , howpublished =.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction 2024 , howpublished =

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.448333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:9cbca8fe9a529f5576b1e6d3009dd3b8880f8214b4849a83b2b07828125acb34

Observation 0024d1c8-e6d0-4bcc-9ca5-e0bd5e3fc405 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:39:48.408958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:69a12580f5a2a7d54cec6ad69eb243a31388137cd23e703279c847b6357f3d78

Observation e92b9024-1312-4650-859e-7071f8a5f753 · outbound

This paper cites \ Terry, M E.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction \ Terry, M E

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.455010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:c726ce03801f11afb9fce4888a155acb526f43589a57b9ff0bcb6b0f1e9f6261

Observation 39d98dc5-3309-4596-8421-75a62949bf42 · outbound

This paper cites APACrefauthors \ 2024.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction APACrefauthors \ 2024

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.458573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:e23c5ea694d7dc1fc52d9262fd6293e051fb40ea516abf9509b8549dc430fd68

Observation b8366d02-27ad-4e39-8c1c-979a30e28278 · outbound

This paper cites MinMo: A Multimodal Large Language Model for Seamless Voice Interaction.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.414015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:0282031ef6ce46d479e90be3efaf8bab762a5776bd5b2b97c5da27fb2ab43e0f

Observation 852657ce-7c19-473e-95bf-10933b2c1e2d · outbound

This paper cites An Enhanced Res2Net with Local and Global Feature Fusion for Speaker Verification.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction An Enhanced Res2Net with Local and Global Feature Fusion for Speaker Verification

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.277460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:d19565f752612f54ab6e75103ea44817fbe3b316ad582fe0d0e05407c4225ed6

Observation 5161f6ab-9758-4d17-9c3e-4ba99d038d7a · outbound

This paper cites Qwen2-Audio Technical Report.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Qwen2-Audio Technical Report

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:39:48.283826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:1cdb4ef502e7e058b565f94a5835796a1d08430aa3c5840b2467bcbc993b98bd

Observation 62792d3a-d3f9-48b5-badd-2459f82d3802 · outbound

This paper cites SpeechVerse: A Large-scale Generalizable Audio Language Model.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction SpeechVerse: A Large-scale Generalizable Audio Language Model

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.291681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:4df4fe9f0be3dedf5160d5d05892554873d307a7ca2aaabe708eceaede903151

Observation e650777f-ee34-4217-a15e-44188e6de56f · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Moshi: a speech-text foundation model for real-time dialogue

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:39:48.298794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:dc2228b111bee3e8cdc2938d0e2249f7121d96b52f1538f07b5cd07c14391151

Observation d59f2916-ea4e-4f3d-a6e4-fbe15deea9f2 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:39:48.305578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:d9b06de9972cb9f8c3a472220fd7dad97c3de21a1536fe7a77f8e3e97b0b88a1

Observation 353f164c-5839-4f53-9a37-9d6825c69551 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:39:48.329726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:8783e8e24d73c8ac412327c7be9f042f5a2b70fdf872df3f45a4a7537f36a225

Observation 1625c794-0f28-4b4f-9768-e6e77f92aca3 · outbound

This paper cites The Llama 3 Herd of Models.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction The Llama 3 Herd of Models

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:39:48.336753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:18aaa964096ae514b91efc8af8f8b8da566aa93b3eb88d76f0b41ba4fc7ef7ec

Observation 2d3ae2fc-a0b0-4740-8024-10f7dbc5bf8e · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.342720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:c6ceb3ee032ffce6d9c7cc23890b21a7d0e0e40dad080671f2e4452adaa806d8

Observation 7d3c2227-45f5-494e-b5d2-7890fd337c21 · outbound

This paper cites LUCY: Linguistic Understanding and Control Yielding Early Stage of Her.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction LUCY: Linguistic Understanding and Control Yielding Early Stage of Her

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.348498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:313c74d0c8a0aecae2a4f31e8d0ca89e5e33452618c82fa400ca7b7f35e17dc7

Observation 828da4f7-53b1-4efe-aa1f-31d64759fe8c · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.353222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:8ef320974642dc96be393830113cf00285f4559227b111749782ecc06324adb1

Observation eb985253-43d0-4221-9e6f-15dd2dab079f · outbound

This paper cites WavLLM: Towards Robust and Adaptive Speech Large Language Model.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.359286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:5e6769ff7a1d5d3f2a3dda5706ba35425ce9033d6fbad7b742137cfdf7d8d3b8

Observation 3d6e2987-5ff5-4967-b155-b0902bf75b4d · outbound

This paper cites an unresolved cited work.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-18T13:39:48.479948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:665f8fdd5fda0e96d0f02d4806ba7a4d4d9f95c6e95d6a411775d13f898d4ba1

Observation 1c1befd4-35e1-49c3-86e6-7597a76b290d · outbound

This paper cites GPT-4o System Card.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction GPT-4o System Card

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:39:48.367331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:a58c8120404a1e668df3289f3dfb65d048a5f192e6cad8df060533e4adff7b19

Observation d578991a-7be3-4bf3-a85b-9d47a8c05ba0 · outbound

This paper cites Cao, Y.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Cao, Y

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.464198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:ac667f285d1e8ed0b01b471ed0d4a6bfe3b7d570915c6c66021b5681da33e824

Observation 72d6ebc0-f803-40cb-830e-3e6e6ebe8195 · outbound

This paper cites ARCON: Advancing Auto-Regressive Continuation for Driving Videos.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction ARCON: Advancing Auto-Regressive Continuation for Driving Videos

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.372483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:20e57555c9a0d1c9ba4e61f17c4386bffc3d1fb0b40bda2dda33506d0c438f8b

Observation 1fe309e4-2f44-4648-913b-fde8697162f6 · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Spirit LM: Interleaved Spoken and Written Language Model

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.378269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:288dac4b2e60c7c9ebc68418826ed46a91b3b77a0ccb0705ab2753d203221e26

Observation 305b1f51-acbd-419f-be71-38d19996f661 · outbound

This paper cites Kim, J W.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Kim, J W

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.473356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:d7806ce4a4631b4181761d0a7a1612f9c3369b70778367bea727ee23e0754e2c

Observation 1bc57ac3-dbca-4126-941b-3dbca3c8a708 · outbound

This paper cites Massa, F.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Massa, F

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.476643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:972dcc9638844d442cdc0f4f9d1e4c771ec8ecfb14431673fd3efcde08b024db

Observation 357ca484-add2-458f-aa01-810c9ea87b08 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Proximal Policy Optimization Algorithms

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:39:48.383998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:19b04bd9cf67eabd17e8c23c1b900c98bbb79ca8c91fd67a521271a0ecc8b73a

Observation 724f732a-8be3-4d29-860c-d7261582ab66 · outbound

This paper cites APACrefauthors \ 2024 1.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction APACrefauthors \ 2024 1

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.467159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:6c8eb0a837288eef1d56016fd31073e3db6c2064eebbf92212e7e81023982b8f

Observation d233337b-4296-40e9-81d4-9167ccf7b470 · outbound

This paper cites APACrefauthors \ 2024 2.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction APACrefauthors \ 2024 2

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-18T13:39:48.470399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:cd7952c4eca5cbd8cf3018d6d7a532d45efcb67043f03f93d1658180fe0186ec

Observation eef0cf8f-ad35-490b-9f52-6c2b723ccec9 · outbound

This paper cites Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.390382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:4afc4a0b9da667aa0b821cd57206a45e997c7f9539ce31aa56ed7859dfa65c75

Observation 4b4fd7f8-9bd0-4b4d-b44b-0a7df5b37aa7 · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.398368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:6f3fe8eab6426b663b361519d7bc12b413f2516a0368c3770d51bd5713a19ab1

Observation b6ccf608-f3df-4855-b006-c7740879e3e6 · outbound

This paper cites GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:39:48.403186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:d789689644006521646accee033d1f9cd4112c974d33ec896506f3599f0ccfe9

Observation d15e40ce-8a13-48cd-95d1-f464427993df · outbound

This paper cites Disttrain: Addressing model and data heterogeneity with disaggregated training for multimodal large language models.

Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction Disttrain: Addressing model and data heterogeneity with disaggregated training for multimodal large language models

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.268558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T13:39:48.225482Z digest=sha256:6e1d77334af1c1db0d8f0210b2b20b38a367acfacacc848c9a4ecd092c9e48b1

Pith citing papers

Observation 13ac8ad6-f421-41ce-b01f-54bfe7b17bbc · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:46144b32d1b77108546e39765c23bb1ae1f45e82ed0e96ac8bf450595c966063

Observation ab37e7fe-3520-4d25-a911-e918eb5b194c · inbound

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models cites this paper.

Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T03:42:44.523919Z digest=sha256:d3cf74b29bd497fab15f84918b292879cadcaa8c327bde934329007242ca8857

Observation a90d7b25-65a4-4383-92f1-9f8da6f8cceb · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:2fb6bfaac7c9702218c9d6fdcf80522a7059e7bb19139c4c4270c34254518f9c

Observation e1b44793-28aa-422b-870b-ede372a11e0b · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:dba569ce5e7006c43379fe3c8fdbafa7d52a1c3570d691a47c9196e26bf41916

Observation 9cfbed63-0060-4540-9caa-d89e840fbd21 · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-18T07:42:43.077644Z digest=sha256:cc925ad74fcb3df46a317042ab5fead41744e4c084ec1644101f4bcc32dc11f8

Observation 7a3b9a6b-048d-4fbd-94c0-3d777db17cb0 · inbound

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models cites this paper.

Cross-modal Consistency Guidance for Robust Emotion Control in Auto-Regressive TTS Models Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-21T21:00:39.109525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-21T20:56:37.533183Z digest=sha256:cf233bdf9c9979b42b31bed751087bedeae0b96c753492ff2ecf5832747d88e6

Observation fbe14310-0d2a-48b2-ae6a-962ebab34970 · inbound

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body cites this paper.

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-16T22:04:07.403410Z digest=sha256:c539073d6eb0f9e846403970707a37b4d33d79ff10de519eda035a2eb18f0d95

Observation 104e889b-1350-4146-acf2-dbed95c018de · inbound

Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models cites this paper.

Style Amnesia: Investigating Speaking Style Degradation and Mitigation in Multi-Turn Spoken Language Models Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-16T19:26:30.384474Z digest=sha256:a784c6cecf11c9458da84b00cd62f9948815711c9e2caceeee92a808146865f1

Observation fa3793dc-dee9-444c-8054-be4de0c92e79 · inbound

Same Words, Different Judgments: How Preferences Vary Across Modalities cites this paper.

Same Words, Different Judgments: How Preferences Vary Across Modalities Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T19:31:51.996480Z digest=sha256:e18a398b046b8a7c7b63fbbea802380251fa86f801bcf90d3ff7618dec082d33

Observation a8a81817-877a-4a6b-9868-775f19da4368 · inbound

Bridging What the Model Thinks and How It Speaks: Self-Aware Speech Language Models for Expressive Speech Generation cites this paper.

Bridging What the Model Thinks and How It Speaks: Self-Aware Speech Language Models for Expressive Speech Generation Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T15:16:10.832592Z digest=sha256:404bcb04a2eabc12e21b593b6b4cf2a0dda2f745ea39a0291d0df4d884301ef1

Observation 8a072de6-a6ef-45d3-aa97-e0e955af4523 · inbound

VERITAS: A Multi-Agent Co-Scientist for Verifiable Image-Derived Hypothesis Testing cites this paper.

VERITAS: A Multi-Agent Co-Scientist for Verifiable Image-Derived Hypothesis Testing Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T21:34:30.078813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T21:34:30.078813Z digest=sha256:9279205943fce0548edc7126d64d81fa9bdd116201a66ecaa69b69001b52ed75

Observation 65ab0000-4558-4b3b-bc61-87cf5914c5db · inbound

Why Your Tokenizer Fails in Information Fusion: A Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization cites this paper.

Why Your Tokenizer Fails in Information Fusion: A Timing-Aware Pre-Quantization Fusion for Video-Enhanced Audio Tokenization Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T14:49:51.481873Z digest=sha256:b020f532457bd855c06ebc857975b0baa287995a12f1a7be8a87369e7c1f0cce

Observation cdd9973b-928b-4ddb-ade5-a614c7dfdb0e · inbound

On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation cites this paper.

On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-10T15:30:58.664375Z digest=sha256:8868d40a93ebc3ecf0df3923cd261770cd520f45f60c8bd0ca336e4794b9efc6

Observation 74e2921e-86fe-472d-9b9f-c95e31ff23d8 · inbound

Sema: Semantic Transport for Real-Time Multimodal Agents cites this paper.

Sema: Semantic Transport for Real-Time Multimodal Agents Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T22:20:55.282442Z digest=sha256:61447b2f2bb096cdbc9f7c7081762e76fc6e6cc18106c9185935fe4d3e846768

Observation 86600281-9dfa-478c-b03d-9c1f35a28d3f · inbound

ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval cites this paper.

ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-07T16:48:54.922606Z digest=sha256:6c14be153e4a8e5df770bc6ac1b225727e0e96050f659f4661f951593156ff45

Observation 72d91187-ece0-4e0d-8c42-5fb3f5bbe05d · inbound

ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval cites this paper.

ReasonAudio: A Benchmark for Evaluating Reasoning Beyond Matching in Text-Audio Retrieval Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-08T18:28:16.549541Z digest=sha256:26d77321841b6efd4d9dedfe7ef36c4a8ead3979bdc1c83b7d79c9461db17454

Observation 72c7adcb-7c33-4a2f-b2c8-e3c7d78203c3 · inbound

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model cites this paper.

MiniMind-O Technical Report: An Open Small-Scale Speech-Native Omni Model Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-09T15:34:50.848124Z digest=sha256:5958fc014f89cd5c6918ea66ff039d04c9e2a7b273f0c1a87bcd4e21dcc33247

Observation c1ee0ef2-4d5a-470d-b565-8b193a0b988c · inbound

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization cites this paper.

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-05-08T18:08:16.051117Z digest=sha256:39b794ce9887940e71b6b6a53593eca5fe6ade425eb76005b978da97607bfa12

Observation d66889ff-1e86-49a8-ad0a-88aae65164a4 · inbound

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization cites this paper.

Sparse Tokens Suffice: Jailbreaking Audio Language Models via Token-Aware Gradient Optimization Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 2

Resolution
malformed identifier
local_arxiv, observed 2026-07-01T13:15:46.063781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-06-30T23:46:46.296060Z digest=sha256:19a0acc0faa03ca393396ccd1fd763c567a958112dd3e2675cf328ac07278c92

Observation 51020bc1-d6ab-477c-b772-dc7bf98b6e5f · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:cc9f3c040add78fd98f97c511742e20eb4e77e2ace4f89ca5a967fa010b32463

Observation a6edd245-4b5b-4bd0-a826-576966cff970 · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T03:49:58.240883Z digest=sha256:31e06a001fa93a2dc78c8760185d3918fe6cf59f3a157aa2d85171b2e09f2aae

Observation 292a4020-df8c-4849-b154-45e59cd1802e · inbound

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation cites this paper.

Boosting Omni-Modal Language Models: Staged Post-Training with Visually Debiased Evaluation Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:39:48.481060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-15T06:06:20.658030Z digest=sha256:25f89f7ff7c1e77b213390b6cb5d585ca5c4ca122ff9517e02d3d5188340ace6

Observation bad9d523-5735-4474-afb4-04a70cb742b1 · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-21T02:43:55.079617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-21T02:41:13.583493Z digest=sha256:56b7db7f1b109fb93661f847347c70c5cb6f463587d4606306855ee20fe16ac3

Observation b8d0dbc3-3147-49a9-84a8-58631c65090c · inbound

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action cites this paper.

DuplexSLA: A Full-Duplex Spoken Language Model with Synchronized Speech, Language, and Action Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-30T17:34:57.326987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-06-30T17:32:58.848455Z digest=sha256:22d54e41a1f99936a5d9cf464059ada52d269909c6c3a279579b00922d0e3efa

Observation 909eb336-2e95-4c26-9119-bc8976c9d00b · inbound

BigMac: Breaking the Pareto Frontier of Compute and Memory in Multimodal LLM Training cites this paper.

BigMac: Breaking the Pareto Frontier of Compute and Memory in Multimodal LLM Training Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.019293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-06-29T23:01:14.074894Z digest=sha256:0d17d4517ce4e3d2f20e34b84d0d1161eab2b223933c0f3f797975435b70a498

Observation 6105dafd-069c-4ad2-bbf2-f004becb9545 · inbound

Audio-Mind: An Auditable Agentic Framework for Audio Understanding cites this paper.

Audio-Mind: An Auditable Agentic Framework for Audio Understanding Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:13:17.925272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-06-29T10:03:54.653164Z digest=sha256:8f396312ed4c254e7b0b51d679af1b7e12b073fa3d43a7ab3033fde54f8b4da4

Observation d4b3731c-80aa-49c6-84d1-d6fe04227c59 · inbound

MOSS-Audio Technical Report cites this paper.

MOSS-Audio Technical Report Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-07-02T00:56:24.977576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-06-28T13:05:29.813707Z digest=sha256:d717ebc017a57fee4150680429b8df9e68307458b5fb76d9cb488b3fe9656431

Observation 5c44acbd-621d-4bbf-b340-5a12e43a4a9a · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-03T13:18:12.830247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:f29bed30f2cba5b8d91e61b89d1abe0b0cb3c003df81ead4f72df27404413055

Observation 989c9785-7c98-4ffd-9537-92a5aa516087 · inbound

M*: A Modular, Extensible, Serving System for Multimodal Models cites this paper.

M*: A Modular, Extensible, Serving System for Multimodal Models Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-03T10:27:56.698567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-06-27T10:01:04.153130Z digest=sha256:545b834b780f31af9c5bc4c55cff292401c4b0df3161f75269ffbdc87b317071

Observation da34a01a-b062-45f0-a7d8-25496182aff1 · inbound

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine cites this paper.

A Survey of Full-Duplex Spoken Dialogue Systems: Architectural Hierarchy, Interaction Ontology, and Decision State Machine Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:25.013734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-06-26T19:02:16.119373Z digest=sha256:2c7b27e62f1f594973829ef6dc20f682d517b8aa7ee318d087b03b022cf32ca5

Observation 1f1ee91c-9617-45ad-a0f8-c85f5226a99a · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 214

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T11:55:42.174147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:0c1f22e2be30721587df16f697339cfb3eb34806aeac07673172b27d2d29f34b

Observation bb9728ae-77b5-48a0-b022-301d27390afb · inbound

Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis cites this paper.

Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-02T06:56:44.326692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-07-02T06:36:42.174254Z digest=sha256:61173880171c1902c6b7f05e4869e7ad6a5097055afdc499149957fb47cbdf7a

Observation 1fb345c3-36e5-48f3-8748-467b166f3ba4 · inbound

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding cites this paper.

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-11T19:34:49.358453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T19:34:49.358453Z digest=sha256:96a3e578551d55705d7b351a522a48e5cf28732056ce12d061774e5ba1a86079

Observation a5608d3f-e015-4d6e-8d61-6f737ab51a14 · inbound

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding cites this paper.

Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T08:43:34.049064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:43:34.049064Z digest=sha256:47441b3e3a82994028fe6173d76c434ee0e23e593e05ac76ded8a4b479f3587d

Observation b1511d11-6966-4856-8b76-de8a36588d31 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment Step-Audio: Unified Understanding and Generation in Intelligent Speech Interaction

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:16.034175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:16.034175Z digest=sha256:5736899bfa2963813b9590c4ed383397e469cc97eb135b77249983c4ef8588aa