Pith. sign in

Paper Citation Record · LEDGER

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception

As of 10 August 2026, this Paper Citation Record lists 10 of 10 outbound references and 0 inbound Pith citation observations for arXiv:2605.31521.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.31521 v1

Coverage vector

measured 10 of 10 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T22:26:20.100595Z

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

10 of 10 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 83e3927a-75e6-426b-9fb7-1caeefb740a1 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T22:32:44.573959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:26:20.100595Z digest=sha256:00a350bc02e8929668e1309dbd7b0a64a9c057adc15ac26216f101676addd61d

Observation 74bfcf32-55f1-46f8-83b2-7c9ea3653a36 · outbound

This paper cites In2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 1–8.

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception In2023 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), pages 1–8

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T22:26:20.100595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:26:20.100595Z digest=sha256:242df7b515eb00d67b28e26fcef2d44b1f5ae43ab48407325595079e6524449f

Observation 71943c2b-405d-41df-b53c-942ac987945e · outbound

This paper cites In Advances in Neural Information Processing Systems, volume 38.

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception In Advances in Neural Information Processing Systems, volume 38

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T22:26:20.100595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:26:20.100595Z digest=sha256:b97e3e6f7c8aa676a59230228524ac6747618885f9258dfc80a397663e5d0051

Observation c1637969-f234-4c78-9fe8-b2521552ec14 · outbound

This paper cites DASB - Discrete Audio and Speech Benchmark.

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception DASB - Discrete Audio and Speech Benchmark

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T22:32:44.576420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:26:20.100595Z digest=sha256:fb7f7d61c97f335a8fadc35d223d30fe68460fee27c81e166c1f8efa1b1bf509

Observation 64c50274-e662-4a36-9526-152e64fe279a · outbound

This paper cites Qwen2.5-Omni Technical Report.

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception Qwen2.5-Omni Technical Report

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-06-28T22:32:44.578700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:26:20.100595Z digest=sha256:0f12778ffede956a89ea9502bc6b448c1d2a18068ceff484e4923d0b9d32e320

Observation cb4ffe68-c56e-4813-9e8f-8f240219b418 · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 6

Resolution
malformed identifier
arxiv_id, observed 2026-06-28T22:32:44.571649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T22:26:20.100595Z digest=sha256:3fdc94287b3c1f70684af3c75cb19fd659e3d0f9a8a4c26a1f1157580842e9b7

Observation 6d92705c-b51f-4932-bc65-5c4c00c3b1a4 · outbound

This paper cites We use the officially released most powerful large, 75Hzvariant in our experiments.

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception We use the officially released most powerful large, 75Hzvariant in our experiments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T22:26:20.100595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:26:20.100595Z digest=sha256:d70f15506606da9b083ea9237b1c2367835cb5f69ed8ff276056328af7b3f181

Observation fc386693-3f73-492d-bf42-d91c96da797e · outbound

This paper cites an unresolved cited work.

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T22:26:20.100595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:26:20.100595Z digest=sha256:991cd13479877a73a46388ddb0436b449c51a658bfe449553b6cddda9fb502a0

Observation fb2df804-748d-496d-afa7-1b58eaa91b9b · outbound

This paper cites It can compress speech into highly efficient discrete tokens at a significantly lower frame rate while ensuring robust semantic preservation.

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception It can compress speech into highly efficient discrete tokens at a significantly lower frame rate while ensuring robust semantic preservation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T22:26:20.100595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:26:20.100595Z digest=sha256:147bff87ab1555ffe5dfd32802f0079d04057a760f453cacadbd26007637f1e9

Observation a9f19da0-24d5-403d-bb7b-81b4b6072ccf · outbound

This paper cites Perfect Consistency.

UniAudio-Token: Empowering Semantic Speech Tokenizers with General Audio Perception Perfect Consistency

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T22:26:20.100595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T22:26:20.100595Z digest=sha256:603393c354b0892c108593955d75968b48f74f75cc582e07657a79b95cdabaef

Pith citing papers

No inbound Pith citation observations are available.