Pith. sign in

Paper Citation Record · LEDGER

Exploring Token-Space Manipulation in Latent Audio Tokenizers

As of 6 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2605.11192.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.11192 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T02:40:05.812426Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact11
  • verified fuzzy1
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26755131-f83a-4d75-9c80-57d418cbb8de · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Exploring Token-Space Manipulation in Latent Audio Tokenizers Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:42:08.298187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:40:05.812426Z digest=sha256:a7c0e00f4a1fadd79f6af09949fc75d8d40f175e620e84ff22711a9e2aab5464

Observation 0aa1a2e2-b0ae-4882-a625-91c5e647e3b5 · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

Exploring Token-Space Manipulation in Latent Audio Tokenizers LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:42:08.303081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:40:05.812426Z digest=sha256:962ce4d2eaefecc69db8eeaae66060f5e941ac26cc8f20b70a4f71b7607f6c43

Observation 3d1b911d-cbbd-4290-b9d7-f848c3a4f4d8 · outbound

This paper cites DeepSeek-V3 Technical Report.

Exploring Token-Space Manipulation in Latent Audio Tokenizers DeepSeek-V3 Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:42:08.293827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:40:05.812426Z digest=sha256:77ba4ecf86284aa12d6102e5eb29a070e2c94290d6c6cc0300dd13fe64db66d9

Observation ca134844-a0ef-4a1c-a352-4d79de6f1246 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Exploring Token-Space Manipulation in Latent Audio Tokenizers Moshi: a speech-text foundation model for real-time dialogue

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:42:08.267606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:40:05.812426Z digest=sha256:154236a4bf3523d44cd5f96154d4a7da7645dc38e8f6f7ce33141665689c60f0

Observation 23dead95-c35c-4e0e-953c-eb86556de61b · outbound

This paper cites The Llama 3 Herd of Models.

Exploring Token-Space Manipulation in Latent Audio Tokenizers The Llama 3 Herd of Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:42:08.272249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:40:05.812426Z digest=sha256:1b65a1f3b1ff5f2188524dd8da11aabc07c55537aba1ac6136c708589697fe6e

Observation f3bc3501-d615-4054-b0a7-ae8f62f292d4 · outbound

This paper cites Mixtral of Experts.

Exploring Token-Space Manipulation in Latent Audio Tokenizers Mixtral of Experts

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T02:42:08.261695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:40:05.812426Z digest=sha256:f3d94e2467a93f2c05db9e924139561a8b8f2779ff079f70dff86ed930a62955

Observation 0a5eae74-5677-4857-a833-0b5290bc3040 · outbound

This paper cites SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound.

Exploring Token-Space Manipulation in Latent Audio Tokenizers SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:42:08.277159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:40:05.812426Z digest=sha256:31db1bb9d775ba4281d4f3023f708d67e28a9d6af2bca9dd6eb7597f469c7318

Observation fdb80f78-d43b-4f80-b29f-4a8abc29354b · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

Exploring Token-Space Manipulation in Latent Audio Tokenizers Spirit LM: Interleaved Spoken and Written Language Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:42:08.307524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:40:05.812426Z digest=sha256:491af657ae04830d6b5728cd8b3fb45d108fdee168472c6ae1d585baee902d04

Observation fe393c50-d099-4981-b76b-827c9783f5e8 · outbound

This paper cites OpenAI GPT-5 System Card.

Exploring Token-Space Manipulation in Latent Audio Tokenizers OpenAI GPT-5 System Card

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:42:08.284880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:40:05.812426Z digest=sha256:4ef63fce8c906435ff9cc55e46f55701319f644828667f7a9aec446051d30419

Observation 03070e2a-ae51-46c4-babf-7980868c548a · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Exploring Token-Space Manipulation in Latent Audio Tokenizers Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-13T02:42:08.281269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:40:05.812426Z digest=sha256:39c59e8625b5992c5ab494c37f615c4c6f6e7f1858461084ef8389fbfffd65a0

Observation 33fa20fe-e4d6-4a0c-9437-fd6e14566f71 · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

Exploring Token-Space Manipulation in Latent Audio Tokenizers BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:42:08.256552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:40:05.812426Z digest=sha256:f78d2256db8fc39391ea8e2b3c4e7617a28d11c10c709690a3db572cd956d9f8

Observation 9ffbe4ae-6f79-4c83-aa28-2e91d2073f44 · outbound

This paper cites Streaming sequence-to-sequence learning with delayed streams modeling.

Exploring Token-Space Manipulation in Latent Audio Tokenizers Streaming sequence-to-sequence learning with delayed streams modeling

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:42:08.289130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:40:05.812426Z digest=sha256:d372282baab5baa5dcc59df1845cb83d41299e8bb5588d8f1e09c95ca643eda9

Observation ca33e5d4-9b3d-4edb-ae14-df30f83c985b · outbound

This paper cites We use the utmos22_strong model loaded via torch.hub from tarepan/SpeechMOS:v1.2.0.

Exploring Token-Space Manipulation in Latent Audio Tokenizers We use the utmos22_strong model loaded via torch.hub from tarepan/SpeechMOS:v1.2.0

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T12:52:46.463087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:40:05.812426Z digest=sha256:1851a06a6692b7950b728276895b8dab4d79a22993775ca4313c3ee588ef256f

Pith citing papers

No inbound Pith citation observations are available.