Pith. sign in

Paper Citation Record · LEDGER

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation

As of 6 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2508.20660.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20660 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:00:14.179533Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9045120-c6bd-4e02-a8ef-0072e4086007 · outbound

This paper cites Sdr–half-baked or well done? In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Sdr–half-baked or well done? In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.022586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T15:00:14.014177Z digest=sha256:ec69af47e59dac8ad3752d5dd1ef2e4a7eb4469c6e9b07c4db788254774ae95c

Observation 0c97b505-afe9-45f2-9ef1-596b508117c7 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Librispeech: an asr corpus based on public domain audio books

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.924421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T15:00:14.032179Z digest=sha256:c4a6c83e8be7b899643cd82a7dfb4eb146a23071c381050472f53c30e961980f

Observation 8f3a2d7c-3c5c-4a82-ad09-b6f4e4ee3e4e · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.045603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.045603Z digest=sha256:f925268af215cab25fc27c132a5099e76f10fa669d2f02354ff9b7a15475fd68

Observation 49e0a2e3-dd50-461a-a6b3-fc6499b242ff · outbound

This paper cites A short-time objective intelligibility measure for time-frequency weighted noisy speech.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation A short-time objective intelligibility measure for time-frequency weighted noisy speech

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.812907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T15:00:14.064215Z digest=sha256:a9ee4f34012ba5a352e231067360f13f75b842242923a635159b7cc31fc08fbf

Observation 92059eaf-e1d3-4fd6-a32d-58e5de8d0d8c · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.073906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.073906Z digest=sha256:e6ebfa2b9f1eba006f65f922257598f1a10dd03eeeabbd3661c12177dca0c0f8

Observation cd815554-0fb7-4877-a0b9-f2406c475e7b · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.080956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.080956Z digest=sha256:550625f2bc67ae9580879bfb622f505d69acbcb8df127685ccd1da2ee82e2857

Observation 379d0d46-c6e6-48b0-b389-e657e0d22e09 · outbound

This paper cites FlowDec: A flow-based full-band general audio codec with high perceptual quality.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation FlowDec: A flow-based full-band general audio codec with high perceptual quality

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.089389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.089389Z digest=sha256:b6305be991248eee46c5dead376d5ae67f7315013c23d0f1d0b2ad8e0d391fb0

Observation dd009f05-b7c0-49a8-bca3-cdaa2670d4a1 · outbound

This paper cites Codec-SUPERB: An In-Depth Analysis of Sound Codec Models.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Codec-SUPERB: An In-Depth Analysis of Sound Codec Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.103442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.103442Z digest=sha256:fe68ae7d3313f3ec446ed2e989b70307e0b13d2e7bc1b18cfcf0dc3efe766a10

Observation ef9e0fda-ef88-4633-b2cc-8b8da6e8ec72 · outbound

This paper cites Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:00:14.381991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T15:00:14.116555Z digest=sha256:ce9f17a7215e1dde728e0f305fca0660322c6ad20ce072c257223b7e3eca6815

Observation d06df0a6-9d79-4011-803e-22139be9f85e · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.133170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.133170Z digest=sha256:528af5f96912e395abb06af2069131cf1f13f17ed2081cfaeafa36111ab14846

Observation ccf870b0-a18d-4c93-abe1-94eff0dcee0c · outbound

This paper cites Qwen2.5 Technical Report.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Qwen2.5 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.143419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.143419Z digest=sha256:b39baa6e03659a768698716b061638223849520682ee223cf64128d360909165

Observation 31ade6f1-8007-4a2f-84af-848fe68826cc · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.150027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.150027Z digest=sha256:bf993a92098b8b92b1ae09660ccf95890bb5c0eaed76db4728d9721e5da69ccb

Observation c53f719a-b5e5-4271-b8d6-e98932066d8e · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.159489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.159489Z digest=sha256:b45bdf21119c03ce6f7e5c3ff2e3b8048645059b50daf88e27eb9c28eb41f0da

Observation 93203482-b7d1-49f0-8ed7-7fd2812cbe20 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.164561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.164561Z digest=sha256:acc45bc59d59978b1b3cce1b2ffea1c80e82fa05c742a017eb49a47469de87f1

Observation 5e2992c9-eb93-4c69-8f52-ea94bc1765ae · outbound

This paper cites The clips within this dataset are manually selected from public field recordings compiled by the Freesound.org project.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation The clips within this dataset are manually selected from public field recordings compiled by the Freesound.org project

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.706512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T15:00:14.173358Z digest=sha256:00ab294a882396f093714de9e944b95dd3e791991916c18939aff471c93a0363

Observation 5445decb-3782-43f6-ab4d-934660afc49d · outbound

This paper cites an unresolved cited work.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Unresolved cited work

Reference 28

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T15:00:14.675893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T15:00:14.179533Z digest=sha256:c73db549d115d0e0b94c87f90b4fabdcab8e72343e577a5fc09c3b0f6cb8be33

Observation 2fb3cf35-e48e-468f-83b3-8bc50d89a366 · outbound

This paper cites VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.058277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.058277Z digest=sha256:358ae309f4c09be950d5fc81c460806a8d5146c7c78bc36eac9ef481ac71a0d0

Observation 71afe80c-116d-42cc-b965-6cffa81d88dc · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Audio set: An ontology and human-labeled dataset for audio events

Reference 2005

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.077517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T15:00:13.991843Z digest=sha256:6b075fa64deae21f640bc708e456e82f9a0240d006925f1044d400167251f28f

Observation e4178831-f504-4363-8fbf-12f1d8439c64 · outbound

This paper cites Visqol v3: An open source production ready objective speech and audio metric.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Visqol v3: An open source production ready objective speech and audio metric

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.100047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T15:00:13.973695Z digest=sha256:028152c83b64704941a48ecc58c3ae23a391d0d11f2af56cd9e747a4986fe162

Observation a1fb8b1a-4fd4-46d0-a271-3b7e7aeec344 · outbound

This paper cites Scaling Transformers for Low-Bitrate High-Quality Speech Coding.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Scaling Transformers for Low-Bitrate High-Quality Speech Coding

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.039324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.039324Z digest=sha256:ccc8d6a62afb5b4ad9a50be124df51a4a1347b4cde6b21fe5e66bb354ccdd43f

Observation 3190eddb-5666-4dde-8754-2cc8bf56d42b · outbound

This paper cites 1632–1636,.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation 1632–1636,

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.741535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T15:00:14.069461Z digest=sha256:01a1afca53d4c2242e7df3d2aa37816554113b52e2803ae30d403563748046e9

Observation ef66f28f-a91b-48ba-9091-de9ec5c49eb6 · outbound

This paper cites V ocalsound: A dataset for improving human vocal sounds recognition.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation V ocalsound: A dataset for improving human vocal sounds recognition

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.059689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T15:00:13.999297Z digest=sha256:b57b6fcaf0ae845547592ca21bbfd32c4c3a80e92cbfb60fa5a6d5c0bba477fc

Observation 40b1610e-2ed8-40ec-b53c-be989790466a · outbound

This paper cites Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.020393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.020393Z digest=sha256:2525db0740c8647a26bab8d51f0fc7ca1411f35351b6ba2b859f6a36bd83dd98

Observation 9db2a1b1-70dc-40ce-a056-72781458b97b · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:13.980353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:13.980353Z digest=sha256:33fd54070967355c173f84d44ff9f90cf3daff44eadee57f04b4788255a7a3b2

Observation 69c6417a-e655-405f-ad0b-77ab70ec46da · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.052024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.052024Z digest=sha256:ab398ddafc3c7340b1f2063cb87923b5aaba64d472a43ed246708bcbab1099ff

Observation 05f7c0ff-f8eb-47aa-8c09-462361b66ef4 · outbound

This paper cites Benchmarking representations for speech, music, and acoustic events.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Benchmarking representations for speech, music, and acoustic events

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.040602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T15:00:14.008369Z digest=sha256:a8d4db364efc8d3793605015596e105d58ac84377b2dbade2ba05b5a47821821

Observation d8fd45dc-73ae-4d33-b7fc-36a93b950a4c · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Moshi: a speech-text foundation model for real-time dialogue

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:13.986387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:13.986387Z digest=sha256:c6c6beeb76059172bddbeeda953d4900fd6adc64713b6df6f6cf73e8b40eac77

Observation c2a2d476-99a5-47f2-b5ce-84675ebd4f29 · outbound

This paper cites Clotho- aqa: A crowdsourced dataset for audio question answering.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Clotho- aqa: A crowdsourced dataset for audio question answering

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.001969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-08-05T15:00:14.026878Z digest=sha256:e4c0123fc9b8a5cf1627f8e3dcbbd4fbb85c4812742be41d9a86637a4d46a62b

Pith citing papers

No inbound Pith citation observations are available.