Pith. sign in

Paper Citation Record · LEDGER

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation

As of 13 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2508.20660.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20660 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:00:14.179533Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved16
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9045120-c6bd-4e02-a8ef-0072e4086007 · outbound

This paper cites Sdr–half-baked or well done? In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Sdr–half-baked or well done? In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.022586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T15:00:14.014177Z digest=sha256:4813b9a1457b4094600675bb41839e452a87564095d44701865ab069d50d703b

Observation 0c97b505-afe9-45f2-9ef1-596b508117c7 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Librispeech: an asr corpus based on public domain audio books

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.924421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T15:00:14.032179Z digest=sha256:dfa986016c823ecf013a7b49c251f25842e04455ce80f6c9b5f12e7c0958d650

Observation 8f3a2d7c-3c5c-4a82-ad09-b6f4e4ee3e4e · outbound

This paper cites MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.045603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.045603Z digest=sha256:df08dfb96e782728cda0aad23fbcf8f9c44f39329798c424ef748ee6be66dd69

Observation 49e0a2e3-dd50-461a-a6b3-fc6499b242ff · outbound

This paper cites A short-time objective intelligibility measure for time-frequency weighted noisy speech.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation A short-time objective intelligibility measure for time-frequency weighted noisy speech

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.812907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T15:00:14.064215Z digest=sha256:efcc6b41ff033c931951e4d02fa55d35fb8d68cbebe5562eabcdc141c6c2304a

Observation 92059eaf-e1d3-4fd6-a32d-58e5de8d0d8c · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.073906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.073906Z digest=sha256:6a95a8ef99d1580ffe42dc57ff775426a65b42227f520fe0a385341f0e964ed6

Observation cd815554-0fb7-4877-a0b9-f2406c475e7b · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.080956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.080956Z digest=sha256:898aee6e1677a5c1e321d7ce8525292012dee7ced061270d96e37b94f1e3b9af

Observation 379d0d46-c6e6-48b0-b389-e657e0d22e09 · outbound

This paper cites FlowDec: A flow-based full-band general audio codec with high perceptual quality.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation FlowDec: A flow-based full-band general audio codec with high perceptual quality

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.089389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.089389Z digest=sha256:16f35c99a77acb856616399275cff4677dae3b06d43f7e100f49cb3c68ece6a9

Observation dd009f05-b7c0-49a8-bca3-cdaa2670d4a1 · outbound

This paper cites Codec-SUPERB: An In-Depth Analysis of Sound Codec Models.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Codec-SUPERB: An In-Depth Analysis of Sound Codec Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.103442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.103442Z digest=sha256:d256ad8f761003a08ee6888eb59b4c042949291698c0b12b3f7282756c36b0aa

Observation ef9e0fda-ef88-4633-b2cc-8b8da6e8ec72 · outbound

This paper cites Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Laughter Synthesis using Pseudo Phonetic Tokens with a Large-scale In-the-wild Laughter Corpus

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:00:14.381991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T15:00:14.116555Z digest=sha256:9e58d39e8be357534bcfdc4a500b90652addb0fc5665b87470bb843157bc00eb

Observation d06df0a6-9d79-4011-803e-22139be9f85e · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.133170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.133170Z digest=sha256:6ed596f7960bee03a46c11a10a258a1ce14bb6a64b94f9d02ec8a126f6f8aacf

Observation ccf870b0-a18d-4c93-abe1-94eff0dcee0c · outbound

This paper cites Qwen2.5 Technical Report.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Qwen2.5 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.143419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.143419Z digest=sha256:4a368cfe934051fb77766fba92a534b2e0ecd44950163172feec4148d10eb54e

Observation 31ade6f1-8007-4a2f-84af-848fe68826cc · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.150027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.150027Z digest=sha256:49f67ebb7ac87f0aa0e36e9f1bbcb43d6812bec9fb4932fb002b34ebf926a19e

Observation c53f719a-b5e5-4271-b8d6-e98932066d8e · outbound

This paper cites SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation SpeechGPT: Empowering Large Language Models with Intrinsic Cross-Modal Conversational Abilities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.159489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.159489Z digest=sha256:1b763d77918301a729ac0aaf8b0925af170c9914c38d0e2d385b1cfab20075d7

Observation 93203482-b7d1-49f0-8ed7-7fd2812cbe20 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.164561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.164561Z digest=sha256:44ae0b25b5e760b0fbb3ece0cb5a3783b94827fe05ed929200af9a8f74365e0d

Observation 5e2992c9-eb93-4c69-8f52-ea94bc1765ae · outbound

This paper cites The clips within this dataset are manually selected from public field recordings compiled by the Freesound.org project.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation The clips within this dataset are manually selected from public field recordings compiled by the Freesound.org project

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.706512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T15:00:14.173358Z digest=sha256:74b9510f9eb11c7d59c35f5c529a7d30cca8ff90770296c5e91d9b877afd4b27

Observation 5445decb-3782-43f6-ab4d-934660afc49d · outbound

This paper cites an unresolved cited work.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Unresolved cited work

Reference 28

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T15:00:14.675893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T15:00:14.179533Z digest=sha256:d00b2056418f29cb23aeda020244fdc1bcf8afe7b593b0dcaad85078a8f66b8c

Observation 2fb3cf35-e48e-468f-83b3-8bc50d89a366 · outbound

This paper cites VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation VERSA: A Versatile Evaluation Toolkit for Speech, Audio, and Music

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.058277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.058277Z digest=sha256:6937d6591b1b56041c0e64b6666286db9cd729c809b6a9f057916cfbaea3f8c2

Observation 71afe80c-116d-42cc-b965-6cffa81d88dc · outbound

This paper cites Audio set: An ontology and human-labeled dataset for audio events.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Audio set: An ontology and human-labeled dataset for audio events

Reference 2005

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.077517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T15:00:13.991843Z digest=sha256:db39a60cdbce7f8b80db1d548253e31108f11e8a6d5cc15d38e1047431f9f481

Observation e4178831-f504-4363-8fbf-12f1d8439c64 · outbound

This paper cites Visqol v3: An open source production ready objective speech and audio metric.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Visqol v3: An open source production ready objective speech and audio metric

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.100047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T15:00:13.973695Z digest=sha256:6732b1e20e5eb4c7a357a8943d0a874fb73103ee9831482a90eddab92709561d

Observation a1fb8b1a-4fd4-46d0-a271-3b7e7aeec344 · outbound

This paper cites Scaling Transformers for Low-Bitrate High-Quality Speech Coding.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Scaling Transformers for Low-Bitrate High-Quality Speech Coding

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.039324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.039324Z digest=sha256:b4ff4724719e5ddd1a48761a0bc34cfed9c9fe396777347f26d57ad9f95037b5

Observation 3190eddb-5666-4dde-8754-2cc8bf56d42b · outbound

This paper cites 1632–1636,.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation 1632–1636,

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:14.741535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T15:00:14.069461Z digest=sha256:c731896190c2e2432502035776023f7c4e16fb995b11f1d17db35d1a3aef8f5e

Observation ef66f28f-a91b-48ba-9091-de9ec5c49eb6 · outbound

This paper cites V ocalsound: A dataset for improving human vocal sounds recognition.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation V ocalsound: A dataset for improving human vocal sounds recognition

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.059689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T15:00:13.999297Z digest=sha256:8e89dfe19a9ceca8cb44943b735590f330797b65ae09aa9aceb775e4a9225691

Observation 40b1610e-2ed8-40ec-b53c-be989790466a · outbound

This paper cites Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.020393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.020393Z digest=sha256:f78c10590526b47a3dccf30c325de085d52ab8107795154cba87fcd9d7162d05

Observation 9db2a1b1-70dc-40ce-a056-72781458b97b · outbound

This paper cites LibriMix: An Open-Source Dataset for Generalizable Speech Separation.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation LibriMix: An Open-Source Dataset for Generalizable Speech Separation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:13.980353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:13.980353Z digest=sha256:f9f38f9ccd2aa7d3377eca378c892382539e03235c75ab091e2ed0cb191c3504

Observation 69c6417a-e655-405f-ad0b-77ab70ec46da · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:14.052024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:14.052024Z digest=sha256:3ab51900ef8126536dd711c4b4a9621d343cdcceae17aa38e5228ad6fe10a44c

Observation 05f7c0ff-f8eb-47aa-8c09-462361b66ef4 · outbound

This paper cites Benchmarking representations for speech, music, and acoustic events.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Benchmarking representations for speech, music, and acoustic events

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.040602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T15:00:14.008369Z digest=sha256:9432a6d0a1fc3399886b4967b2f89d096c73fc53b7ac771ed80572d8c3d8ba4e

Observation d8fd45dc-73ae-4d33-b7fc-36a93b950a4c · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Moshi: a speech-text foundation model for real-time dialogue

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T15:00:13.986387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:00:13.986387Z digest=sha256:2513d33a0e8ea322c3231ac7a8d7199b3ce5116175fe47894b2aedcd5c67e966

Observation c2a2d476-99a5-47f2-b5ce-84675ebd4f29 · outbound

This paper cites Clotho- aqa: A crowdsourced dataset for audio question answering.

CodecBench: A Comprehensive Benchmark for Acoustic and Semantic Evaluation Clotho- aqa: A crowdsourced dataset for audio question answering

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:00:15.001969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T15:00:14.026878Z digest=sha256:5a3fcb3603c458d68ee7ea5701f9d9ce8e3618f290124f0bfb4bbeb6473424f2

Pith citing papers

No inbound Pith citation observations are available.