Pith. sign in

Paper Citation Record · LEDGER

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing

As of 21 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.06424.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06424 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:29:50.822607Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact4
  • verified fuzzy7
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 493e87f9-4f2a-400b-9ce8-e63e37ce9c30 · outbound

This paper cites Mapache: Masked parallel transformer for advanced speech editing and synthesis.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Mapache: Masked parallel transformer for advanced speech editing and synthesis

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.183649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T04:29:50.730382Z digest=sha256:4a4158e3649b28bbda4bc818bef5cc55e9ccf3720a6cea9d7dbee1eeab3e7d54

Observation a92a2c7b-0e8b-4805-942e-2373fc3c6952 · outbound

This paper cites High Fidelity Neural Audio Compression.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing High Fidelity Neural Audio Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.739671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.739671Z digest=sha256:b8134536912192f549f0c098f1cc1a8c5341a67501a18e1aabf23655355e753a

Observation e1e26ccf-f13c-4ce4-8a66-ff0e77653a3f · outbound

This paper cites For substitution and insertion operations, the system must determine the number of codec frames allocated to the edited span before generation.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing For substitution and insertion operations, the system must determine the number of codec frames allocated to the edited span before generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.112349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T04:29:50.822607Z digest=sha256:00438851c461ec9999c01487568ef6ff02584a6ce0697577549adbf774bfb4f3

Observation 870693b7-d270-4e07-be0a-43b7cf31c432 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.753985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.753985Z digest=sha256:63217ac3fa31309f5bd2e5368a4ee36a50506689cdb89ac96f3fc96995270a9a

Observation 13167d88-46d8-4538-a913-7a85fd875129 · outbound

This paper cites Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.768197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.768197Z digest=sha256:6f659dd6b9b2ecb132153a4be1d6e2af917abee05ca948b3f73996fde03d238d

Observation 306adf4f-176d-4628-89a5-b0a44f456ae9 · outbound

This paper cites Speech inpainting: Context-based speech synthesis guided by video.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Speech inpainting: Context-based speech synthesis guided by video

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:50.991444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T04:29:50.772653Z digest=sha256:15c44ff92b35af785bf371ab342e35e324bef8f9af36635bc113763c105b46e2

Observation f294eefb-200d-46d8-8e4c-7adfd03ee704 · outbound

This paper cites Transient Noise Removal via Diffusion-based Speech Inpainting.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Transient Noise Removal via Diffusion-based Speech Inpainting

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:50.970902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T04:29:50.777180Z digest=sha256:63382e497f545facb001c64beab96439387ae0c3b0196ea9f8d5b6ca382b99c4

Observation 68e1f938-7d18-4350-9453-50e95ab20872 · outbound

This paper cites Large Language Diffusion Models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Large Language Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.786020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.786020Z digest=sha256:cf976f32b341484af0ac5116b33c78c63ca66aac5b4454463a9634e57a5debcf

Observation cff0339d-8e44-4e09-9fed-bc787adb64e1 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Robust Speech Recognition via Large-Scale Weak Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.790349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.790349Z digest=sha256:995591797dbadc29a53d7bf0f087c22bf13f1593cd54b5880e080f9b18cabae3

Observation 4460e2a5-e242-44b2-ade8-5b385c264709 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Score-Based Generative Modeling through Stochastic Differential Equations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.799726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.799726Z digest=sha256:44c0890cfb1b9e9a42faf19105b796084ffa89b96124005b1e324b2ca695489b

Observation 695dd7ba-2dbb-40ff-bdaa-3e17869b4416 · outbound

This paper cites Score-based Continuous-time Discrete Diffusion Models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Score-based Continuous-time Discrete Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.804323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.804323Z digest=sha256:ea368a4d997fa9a9a5999b8a5440dac8458f61ee959143a60573835b2c5d513f

Observation ec8c193e-abee-4711-86df-568a36703bc2 · outbound

This paper cites usee: Unified speech enhancement and editing with conditional diffusion models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing usee: Unified speech enhancement and editing with conditional diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.126709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T04:29:50.813283Z digest=sha256:3306a6806e74769fe8282d4455a52657ce3a8c0038a1a7cbef84b2fffb11eafb

Observation 496d843a-deeb-4b8e-bc93-2d289e9be7c4 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.817951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.817951Z digest=sha256:88d8d23519b8660d24a85d93eac3ff8bfc35ad808e2ace842426ba803e848ed7

Observation ae67ffad-d65e-4189-a488-2f004e558dde · outbound

This paper cites Discrete diffusion for generative modeling of text-aligned speech tokens.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Discrete diffusion for generative modeling of text-aligned speech tokens

Reference 1981

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.155056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T04:29:50.758646Z digest=sha256:1f7c1b904799bbb84b82bb7421e505604b6347b21a8df6cf0e612fc0cbb5b90a

Observation aad7ddf9-75d8-4283-804b-d82883370553 · outbound

This paper cites Block diffusion: Interpolating between autoregressive and diffusion language models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Block diffusion: Interpolating between autoregressive and diffusion language models

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.197422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T04:29:50.720825Z digest=sha256:710631b9ef31c80c5a56a46b30c9f5e6cc6eff35ffe52bb775c5ee54d5ab1dcd

Observation 8231d0da-6716-4abe-887c-ea922dd5d965 · outbound

This paper cites FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.763054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.763054Z digest=sha256:5e4cd5fcfaa66de6da57d532eba4d08c1fb34dada3391a6195fa9e1ec55286f3

Observation 5f7e274b-4338-4875-be56-f1de170b042d · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.795291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.795291Z digest=sha256:5f4914cba41c478462022ea5576969055ce8d85812ede30b63809ab8b918533b

Observation 377a4f7b-9223-4b5a-b392-376cf5a50bb7 · outbound

This paper cites Token-based audio inpainting via discrete diffusion.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Token-based audio inpainting via discrete diffusion

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.169124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T04:29:50.744507Z digest=sha256:b527f85d872516dc771f4f6f77a61a30f636cefafad695438506ec936b7e37a1

Observation 7b322f63-27e0-4658-895b-cc44b133efae · outbound

This paper cites SpeechPainter: Text-conditioned Speech Inpainting.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing SpeechPainter: Text-conditioned Speech Inpainting

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:51.098061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T04:29:50.725712Z digest=sha256:36dcc7957ec904b35d3e913189454cbc5aa76db58933ca19d5e41dd46dedff93

Observation 433897b6-e47c-4e2e-a46f-19001d27301c · outbound

This paper cites Classifier-Free Diffusion Guidance.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Classifier-Free Diffusion Guidance

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.749133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.749133Z digest=sha256:1fab225a5f18a61538ba12f85e956b356afc3f51d9725ede24f0b3e6c85d6606

Observation a755dd70-14c7-4da9-a1b6-0ed318841856 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.734798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.734798Z digest=sha256:5b18ae48e286d6230aec0ebab7577e3cc9dee1d8e40a89e2027abbbe2906d134

Observation e7f8c054-2f42-463c-98ce-d9cd246d7ac6 · outbound

This paper cites XPhoneBERT: A Pre-trained Multilingual Model for Phoneme Representations for Text-to-Speech.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing XPhoneBERT: A Pre-trained Multilingual Model for Phoneme Representations for Text-to-Speech

Reference 2025

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:50.950424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T04:29:50.781601Z digest=sha256:021be24dfc3f84d39b10c9194458d0be986024c6cddcf73e051195a5b16dd726

Observation 0d2d4c22-7347-4b15-8700-235ef0533c0e · outbound

This paper cites Ssr-speech: Towards stable, safe and robust zero-shot text-based speech editing and synthesis.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Ssr-speech: Towards stable, safe and robust zero-shot text-based speech editing and synthesis

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.141461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-10T04:29:50.808904Z digest=sha256:6b96e1bacd95266a032a4566e0e2ed5283a2e61e35e6c08713d52ef8cb934ddb

Pith citing papers

No inbound Pith citation observations are available.