Pith. sign in

Paper Citation Record · LEDGER

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing

As of 10 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2608.06424.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06424 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:29:50.822607Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact4
  • verified fuzzy7
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 493e87f9-4f2a-400b-9ce8-e63e37ce9c30 · outbound

This paper cites Mapache: Masked parallel transformer for advanced speech editing and synthesis.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Mapache: Masked parallel transformer for advanced speech editing and synthesis

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.183649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.730382Z digest=sha256:cded0d88493095699258a8b706e8f595ecfdc6c691e0ef70c739c570bd048d36

Observation a92a2c7b-0e8b-4805-942e-2373fc3c6952 · outbound

This paper cites High Fidelity Neural Audio Compression.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing High Fidelity Neural Audio Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.739671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.739671Z digest=sha256:8c5e1d5e27389e19684d43f42b9379ba5754e972a364e4f583ad860467c2668f

Observation e1e26ccf-f13c-4ce4-8a66-ff0e77653a3f · outbound

This paper cites For substitution and insertion operations, the system must determine the number of codec frames allocated to the edited span before generation.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing For substitution and insertion operations, the system must determine the number of codec frames allocated to the edited span before generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.112349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.822607Z digest=sha256:e3c129703f2222e2b89c831b646ed0859702e4d06ae160c5d53b351efa329b40

Observation 870693b7-d270-4e07-be0a-43b7cf31c432 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.753985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.753985Z digest=sha256:7ed7b6ba11c4f9d7f5b448d98c5e396208729a8b9f3d318d856db939c79c84d7

Observation 13167d88-46d8-4538-a913-7a85fd875129 · outbound

This paper cites Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.768197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.768197Z digest=sha256:c961dbfc5bd5bf0ad5ee264b4d41ce2dc9250ad833d38155315b95ef8babb190

Observation 306adf4f-176d-4628-89a5-b0a44f456ae9 · outbound

This paper cites Speech inpainting: Context-based speech synthesis guided by video.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Speech inpainting: Context-based speech synthesis guided by video

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:50.991444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.772653Z digest=sha256:f21a22c668f3f5172fbe12f01687596901732703eec643fed06cceeabbeee7a1

Observation f294eefb-200d-46d8-8e4c-7adfd03ee704 · outbound

This paper cites Transient Noise Removal via Diffusion-based Speech Inpainting.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Transient Noise Removal via Diffusion-based Speech Inpainting

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:50.970902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.777180Z digest=sha256:b196bd43ab9336198d6cd0b08defca71de6ecba3d3921343ce34d6525c5d2f6b

Observation 68e1f938-7d18-4350-9453-50e95ab20872 · outbound

This paper cites Large Language Diffusion Models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Large Language Diffusion Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.786020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.786020Z digest=sha256:a6d789b052d4bf0afdcb15c942301138d6e2df9243945b107f432581f7fe46e5

Observation cff0339d-8e44-4e09-9fed-bc787adb64e1 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Robust Speech Recognition via Large-Scale Weak Supervision

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.790349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.790349Z digest=sha256:78c7e2536856aa9c373d5d00e28cf525fbd5abe29ce82c5245dc565318fe5092

Observation 4460e2a5-e242-44b2-ade8-5b385c264709 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Score-Based Generative Modeling through Stochastic Differential Equations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.799726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.799726Z digest=sha256:28deaec28f080bc901534f19c79a7f0c0b6293cf78c006af5e1be416fc61c462

Observation 695dd7ba-2dbb-40ff-bdaa-3e17869b4416 · outbound

This paper cites Score-based Continuous-time Discrete Diffusion Models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Score-based Continuous-time Discrete Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.804323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.804323Z digest=sha256:9ea390db07da40b756440fa70bc60439b994ea3fa2eb8739b3b3f35a3a244f06

Observation ec8c193e-abee-4711-86df-568a36703bc2 · outbound

This paper cites usee: Unified speech enhancement and editing with conditional diffusion models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing usee: Unified speech enhancement and editing with conditional diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.126709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.813283Z digest=sha256:c243d16c3ce15e836da45a4b249d404ba5d49783337fe485cbd268b72e33bcd0

Observation 496d843a-deeb-4b8e-bc93-2d289e9be7c4 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.817951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.817951Z digest=sha256:7ef3e2caf19f94d9453a8b6f7ad7a3ebc3aa0c53a0aab2fb97e2446c1daaa86a

Observation ae67ffad-d65e-4189-a488-2f004e558dde · outbound

This paper cites Discrete diffusion for generative modeling of text-aligned speech tokens.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Discrete diffusion for generative modeling of text-aligned speech tokens

Reference 1981

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.155056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.758646Z digest=sha256:a0e2b4a36db7fde62a035955b1548d357b5d71b4122c541f82c7f404edbef407

Observation aad7ddf9-75d8-4283-804b-d82883370553 · outbound

This paper cites Block diffusion: Interpolating between autoregressive and diffusion language models.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Block diffusion: Interpolating between autoregressive and diffusion language models

Reference 2011

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.197422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.720825Z digest=sha256:d69ccc73b14846cb0caf786826a21b9d24ef9edbca5020416ff43baba599a53d

Observation 8231d0da-6716-4abe-887c-ea922dd5d965 · outbound

This paper cites FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.763054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.763054Z digest=sha256:266d2b672670a3b38d015d01bb0c88f2c2be701db8bd27715f186e00afbf47fe

Observation 5f7e274b-4338-4875-be56-f1de170b042d · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.795291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.795291Z digest=sha256:6416e0269780c4fb04e67fd14980065457771d3785beb2bcccdfedc60fe47c55

Observation 377a4f7b-9223-4b5a-b392-376cf5a50bb7 · outbound

This paper cites Token-based audio inpainting via discrete diffusion.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Token-based audio inpainting via discrete diffusion

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.169124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.744507Z digest=sha256:ef0f475ea12881ebff04a854898a38fe2dc6544521f53b37ab5f287eba1230ef

Observation 7b322f63-27e0-4658-895b-cc44b133efae · outbound

This paper cites SpeechPainter: Text-conditioned Speech Inpainting.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing SpeechPainter: Text-conditioned Speech Inpainting

Reference 2022

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:51.098061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.725712Z digest=sha256:6dc81fae5bbdca40fbda24d9695b76ddd66f9fb9efb0cd8c6558b8b509235489

Observation 433897b6-e47c-4e2e-a46f-19001d27301c · outbound

This paper cites Classifier-Free Diffusion Guidance.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Classifier-Free Diffusion Guidance

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.749133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.749133Z digest=sha256:f1af123859033eabb88477c26d1ae654ec776131a19ab2e06e7083b4a3efbff7

Observation a755dd70-14c7-4da9-a1b6-0ed318841856 · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T04:29:50.734798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T04:29:50.734798Z digest=sha256:55e43dbcb98745b35152d96a6023ff3e8f4e2fb758c911f0a5c01cc98a7b5998

Observation e7f8c054-2f42-463c-98ce-d9cd246d7ac6 · outbound

This paper cites XPhoneBERT: A Pre-trained Multilingual Model for Phoneme Representations for Text-to-Speech.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing XPhoneBERT: A Pre-trained Multilingual Model for Phoneme Representations for Text-to-Speech

Reference 2025

Resolution
verified exact
local_arxiv, observed 2026-08-10T04:29:50.950424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.781601Z digest=sha256:50b38d2851e4e038439f88210dc832dbf64aba07d910905db961829481cd3df0

Observation 0d2d4c22-7347-4b15-8700-235ef0533c0e · outbound

This paper cites Ssr-speech: Towards stable, safe and robust zero-shot text-based speech editing and synthesis.

Multi Codec Discrete Diffusion Model for Text Guided Speech Inpainting and Editing Ssr-speech: Towards stable, safe and robust zero-shot text-based speech editing and synthesis

Reference 2026

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T04:29:51.141461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-10T04:29:50.808904Z digest=sha256:30f4c305cc797583bc79c9cc40e3bdaabd37750dc7a1e2c87635c876521a63d0

Pith citing papers

No inbound Pith citation observations are available.