Pith. sign in

Paper Citation Record · LEDGER

Vision-Integrated High-Quality Neural Speech Coding

As of 22 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 1 inbound Pith citation observation for arXiv:2505.23379.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23379 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:36.199609Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:50:33.692524Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:50:36.461320Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4fbb2ba6-3512-4f00-9f03-ff89b0e0c602 · outbound

This paper cites Vision-Integrated High-Quality Neural Speech Coding.

Vision-Integrated High-Quality Neural Speech Coding Vision-Integrated High-Quality Neural Speech Coding

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:50:36.534950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:33.692524Z digest=sha256:9b272c727c3b16942626a14eb0a63bb6c099c5a616078a71a20f8a212b7677e5

Observation d6a41dc8-bbb7-4387-8cf5-a2c5edfb9303 · outbound

This paper cites Overview As shown in Figure 1, VNSC consists of a speech coding mod- ule, an image analysis-synthesis module and a feature fusion module.

Vision-Integrated High-Quality Neural Speech Coding Overview As shown in Figure 1, VNSC consists of a speech coding mod- ule, an image analysis-synthesis module and a feature fusion module

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:40.297099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:33.772725Z digest=sha256:7cb4e72dd1103ec4ab5ebcec2a47662c34a462c7a213d648af2f03333bd963d0

Observation 25c20d15-83cc-4562-a600-fb73f8e5d5a4 · outbound

This paper cites an unresolved cited work.

Vision-Integrated High-Quality Neural Speech Coding Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:40.079579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:33.838446Z digest=sha256:b4a2ea2816b47c0cef6afe0020cfef918df0f78bb8dae61e16223d335300f3b9

Observation 98a283bd-3523-440a-b593-dedb784c6c8a · outbound

This paper cites an unresolved cited work.

Vision-Integrated High-Quality Neural Speech Coding Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:50:39.938610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:33.954738Z digest=sha256:1f8922fa2708cbb8cccb6681037f31660875d49d6ff8ebebf0ac17020b0d01cf

Observation 250044a9-f627-49dd-a427-5d32286a1307 · outbound

This paper cites VNSC is built upon the speech-modal MDCTCodec, with visual information extracted from lip images flowing into the speech coding process.

Vision-Integrated High-Quality Neural Speech Coding VNSC is built upon the speech-modal MDCTCodec, with visual information extracted from lip images flowing into the speech coding process

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:39.716689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:34.058108Z digest=sha256:df19890d54ed137bea512026aba5b950116c52560fb6b228dc891412b54eb5ed

Observation b63097c6-e681-49e1-87d4-7735c2a88809 · outbound

This paper cites Recommendation G.711: Pulse code modulation (PCM) of voice frequencies,.

Vision-Integrated High-Quality Neural Speech Coding Recommendation G.711: Pulse code modulation (PCM) of voice frequencies,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:39.517538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:34.158951Z digest=sha256:3a8a0e8e0f276b261e7b1337425bf8bd04991f001dd9e51c537afe27048d9607

Observation 09ef0972-0153-4e21-96e7-9c8649d87c8c · outbound

This paper cites Recommendation G.723: Speech coders for multimedia communications: dual-rate coder (5.3/6.3 kbps),.

Vision-Integrated High-Quality Neural Speech Coding Recommendation G.723: Speech coders for multimedia communications: dual-rate coder (5.3/6.3 kbps),

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:39.303607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:34.228261Z digest=sha256:4cf0ecffb41bcbeaf740f698f86d624db1482470722113b9dffcb1cdbe495a52

Observation 511bda03-b79d-4c90-acbf-19c4e09ddf7c · outbound

This paper cites Recommendation g.726: 40, 32, 24, and 16 kbps adap- tive differential pulse code modulation (ADPCM),.

Vision-Integrated High-Quality Neural Speech Coding Recommendation g.726: 40, 32, 24, and 16 kbps adap- tive differential pulse code modulation (ADPCM),

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:39.132658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:34.314412Z digest=sha256:e14bf1ede2a6ab188afbe5c5a703a6aabab5372ad0bbc4cf4277f306741c6360

Observation deecb13e-b03e-47be-aa0d-0ddfc3c2cb34 · outbound

This paper cites Recommendation G.729: Coding of speech at 8 kbps using conjugate-structure algebraic-code-excited linear prediction (CS-ACELP),.

Vision-Integrated High-Quality Neural Speech Coding Recommendation G.729: Coding of speech at 8 kbps using conjugate-structure algebraic-code-excited linear prediction (CS-ACELP),

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:38.977607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:34.388617Z digest=sha256:111d9156733955ade6b9dc8b1bba529e8b2dfc5ede4b2226e5b0b254496b5676

Observation 31808cef-cef9-4d63-beca-2c361bb3c11d · outbound

This paper cites SoundStream: An end-to-end neural audio codec,.

Vision-Integrated High-Quality Neural Speech Coding SoundStream: An end-to-end neural audio codec,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:34.461142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:34.461142Z digest=sha256:a30d3a612ad416ef92c59ce0dc0e2b8e78f2e5aef70f9504f6170cc88f1698af

Observation 5091cf6c-f157-4820-8046-c2724f6058c4 · outbound

This paper cites A review of vector quantization tech- niques,.

Vision-Integrated High-Quality Neural Speech Coding A review of vector quantization tech- niques,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:38.841316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:34.548372Z digest=sha256:936d5119bf9f023dfed6292f4d795168f4305ad959e959802f4d952a8fb01051

Observation 52c5127c-ac04-4d30-9de3-ce38584d7de3 · outbound

This paper cites High fidelity neural audio compression,.

Vision-Integrated High-Quality Neural Speech Coding High fidelity neural audio compression,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:34.631712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:34.631712Z digest=sha256:e0dd115dc8b75fb1ef82daa441025ff6b3dc9655409998109cf8394c20bd7149

Observation cb3a9791-83df-42c5-b990-e4fc8f7d12cd · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

Vision-Integrated High-Quality Neural Speech Coding HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:34.712516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:34.712516Z digest=sha256:3515ef4fa838e290707fe1009328bd757be5963b2b180f81ab3d81c704eb6a1c

Observation 8551fd59-4311-4472-a9fa-9bd0a5ed5f46 · outbound

This paper cites APCodec: A neural audio codec with parallel amplitude and phase spec- trum encoding and decoding,.

Vision-Integrated High-Quality Neural Speech Coding APCodec: A neural audio codec with parallel amplitude and phase spec- trum encoding and decoding,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:34.799104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:34.799104Z digest=sha256:836e994c405a531967fe54a9c2a82804d8955e55d86de86ce1a92e7a5873710a

Observation aa39616b-e0e7-4ed1-827d-2321e6f97fc5 · outbound

This paper cites Mdctcodec: A lightweight mdct-based neural audio codec towards high sampling rate and low bitrate scenarios,.

Vision-Integrated High-Quality Neural Speech Coding Mdctcodec: A lightweight mdct-based neural audio codec towards high sampling rate and low bitrate scenarios,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:38.627209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:34.902677Z digest=sha256:f82a22f1db2ba8f36128e54e7202b49b8c7730e4fc823d8d1c099957ee1609fc

Observation f84ec758-6146-4190-922d-4311a50b54c1 · outbound

This paper cites DM-Codec: Distilling multimodal representations for speech tokenization,.

Vision-Integrated High-Quality Neural Speech Coding DM-Codec: Distilling multimodal representations for speech tokenization,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:35.013098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:35.013098Z digest=sha256:c34604c85bd29c28302a2783a9a404dc78d97cf096b22b7715918d2fc62469ba

Observation cf1d0463-7a56-4f46-87d5-20177b1622bf · outbound

This paper cites The conversation: Deep audio-visual speech enhancement,.

Vision-Integrated High-Quality Neural Speech Coding The conversation: Deep audio-visual speech enhancement,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:38.440004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:35.070455Z digest=sha256:120139382186003ff187a6c61dda38a72a80942e385acd1f2a6cb63e131d4569

Observation 2c7aecbf-bb65-4c15-890d-12399e068b1e · outbound

This paper cites Vsegan: Visual speech enhancement generative adversarial net- work,.

Vision-Integrated High-Quality Neural Speech Coding Vsegan: Visual speech enhancement generative adversarial net- work,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:38.250109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:35.180640Z digest=sha256:668e850f2e1c396cfbf2604f0a5610525e0e569894521e6dc889995e17715825

Observation 56d598f4-c43e-4b42-8d91-b09d2cb635e2 · outbound

This paper cites Incorporating ultra- sound tongue images for audio-visual speech enhancement,.

Vision-Integrated High-Quality Neural Speech Coding Incorporating ultra- sound tongue images for audio-visual speech enhancement,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:38.053040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:35.288448Z digest=sha256:4cea0e82286e3fc85517bf5d58dd8f907800a62a34be6b11587cefb0939065a8

Observation a31dba6f-16c4-422e-b95e-deba50925826 · outbound

This paper cites Improving visual speech enhancement network by learning audio-visual affinity with multi-head attention,.

Vision-Integrated High-Quality Neural Speech Coding Improving visual speech enhancement network by learning audio-visual affinity with multi-head attention,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:37.910741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:35.366485Z digest=sha256:c17d94594ad175ed92c48ca45037cd34020c4eba90a11cb0e1b1c6d6f464cba3

Observation fd840dd6-135b-4e53-a130-f89dff2afe3d · outbound

This paper cites TaL: a synchronised multi-speaker corpus of ultrasound tongue imaging, audio, and lip videos,.

Vision-Integrated High-Quality Neural Speech Coding TaL: a synchronised multi-speaker corpus of ultrasound tongue imaging, audio, and lip videos,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:37.731528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:35.413134Z digest=sha256:898b253511df37efaf9c39a88e0cece5edb3a2fb2964c55009144ef7b6ea107a

Observation 5e24414c-2326-4c23-9d60-a652e68831ae · outbound

This paper cites Deep audio-visual speech recognition,.

Vision-Integrated High-Quality Neural Speech Coding Deep audio-visual speech recognition,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:37.563871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:35.521183Z digest=sha256:ccb1713f0397631f593b3ffbe1dc9420e0fc98047589ca798964f47d362376a3

Observation da2cc957-4803-4b3f-b84b-4bd11fc504d5 · outbound

This paper cites ConvNeXt v2: Co-designing and scaling convnets with masked autoencoders,.

Vision-Integrated High-Quality Neural Speech Coding ConvNeXt v2: Co-designing and scaling convnets with masked autoencoders,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:35.615922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:35.615922Z digest=sha256:576f00d99822948e2bf863a2fd9b2934da767f0443a9963abc05c5d4b609f50e

Observation 1ae59075-736a-471a-88ad-726a3c42d508 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Vision-Integrated High-Quality Neural Speech Coding Gaussian Error Linear Units (GELUs)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:35.712227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:35.712227Z digest=sha256:e163b318a9b8553595eb318a9f39f4ad2ceb7349e1428d11fe73a833e530c498

Observation 53a7abad-c408-4b81-b797-fbde56d063a0 · outbound

This paper cites Converting video formats with ffmpeg,.

Vision-Integrated High-Quality Neural Speech Coding Converting video formats with ffmpeg,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:37.330103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:35.835330Z digest=sha256:b4575149ccbd49f8906ed7cce6683dbedc12735447ab4c36438c89dca31d040b

Observation b4f915c7-eaa2-4678-bab0-e134a6df0b8c · outbound

This paper cites Decoupled weight decay regulariza- tion,.

Vision-Integrated High-Quality Neural Speech Coding Decoupled weight decay regulariza- tion,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:37.107644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:35.961443Z digest=sha256:58d483af712c71f37845219bb2e4b74d7b218e9994b5c5111e8175f5c36e34d8

Observation 2b135383-6ef9-4d95-9d66-31f8f24ec55b · outbound

This paper cites P. 862.2: Wideband extension to recom- mendation P. 862 for the assessment of wideband telephone networks and speech codecs,.

Vision-Integrated High-Quality Neural Speech Coding P. 862.2: Wideband extension to recom- mendation P. 862 for the assessment of wideband telephone networks and speech codecs,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:36.921223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:36.014763Z digest=sha256:ae1eeb398adbe7a7f81345df5a7d5370041e80f31441933411d59751a9feea50

Observation 42100a44-3988-456e-b4d4-04bc84b8c6f0 · outbound

This paper cites Evaluation of objective measures for speech enhancement,.

Vision-Integrated High-Quality Neural Speech Coding Evaluation of objective measures for speech enhancement,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:50:36.696296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:36.121916Z digest=sha256:4c99dab06981fed64aedadd075d2ad7fd875241c72a5de46684478170a25be14

Observation 66c807d6-df42-4ff4-870e-d991b8c967a8 · outbound

This paper cites A short- time objective intelligibility measure for time-frequency weighted noisy speech,.

Vision-Integrated High-Quality Neural Speech Coding A short- time objective intelligibility measure for time-frequency weighted noisy speech,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:50:36.199609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:50:36.199609Z digest=sha256:e6e67d0508634af23ebf1048d2e0c0d5afc23278245d32100475483c432c4bc2

Pith citing papers

Observation 4fbb2ba6-3512-4f00-9f03-ff89b0e0c602 · inbound

Vision-Integrated High-Quality Neural Speech Coding cites this paper.

Vision-Integrated High-Quality Neural Speech Coding Vision-Integrated High-Quality Neural Speech Coding

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T12:50:36.534950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T12:50:33.692524Z digest=sha256:9b272c727c3b16942626a14eb0a63bb6c099c5a616078a71a20f8a212b7677e5