Pith. sign in

Paper Citation Record · LEDGER

Voxtral TTS

As of 19 July 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2603.25551.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2603.25551 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T11:50:26.030339Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-19T06:30:13.599613+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T21:18:22.911332Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-02T19:47:19.722749Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact6
  • verified fuzzy2
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch12

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e099125-5d51-42d8-9179-f8b163ee6863 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Voxtral TTS Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T12:26:37.526122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:e832429d324b4ca5cf338e2633f725bae56b2ea2be0edbe4942052520400cc81

Observation 93308bf7-4b84-480f-88b4-8949f5fdbb8a · outbound

This paper cites Chao, W.-H.

Voxtral TTS Chao, W.-H

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.572678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:14bae0fe541fc607c021d38d52bc6086487b8e645c1ba1410f67170b076c67e6

Observation e5aa4e0a-d80a-4e22-ac8b-8a9a02fcc81c · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing doi:10.1109/TASLP.2023.3288409.

Voxtral TTS IEEE/ACM Transactions on Audio, Speech, and Language Processing doi:10.1109/TASLP.2023.3288409

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.576135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:773f4ef42141d6e5ec4cea04998e53b59266091f5420fb9b296f23acc2455a1e

Observation 009e605d-86d5-48b8-90b3-39a3e07fa395 · outbound

This paper cites High Fidelity Neural Audio Compression.

Voxtral TTS High Fidelity Neural Audio Compression

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:39:35.889040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:32a70d6ccc172253a5b595f53770c804fa65d02389a721866c80b716661e5cbe

Observation 8b662cfa-d9e5-4c14-bfb7-c231a98b59ec · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Voxtral TTS Moshi: a speech-text foundation model for real-time dialogue

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:39:35.893639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:4d9baa02e8415ffb758d48a7bb7092a0a69d5ff4d195af55a70d366f13525495

Observation 33eec127-4b32-4e78-be6d-f7889fb1f727 · outbound

This paper cites ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification.

Voxtral TTS ECAPA-TDNN: Emphasized channel attention, propagation and aggregation in TDNN based speaker verification

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T00:39:36.281758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:3a3a5db017c95cc9a8c03101cb203b88ceae3e84b64c1c751a9f075c69485790

Observation a2ea90cb-c396-4a88-ac44-bda34d8268de · outbound

This paper cites Classifier-Free Diffusion Guidance.

Voxtral TTS Classifier-Free Diffusion Guidance

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T00:39:35.561776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:deec2fc54455eb390bea901708d01b9574eee7d11c00f660a9ffd6e727b92611

Observation fe942684-64d2-445d-8bd2-16a0564a019f · outbound

This paper cites Classifier-Free Diffusion Guidance.

Voxtral TTS Classifier-Free Diffusion Guidance

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T00:39:35.889997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:a17807afd8e249c6de26ddc09d1b1e06663ba3acc6ea53c679a6c690e7e1340e

Observation f3f11dd4-4422-4421-82ab-e01f5fd3a5a2 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =.

Voxtral TTS Efficient Memory Management for Large Language Model Serving with PagedAttention , booktitle =

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.565964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:e90397139b07e2ae7ea50ba250e4ace2ac3709b520dd563e539504a9bebf2b08

Observation b181bd6f-437d-4329-a365-cbee707a207d · outbound

This paper cites Alexander H Liu, Sung-Lin Yeh, and James R Glass.

Voxtral TTS Alexander H Liu, Sung-Lin Yeh, and James R Glass

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T00:39:36.279560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:500a1497ab75ddac87ac0b2e68a601e2c064055862395b36e7e10401770bef6c

Observation 1ed46d8e-0b15-47c3-a503-5d2f63050ffb · outbound

This paper cites Voxtral.

Voxtral TTS Voxtral

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.904186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:11c545ad0c3357638b16b762e73f9158f8c0421d29a05b67b89f50ff4354b21c

Observation d8ea44ac-40cb-4828-8173-552871608972 · outbound

This paper cites EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis.

Voxtral TTS EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.896911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:f5106aa17efe9ff2797e241f6adb1892d8d7c7d84f8e558f2804f911f6d5b00c

Observation ec8e546d-e712-4730-9e1f-5c7fb70d555c · outbound

This paper cites Scaling Transformers for Low-Bitrate High-Quality Speech Coding.

Voxtral TTS Scaling Transformers for Low-Bitrate High-Quality Speech Coding

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.910557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:f0fe48772df92bf3a81544acac365e101f9ebdcc8492cee526e775989b7a64b1

Observation b16f17ab-2d64-4dd7-b86e-2efe76bc2476 · outbound

This paper cites Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation.

Voxtral TTS Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-15T00:39:35.906928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:02b9a8d314bf2f97112001b974aee4d44dcc16a37c6c48ab65cc33c57b2cfac1

Observation 33657dbe-719b-47bd-b6b8-78e3e3089ee3 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Voxtral TTS Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T00:39:35.899779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:0b9e4ab2aa12fc982f42637e66e095b93fe7f71c4aff7560a406a24236d14991

Observation e003bb2c-f50a-4ffc-bc02-cab6c35ad609 · outbound

This paper cites STAB: Speech Tokenizer Assessment Benchmark.

Voxtral TTS STAB: Speech Tokenizer Assessment Benchmark

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:39:35.914011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:b021532184592e18df1d8bec8509e9942d34e8881843100812a69555e71683b0

Observation 59891d00-1594-4f9d-ba16-054ea3478155 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Voxtral TTS Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T00:39:35.900889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:5b5d699576989c717c5c4ba4dfa15054f3bbafcc5c6fe79e3889e5b6514afd64

Observation 1577ae56-63b0-4d02-a211-014364a9fed7 · outbound

This paper cites vllm-omni: Fully disaggregated serving for any-to-any multimodal models.

Voxtral TTS vllm-omni: Fully disaggregated serving for any-to-any multimodal models

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:39:35.892375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:25e43d7bf1cba435cb5dce320f3e5b5f7e7958268db7015aab13ca1442ba7702

Observation 0337724a-9c59-481c-9ea8-1557a43548e9 · outbound

This paper cites MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder.

Voxtral TTS MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:39:35.876040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:0488274dbe38c030578d4b57c169153ba783619d4858ad6c69cfb8848040198e

Observation cec73bd5-6426-4c71-9705-fc42a7a5ad2c · outbound

This paper cites arXiv preprint arXiv:2512.10264 (2025) 4, 5, 6.

Voxtral TTS arXiv preprint arXiv:2512.10264 (2025) 4, 5, 6

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:39:35.872905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-15T00:38:42.441340Z digest=sha256:d3e434dfec43e94d3ffa2eac6b0f226695882f0f5c91c631ecb5a65ad3e19ce4

Pith citing papers

Observation a25acf22-ec9e-41c3-b622-2d1b39e2fb65 · inbound

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech cites this paper.

Raon-OpenTTS: Open Models and Data for Robust Text-to-Speech Voxtral TTS

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-21T02:33:55.495102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-05-21T02:32:26.122526Z digest=sha256:5d88f1bdb1d89ad36278f81d3ff53770b49f8a3a8f3c446d8273f25f0f160ce8

Observation 6051f762-2d09-4799-a60a-d30249337f90 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report Voxtral TTS

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-02T19:47:19.724003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-19T06:30:13.599613+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:7cb6f057d844e29fcde22b8a02199f0fbdb5d03f45e0cde2b5b9a2fd1a9aedc9