Pith. sign in

Paper Citation Record · LEDGER

HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2410.04380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.04380 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:19:56.181491Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T19:02:43.552535Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 992085bb-1e13-4e77-a897-c0c0e2e2511d · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.181491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.181491Z digest=sha256:f28eaf771235b0bf03103bfd2113ed7fb4b4ce599b24352fe1f2fcd096082523

Observation 7f4a5b07-c22b-41c6-bf22-5d42a07dcd83 · inbound

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation cites this paper.

Speech Token Prediction via Compressed-to-fine Language Modeling for Speech Generation HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:27:28.424043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:27:28.424043Z digest=sha256:097af0a8a42f3c8d812e9efe67c097dbffc4e4311b9e47dfbb7fb74add7556f1

Observation ca2f0a00-cd01-453c-ab18-6e562a1fbf90 · inbound

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis cites this paper.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.722764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.722764Z digest=sha256:996983b277adeeb35d5cd409f0010e8520a700cd965045c802337eceb422483b

Observation 2347a4f6-b275-49d8-9020-0220823e8fcd · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.108916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:ee665895914e106c8ef3d1c2d3ea1952c0738294bf0c5f2bd82edb3231ed2dd5

Observation 148aecbb-bd2a-4eb1-b797-1cb8b87f8337 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 84

Resolution
unresolved
no resolver link, observed 2026-07-12T22:04:31.302192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:04:31.302192Z digest=sha256:38f06d6acf7473c0ffce79db96cbf2b3efe7038de9684bcfdfbc253759086de8

Observation cb1e853e-ccfe-4e78-a613-bd035821d010 · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:02:43.554199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:695708a87a72ce87cfe8fae481e90b4cab49b8222c049a2eb9ce49122c84059f