Pith. sign in

Paper Citation Record · LEDGER

Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2412.04917.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04917 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:03:30.685319Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T20:45:08.132130Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4f425dfd-1a03-4aa3-8f97-c1519f08c345 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:45:08.135585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:5d907472e6c3ad70941fc288631bf72470e4e8bfa338941d4b10f37a983592e6

Observation a83a50da-cb83-4dd3-a6cc-78708deb9984 · inbound

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling cites this paper.

IMPACT: Iterative Mask-based Parallel Decoding for Text-to-Audio Generation with Diffusion Modeling Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:03:30.685319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:03:30.685319Z digest=sha256:18b19ca3aec00503088e698d58bc78759c137da4b9f988dc03619d15d2d578b2

Observation 57703305-fd3c-4bd1-8837-dd90a5334c5e · inbound

Next Tokens Denoising for Speech Synthesis cites this paper.

Next Tokens Denoising for Speech Synthesis Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T11:22:27.693175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:22:27.693175Z digest=sha256:c0750c8ade6889313ee9c0fd64b57b73112fb575a94ee05880ec3aaf91df1d9b

Observation ebd68a4c-3872-4398-9419-19b1b8df4f0f · inbound

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation cites this paper.

Fast Text-to-Audio Generation with One-Step Sampling via Energy-Scoring and Auxiliary Contextual Representation Distillation Continuous Speech Tokens Makes LLMs Robust Multi-Modality Learners

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:46:42.075264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T19:17:09.247932Z digest=sha256:226be0f5bfeda632a16c614c463cdeb2dca74fa1d1d3d265480c91b7ec59893b