Pith. sign in

Paper Citation Record · LEDGER

SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2412.15649.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15649 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T11:29:30.772401Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T13:18:12.775752Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 370d4204-988e-4b50-a439-f59ff1b8a613 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:45:08.035913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:1484cda1938338a3db2c59a9e5445aa06cd9f5fef9f557325a9b158fd8286394

Observation 81f83c4b-45f8-43b8-9d37-016b0153cecf · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.772401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.772401Z digest=sha256:e78bff43c7b4b280267da97bc739fe0338965d7bf1328357f0278527cc3db848

Observation 619f1271-d061-4ea1-9907-7f6fc93723ca · inbound

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning cites this paper.

Character Beyond Speech: Leveraging Role-Playing Evaluation in Audio Large Language Models via Reinforcement Learning SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:25:30.001718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T14:22:25.660785Z digest=sha256:aee740770341219257cc6b41df75790308616840e50940f6af1a249c70e36573

Observation 9cb079e0-7989-4bc5-bb9d-7dc195febb24 · inbound

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing cites this paper.

VITA-QinYu: Expressive Spoken Language Model for Role-Playing and Singing SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:50:56.237941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-11T01:03:09.942984Z digest=sha256:91494850b35da44f1f2e22470ad94150ecd9ab2a538c3ebdece24cb5201efb3f

Observation babbd305-f199-4614-9712-658ef4f26b97 · inbound

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation cites this paper.

Which Speech Representation Better Matches Text-Native Reasoning? A Study of Speech-Text Alignment on Frame Rate and Representation SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:18:12.777326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T08:18:23.182355Z digest=sha256:2218d660cb6bd2e183c03b8a7860e0d992767567bc5cf280704c748022e2dce0

Observation 46fb2046-e84d-4c07-840c-83b750e7cd33 · inbound

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment cites this paper.

X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment SLAM-Omni: Timbre-Controllable Voice Interaction System with Single-Stage Training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T07:12:08.148931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:12:08.148931Z digest=sha256:5a6706bdfaddcfcf9b9643dfa492f2e9a25d5e7a6bec6fe83f938eec3e496b3c