Pith. sign in

Paper Citation Record · LEDGER

Taming Data and Transformers for Audio Generation

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2406.19388.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.19388 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:59:27.912849Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.552662Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ca402966-e9f1-484e-8561-a58021840176 · inbound

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models cites this paper.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Taming Data and Transformers for Audio Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.581165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.581165Z digest=sha256:591e6964bdf94e3a7e0197f4dfbaed5cba56d964190c71c196dc54fdb931ba19

Observation 62c79e91-2061-402c-a6bf-6f4d89e674ee · inbound

HunyuanVideo: A Systematic Framework For Large Video Generative Models cites this paper.

HunyuanVideo: A Systematic Framework For Large Video Generative Models Taming Data and Transformers for Audio Generation

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.395027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-23T07:41:58.617477Z digest=sha256:06050eb67898395360dbfa9f97c4f535367cf025806f2a16ccf9a7a0b31482e4

Observation 1077e4aa-094b-498e-829b-60aa55850084 · inbound

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation cites this paper.

AV-Link: Temporally-Aligned Diffusion Features for Cross-Modal Audio-Video Generation Taming Data and Transformers for Audio Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T11:38:08.253139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:38:08.253139Z digest=sha256:082b5d79057639d949f3f762ed8cc1c8e1dc8de77e49357cfae3b2d83676d2f2

Observation 5a5f7d87-ce14-433a-bbdb-35d30f0e8508 · inbound

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis cites this paper.

MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis Taming Data and Transformers for Audio Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:35:57.685416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:35:57.685416Z digest=sha256:31931bd2320dfadb048c476a87df4c0d9b00c05acaf2f4c34d3412231f8e221d

Observation 252aef73-a3a0-4225-925a-957d08ea0c46 · inbound

ETTA: Elucidating the Design Space of Text-to-Audio Models cites this paper.

ETTA: Elucidating the Design Space of Text-to-Audio Models Taming Data and Transformers for Audio Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.734377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.734377Z digest=sha256:f65d682ee29f776afd3b8a53da9ea6f99c9a6053177b40b4eeabd70e79e025fb

Observation 6c6ef354-ca3f-4c70-aeca-c4fac11cd4bd · inbound

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport cites this paper.

LAVCap: LLM-based Audio-Visual Captioning using Optimal Transport Taming Data and Transformers for Audio Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:14:24.048855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:14:24.048855Z digest=sha256:95a90a17beebe3e060af95f428eeaa39b1a616eefdb2066837dbaef30284cbe7

Observation d902a54c-7d83-4850-9ab8-e32caae0b953 · inbound

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey cites this paper.

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Taming Data and Transformers for Audio Generation

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-10T14:36:19.677624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:36:19.677624Z digest=sha256:4582707a5bffe9f43972d9337b23225330a1de96a4d66a65ba996d98635711a7

Observation ab9921c6-7f27-4813-8fdb-9715031c8ae5 · inbound

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer cites this paper.

CLAP-ART: Automated Audio Captioning with Semantic-rich Audio Representation Tokenizer Taming Data and Transformers for Audio Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:34.136572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:34.136572Z digest=sha256:dcd56a344710e2a61638b4161bf29b4f6d92883b95247b6bd6875bf2c8ba8e39

Observation 8b22a89d-8fc6-43e5-8b3d-9b9d56dbf6b8 · inbound

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation cites this paper.

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation Taming Data and Transformers for Audio Generation

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:03:15.228449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-17T01:03:15.183360Z digest=sha256:b903f5f2e36ce33e25f8b348b74e19179f8a86e1688f56c244c2fdd19b65ed7b

Observation eb5cee6c-a1db-4e75-8407-12de4b8c8d7c · inbound

Omni2Sound: Towards Unified Video-Text-to-Audio Generation cites this paper.

Omni2Sound: Towards Unified Video-Text-to-Audio Generation Taming Data and Transformers for Audio Generation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:28:10.003947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T17:25:38.591071Z digest=sha256:bb663b4520d198273491712448ba6e86fcb895c363fd4e319a3341d14dd9a4f8

Observation c4e9a791-cfa6-48e6-9005-ff73212cf8d5 · inbound

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion cites this paper.

UNISON: A Unified Sound Generation and Editing Framework via Deep LLM Fusion Taming Data and Transformers for Audio Generation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-06-28T20:52:37.898279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-06-28T20:44:20.190064Z digest=sha256:d400f7b0037aef1d96fb44689e960350f29adf8b6d103f7a6f5d7d77b5c540bf

Observation aa35f474-cc64-41c6-bfd8-57ea10a6eed6 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Taming Data and Transformers for Audio Generation

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.554191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:579912ee6aaf86ebb3a4668aeb1125b191816449f6bc8c2c62069dcdcd5e2b5a

Observation 0f5bfcfd-e825-45fc-ad6d-8e3cd4f569ec · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Taming Data and Transformers for Audio Generation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:0dbb4310761418b413f1bd2b74c66b0cf7445b5cde39d494537b693c9d8fd675

Observation 0098b65d-aebd-41f6-95de-d804a312b0de · inbound

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching cites this paper.

VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching Taming Data and Transformers for Audio Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:59:27.912849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:59:27.912849Z digest=sha256:4846da75c385c8502e18b653cd7192e191efe18bcc203635a541d9458df980d5