Pith. sign in

Paper Citation Record · LEDGER

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts

As of 23 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2605.28063.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.28063 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T10:28:18.202974Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact15
  • verified fuzzy0
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch14

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ece5bcda-246a-4892-8c05-8ffd6a467d9c · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:23:28.661916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:578ba1ff44fb18faa5c4067c017d4985aae0e582ba37d462ad0932c53ae167ee

Observation 0d413b2f-096d-407f-9f01-caefb6e09dae · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T10:33:17.883486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:4ffdd7c5805f5ac241a9a2febac1b3592cffdc84fdd84dbc64dafc3cd9a843b2

Observation 371d6347-b1b1-4839-9dc8-6ea913c27278 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:33:17.880668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:ee7c6f0310caba4f0f3cbae3aa7409bea4f4334502095b25185e846efaaf0220

Observation f534a123-42ed-408f-8bca-eecc927ea264 · outbound

This paper cites Qwen3-TTS Technical Report.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Qwen3-TTS Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:33:17.875268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:40b624fb3908dc30534ab827a82d0c674a4e505a319d2239ce78a8dc49869555

Observation bfea9cc9-0dc8-4d42-8dc2-c4457ece23b4 · outbound

This paper cites Audiogen: Textually guided audio generation,.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Audiogen: Textually guided audio generation,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T10:28:18.202974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:7e1bb18c548e5f3c2204f9b77eb4f230df4dfe61fe6404c76ea99eb234e08c08

Observation e04477e7-b9e9-4068-a968-b89a18add4f8 · outbound

This paper cites Plumbley.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Plumbley

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:33:17.869872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:08766fe6a2e6831c6b3aa6a7213db8c08bc5043b7a5f4ced51090226992551af

Observation 46e73937-64d8-44ac-9a81-f22cf4d1f8ab · outbound

This paper cites Masked image pretraining on language assisted representation.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Masked image pretraining on language assisted representation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:33:17.876811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:84a8d8958366aad9f58ed0ae43d9539cded5bc40aaab4cbfee791a821108d92c

Observation a9413d29-44bb-4b6b-9150-39d4fdc32e0b · outbound

This paper cites URL http://dx.doi.org/10.1109/ICASSP48485.2024.10447579.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts URL http://dx.doi.org/10.1109/ICASSP48485.2024.10447579

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:33:17.878156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:bee91e9c1fe667c21f796abb1e081a115569e03c9bc49cc7ad1ac4beb1420b24

Observation e4932df2-b7c3-4b7e-adab-71f687d3de51 · outbound

This paper cites Masked image pretraining on language assisted representation.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Masked image pretraining on language assisted representation

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:33:17.928783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:5e1e550ce798b05527dbfacfc0ba932cceba8b04c2936e89ea11a4a8b54ed783

Observation b0d47770-c9c5-4768-85e5-a4b28b290007 · outbound

This paper cites ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts ControlAudio: Tackling Text-Guided, Timing-Indicated and Intelligible Audio Generation via Progressive Diffusion Modeling

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:33:17.901185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:dbf3a9dbb6767fb9b3b231d6734fe62a1f4e709d875ff599ba94a54f1f859f29

Observation 444d884a-4575-45da-9dac-ac58abdf689f · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:17.930200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:918dd4ff2cb09ace9c9902e301bba19fd08a75a19963a97b28b72316299b9e59

Observation b30c9e72-229f-4392-99ef-d8ffa8b9726a · outbound

This paper cites Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu

Reference 12

Resolution
metadata mismatch
doi, observed 2026-06-29T10:33:17.909901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:6a10fec61feefd88c0f316f92ba338dc1bd0d2b10ef63749993ccc79c5fb458e

Observation 0741283b-44c7-485a-b48f-3f377bb160fa · outbound

This paper cites A udio C aps: Generating captions for audios in the wild.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts A udio C aps: Generating captions for audios in the wild

Reference 13

Resolution
verified exact
doi, observed 2026-06-29T10:33:17.911656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:2f7ffd9cd2ebe9376c62d9f8739eb1826c5a1984b92530b6a89f231332a8f0c0

Observation dfbf643f-4677-4750-9cc7-a5636ba2e02e · outbound

This paper cites Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Wavcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T10:28:18.202974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:7434ea4426f1d5911a9c03324a0f1a25fe4e97dc25284d6f4b93325000f9743b

Observation 57ac3978-4a6f-4105-9708-88906345c40e · outbound

This paper cites WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio- Language Multimodal Research.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio- Language Multimodal Research

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:33:17.931394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:d258f43560a56d411146b3ff98faf8d2753373f808b021c8729add5f24858201

Observation 0fec8409-fc71-4ffe-8518-67e7d803d2bf · outbound

This paper cites Masked image pretraining on language assisted representation.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Masked image pretraining on language assisted representation

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:33:17.946571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:5b94007539ce9ed802db7a1338c669dba38b965fbf187492009a7f37e1f5b75c

Observation c1bc5f90-f798-4570-9dff-5051fbf68911 · outbound

This paper cites Freeaudio: Training-free timing planning for controllable long-form text-to-audio generation,.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Freeaudio: Training-free timing planning for controllable long-form text-to-audio generation,

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:17.901912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:dd9574b6f6c43fc0807957104fb66dc1a4c1dc1ca175c17a0def9db6fcd16fb6

Observation 4a491638-ff52-4895-90a5-e3ffaea17fac · outbound

This paper cites Masked image pretraining on language assisted representation.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Masked image pretraining on language assisted representation

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:33:17.907854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:648a19f03c25205d05335c54007325225e4d146380e70ac4814fa8358a8f87ad

Observation c8d226f0-b228-4cda-8687-c8de081ba039 · outbound

This paper cites model-predicted CoT.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts model-predicted CoT

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:33:17.937711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:73279732be1074c43af7f5fc46394af363017fd95ecf48ff9f04187e4ea8cef2

Observation 15ef98f5-b922-4d06-a182-15a3368a65c9 · outbound

This paper cites Flexivoice: Enabling flexible style control in zero- shot tts with natural language instructions.arXiv preprint arXiv:2601.04656,.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Flexivoice: Enabling flexible style control in zero- shot tts with natural language instructions.arXiv preprint arXiv:2601.04656,

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:17.913196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:c5d8b4ff8f731496397cd1a39d1b4dee11f7129633178a94bbed715fa9bbf25c

Observation 5c51c5e4-1523-44d3-a286-74cf8ac7a6d0 · outbound

This paper cites MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts MoE-TTS: Enhancing Out-of-Domain Text Understanding for Description-based TTS via Mixture-of-Experts

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:17.871334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:542f7cb7cc7ec75bbd02b55d0ec6b9808d49e9bd0c81c33adde1cd4a51da4b86

Observation e2101993-a06f-4051-818d-136ce8d283bc · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:17.934369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:5e9374982e71b267b21e62fb4207570409a639dac46676f5c09a55f4407d73c3

Observation d90cdc08-f892-4acc-9532-5626d9dd7b9e · outbound

This paper cites Fugatto 1: Foundational generative audio transformer opus 1,.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Fugatto 1: Foundational generative audio transformer opus 1,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-29T10:28:18.202974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:70fbe86cc3b03701fb6f71c5d7f8175da16cbdbb5821210abb0f6226a1727316

Observation 312a9d9a-08e2-4680-8c09-50e62f97e871 · outbound

This paper cites Available: https://openreview.net/forum?id=B2Fqu7Y2cd.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Available: https://openreview.net/forum?id=B2Fqu7Y2cd

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-29T10:28:18.202974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:77c039c0c11ee56694037fe55803df5b64fa67dac1bd13e80b4b893d97298627

Observation 032ef048-05c2-49d6-bc39-abfb96283acf · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models,.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Chain-of-thought prompting elicits reasoning in large language models,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-29T10:28:18.202974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:c856d777590e9b9cd8165ca2d95eb05c5635d18a5667e4a28cc7084157d7bbb1

Observation a7d00847-865e-4617-89f8-b018de417282 · outbound

This paper cites Cot-vtm: Visual-to-music generation with chain-of-thought reasoning,.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Cot-vtm: Visual-to-music generation with chain-of-thought reasoning,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-29T10:28:18.202974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:90247e7a7074b14f83e4c94dcc30c13a0f02eec9ccf475d693f4e469c77d7f93

Observation a775bd8c-8635-4455-abaf-a63e67bd9e69 · outbound

This paper cites In: Pro- ceedings of the 33rd ACM International Conference on Multimedia (ACM MM).

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts In: Pro- ceedings of the 33rd ACM International Conference on Multimedia (ACM MM)

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:33:17.922875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:11c8cdcdfa419d96f5a9d443751735b82919248f2685991fbf60e42ba3b5ad9b

Observation 4f459bcf-c373-403e-8527-2132986c57b2 · outbound

This paper cites Ov-instructtts: Towards open-vocabulary instruct text-to-speech.arXiv preprint arXiv:2601.01459,.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Ov-instructtts: Towards open-vocabulary instruct text-to-speech.arXiv preprint arXiv:2601.01459,

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:17.895673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:2b1e3084b7c4cd79102c0fbaab9b204bacf182f570f54890e32674d85e6adf0b

Observation 4f9adc71-74b0-4299-bd3e-49e20a41e968 · outbound

This paper cites Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Scaling up Test-Time Compute with Latent Reasoning: A Recurrent Depth Approach

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T10:33:17.927585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:00fcbaf33c381ae8872d87f5bcd6684fd24c9b9e7461375aad634fb7b3703daf

Observation 69930641-b79b-4333-8aa6-b23ca2d6a1b4 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Training Large Language Models to Reason in a Continuous Latent Space

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:33:17.940347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:1fa1f0efa94101e3f55f6011bea1f3b2696d656f1e22bb9bc2898febaf32b0f1

Observation 53010181-7d25-4ccd-b174-76a5bbb84785 · outbound

This paper cites Reasoning beyond language: A comprehensive survey on latent chain-of-thought reasoning.arXiv preprint arXiv:2505.16782, 2025a.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Reasoning beyond language: A comprehensive survey on latent chain-of-thought reasoning.arXiv preprint arXiv:2505.16782, 2025a

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:17.922556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:16a1c02878c745e5a4ecc5d1ef46b93b46ae825f48005435aeae188216d9c63a

Observation a369a4ab-51fb-4b34-9c01-f58b2a0c16d6 · outbound

This paper cites Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Audio Flamingo 3: Advancing Audio Intelligence with Fully Open Large Audio Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-06-29T10:33:17.925822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:bebdd94c4a441bade56c8df0dda92742aff8c8a9c4c2b9b77df7ddb383bbc26e

Observation 6bb8d847-8a32-48fa-b849-63909e5dacac · outbound

This paper cites Gemmeke, Daniel P.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Gemmeke, Daniel P

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:33:17.909848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:b2d6f0887535557c2b370a16a6ba641cd51c33e78ca7c5f3267f84bead046ea3

Observation f1828948-1ee8-4a80-b90a-2366ec50d8d0 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Robust speech recognition via large-scale weak supervision,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-29T10:28:18.202974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:40b8a32639f431c2271e8ab28f6639248d35baa9fe2c5922207d293f6e6d95dc

Observation cf62ae6b-102a-437f-8fa0-7184b8bbf937 · outbound

This paper cites Simple and controllable music generation,.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Simple and controllable music generation,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-29T10:28:18.202974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:edd57976a0c224b11bcac7e1bc7233415954673b0de937ffbca5344027f60f65

Observation 08905fd2-4446-4e2b-ac4a-c5d44b33ca3a · outbound

This paper cites Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Text-to-Audio Generation using Instruction-Tuned LLM and Latent Diffusion Model

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-29T10:33:17.943293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:2cb824c9c38ee197b48365c2eb56730814dd2abbe2ff38b418cf350531c69dbb

Observation ec4ee320-d471-4e3b-aafb-48d439831d16 · outbound

This paper cites Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-29T10:28:18.202974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:d855d89ec1d7333b89e81d582e476f45926b81937a29cc09a62b1851b7b4c67d

Observation 77870d47-ec8f-40a9-8618-a66372ee6102 · outbound

This paper cites URL http://dx.doi.org/10.1109/ICASSP48485.2024.10447579.

Unified Synthesis of Compositional Speech and Sound from Free-Form Text Prompts URL http://dx.doi.org/10.1109/ICASSP48485.2024.10447579

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:33:17.925141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-29T10:28:18.202974Z digest=sha256:010771a9739f3781602e5088ffab3f7577c24eb08a5470f0b67c664cd5ff4e78

Pith citing papers

No inbound Pith citation observations are available.