Pith. sign in

Paper Citation Record · LEDGER

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

As of 22 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2506.09874.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09874 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:45:19.307090Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T21:12:19.893944Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:26:12.658368Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78962cc4-95b9-4726-bf81-1735024ca5ba · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.234441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.234441Z digest=sha256:7808f3c32b5a102c5362c71978463dee84e45c12010270f627fd13f432ba9395

Observation 41accd41-1833-4715-b628-8871c45e6645 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.687085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.687085Z digest=sha256:a027aaf6187d254909460def2f5869801e636500157330fb12b8d1876b91239c

Observation dd1ab871-4070-40d1-bc54-9ce02b7be156 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching AudioGen: Textually Guided Audio Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.835009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.835009Z digest=sha256:7d4c2cb585ac05ea00fbf746f37754e0e3e4402026fbcfee67a82de5a1be76d2

Observation 61ac979c-25d9-4ad5-aa99-8b313ae3ec8d · outbound

This paper cites Flow Matching for Generative Modeling.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Flow Matching for Generative Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.891032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.891032Z digest=sha256:fa4b46176e61d7d1171aa2d50903d897548f72cdc7199e30058a1d62067fb64c

Observation 87fd9334-847c-4c91-abcb-77ad8bf938a9 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.973851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.973851Z digest=sha256:408f97d3540fd2dfca2f39b984ab1125b7cba0557d37946c9ece8862bc85c011

Observation d4c66cb4-ace9-4a88-8fcf-9256a36ca013 · outbound

This paper cites FlowTSE: Target Speaker Extraction with Flow Matching.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching FlowTSE: Target Speaker Extraction with Flow Matching

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:45:19.443377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:45:19.031576Z digest=sha256:abd5f941bedeb5e8a7493c7c68063f7f6cf9130f11ac39fe9a971a190bfd3bc2

Observation e0cd1bd0-62d1-4a9c-8e8c-c581427ab987 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.164512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.164512Z digest=sha256:3003480756a8169b34063122a43799ad804851e143a712f256a4aebb4c95026e

Observation 397edf0a-3a63-4222-8951-32e8d2ce87a7 · outbound

This paper cites A Survey on Neural Speech Synthesis.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching A Survey on Neural Speech Synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.182281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.182281Z digest=sha256:cd166dc49e1163ecafe6ce87e7ad55752a18ec7627c5c9416f700ed31254e69d

Observation 32340c6e-8d6a-47fb-92eb-cd7e6ca085cd · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:19.578299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:45:19.283562Z digest=sha256:e86aa375067fd8cdb3d20d99e8b55e3b7d9194de5ec338cf12028360dd4ba92f

Observation ba40f52d-e387-48a7-9d1b-dcec9d791bdd · outbound

This paper cites A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.307090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.307090Z digest=sha256:8686d64b48e9811a30da53fffa086c4479f03bf3f71943970bfff840b27e39a7

Observation 453f421e-7868-45c0-af7c-3bc0739412a9 · outbound

This paper cites Enhancing Speech Intelligibility in Text-To-Speech Synthesis using Speaking Style Conversion.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Enhancing Speech Intelligibility in Text-To-Speech Synthesis using Speaking Style Conversion

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.104573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.104573Z digest=sha256:e1c19b0f84c15316777308ad8b2fca8148eb6d4cfa0ed3dea937c21e8a51b3bd

Observation 80e90982-39fd-4958-8789-b936820b4f61 · outbound

This paper cites Emilia: An extensive, multi- lingual, and diverse speech dataset for large-scale speech generation.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Emilia: An extensive, multi- lingual, and diverse speech dataset for large-scale speech generation

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:19.593624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:45:18.667628Z digest=sha256:f2b371b7218e4137893666cd99e710b493f48c0012a8fa6577ebac0e7831a83a

Observation 850f0b23-de79-4904-b426-71fde83fc81d · outbound

This paper cites Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.755238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.755238Z digest=sha256:a242078915d4e1496f212f2ad184b3f2cca22e491fbe27fcbbed3234dcf62975

Observation f960d625-ac9f-4934-ac86-60a071372484 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.608842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.608842Z digest=sha256:3fea5fc842703cba05a532636a0b9f29bd31abcfb754e8f7b19ffc88673fcef8

Observation 0176a76c-24a2-4bec-b123-ff19bb18f4bf · outbound

This paper cites VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis

Reference 2023

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:45:19.546105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:45:18.672137Z digest=sha256:0ec6340303d25d7f18aa889cc61eaeef386cdce28912ea1719f7cabd5c06723d

Observation 7b16013c-daf3-4120-bcbe-ffb91856b964 · outbound

This paper cites E., Wang, X., Thakker, M., Li, C., Tsai, C.-H., Xiao, Z., Yang, H., Zhu, Z., Tang, M., Tan, X., et al.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching E., Wang, X., Thakker, M., Li, C., Tsai, C.-H., Xiao, Z., Yang, H., Zhu, Z., Tang, M., Tan, X., et al

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:19.608252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:45:18.647201Z digest=sha256:7bc07cb92360c99c7878c7b36a845df044edff07f8e86b3e3e0fdc001cb618a6

Pith citing papers

Observation 48b72c44-7294-484b-9ac6-93694d3d2945 · inbound

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment cites this paper.

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:12.661007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T21:12:19.893944Z digest=sha256:df192853624fffcde07754348cece1ba3a3352d1acfb5fdc2af5a3c80c7253f3