Pith. sign in

Paper Citation Record · LEDGER

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

As of 6 August 2026, this Paper Citation Record lists 100 of 128 outbound references and 63 inbound Pith citation observations for arXiv:2410.06885.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.06885 v3

Coverage vector

measured 100 of 128 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T06:06:41.348728Z

measured 163 of 163 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 63 of 63 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T06:04:29.398373Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T18:15:21.415875Z

Reference resolution

100 of 128 outbound references displayed

  • verified exact25
  • verified fuzzy26
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 184cfbaf-e39d-48c6-83f9-6f6eff009aea · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.621101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:03aeb6e41a0798736b99ba7b3f54f4daca21a791e55d3ce2dff1aaa991d68ced

Observation 382fd074-effa-4c31-9af4-e448e24c724a · outbound

This paper cites International Conference on Machine Learning , pages=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching International Conference on Machine Learning , pages=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.689945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:13be939d93c666c2e4a7f51df8947e300bea7a91c6f57d01aceed1b4ce683500

Observation e4fc6877-2055-4306-bbe2-c7caf6cbca81 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Advances in Neural Information Processing Systems , volume=

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.692626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:c24f59a745504e1ba7257208a85c24d3b039434ed8c90b888e794620b17f7331

Observation 8abdeed1-e280-44ca-a245-f800b747dad0 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.695104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:e2feb58a2a1c73c1a27e91c00968be63c5f12e1f636b028f65fd4653ffddf6b8

Observation e917f49c-d9aa-40ee-8f05-c23cacc64b03 · outbound

This paper cites 2023 , organization=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching 2023 , organization=

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.697698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:4fb437cfa1221716a41953973472a6e1446d4b300a3e78561788b2edd34dcaf0

Observation 056f0b72-4f3f-4b58-8495-87fd015192c8 · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.700556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:b5bce90201fb80ceab344b9a041689e58ee1d3cb026bbd958477329616dc2831

Observation 2cb7ccd3-b14c-44fa-ac05-21b2498c54ec · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.703242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:fdbe0d003cace25a4e029a035f6cebcafbcd8dfed5ef00cdb11b0ed59fa0076c

Observation 7dfcc645-cf50-47d9-ba3c-7cdca229eb9e · outbound

This paper cites 2011 IEEE international conference on acoustics, speech and signal processing (ICASSP) , pages=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching 2011 IEEE international conference on acoustics, speech and signal processing (ICASSP) , pages=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.706882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:6a5071568ae5afee780407ef52ba424025b653a60375ff1f0150c0c7fa476a4b

Observation 17cf1d67-ab72-411a-88a0-2b24f112b430 · outbound

This paper cites International conference on machine learning , pages=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching International conference on machine learning , pages=

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.710190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:1995afdf082abdee637427b143bc36dd190c88ab98d53139180df9be3a0a1e32

Observation 39698f26-7f55-4daf-a145-f8b823fab942 · outbound

This paper cites Libri-light: A benchmark for.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Libri-light: A benchmark for

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.712460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:c13aae174c774a4f7b450392cfc1c76bfe9b7bab27ed114fe1189ffa2a19e6b7

Observation 135f4a0e-01c2-4d0f-b501-ced8e810d3a4 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.714908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:cbb642a89917dbbd35918d190e9c5dbe3b2e9d32b5b384bb6bbccf06cf86fe05

Observation cfe835e8-e1b0-420e-a0b4-e39d268909fb · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.717497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:2c17f60aa05c430721a7155b959f99be8854c75530e7e6c93362a2ac9a5b8b35

Observation 49df7fdf-d8d3-4516-89cc-b724c44ebe24 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.719930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:82ff7845c0e66faadfa54bad1b7592f4996a8e8c05ef4b0738d480af605066e8

Observation 7e59a5f8-bd05-402d-ab2c-8d32986d0b71 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.722469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:a5374ee456bfc510799d249600f43cc405bc7c95001e2eb3373ed206f0a33bb9

Observation e074e560-a14d-4148-80f0-018200c97e26 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.724877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:1bb3556d07c12a51841ff0b5da6c42cbd8ac2941d09bdc0dd1058b0fbf453831

Observation f084290c-b028-468b-a07a-085dc9dee7af · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.727302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:a0248dba5ab7883317b2cddbbb6ae239fd28c2366e99a6988256627b27b74a79

Observation fa4cbe13-13c7-4969-9e93-c12ef1895dea · outbound

This paper cites Librispeech: an.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Librispeech: an

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.729514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:f4264e863bb560e8d04273cd6ff898d879a731d3e68106de37ce835072c327fd

Observation b71c7846-79c1-4cf1-b29a-bf15c7334010 · outbound

This paper cites Didispeech: A large scale.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Didispeech: A large scale

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.731705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:8ecce044f0c90e2a9f17cc44273083e1a0451292351d46bf540ca016d118afc7

Observation 45b01911-5853-436d-8bf3-93e61038a9ab · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.733867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:ce398a95177b40ad149c2fe1dd19e29ba5d2e189b2fba276cb882079ee42eea4

Observation 62dce785-d519-44eb-8282-43ff8a5411dc · outbound

This paper cites Biometrika , volume=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Biometrika , volume=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.736130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:d142576f8d0a03538d24c3526037f3ac7b16c4b71b2f1922dc1ea081907a2e5b

Observation ae940ce3-e756-42e4-8b3a-a1ec40ed3592 · outbound

This paper cites Neurocomputing , volume=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Neurocomputing , volume=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.738636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:39bf64005974548da153b7050351fab020c8f9c92b2f4770ceaf81a8ab7bf324

Observation 37f607f3-44bf-462c-92b5-4a4dca306cca · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.741422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:9edff3eb489c4f069f297af3084ade191ec389850267d93f34c7b62b6ac7f768

Observation e15fe887-ef0e-40cb-9282-9abb977f74b5 · outbound

This paper cites 2024 , organization=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching 2024 , organization=

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.743792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:764e0916771bd0b1cd0eca3234d229d4a7ef559199e0679aed5dc58dd76e443c

Observation 05122b77-7a9b-4555-93ce-0af5007b8c7a · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.746045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:0c379a31db127c4e4141b7cefec32f288e9aa31686c54d051a44119a642784e4

Observation 770432b2-838d-4cdf-803b-9fc64a7b259d · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.748329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:9295b51b7dcd7fd5383ecc921f14c36f8bb3f0208d5b391aeb846e792bd8fdbb

Observation 138b71b1-ec08-46f3-80cf-9f2b003aa636 · outbound

This paper cites Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.751225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:9cc8e7197b0a2a73aca91034f92480ffd4af5b3b8cc6565b398b596b735b65f2

Observation 31ef130b-2706-4813-829b-00db111ff3aa · outbound

This paper cites Advances in neural information processing systems , volume=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Advances in neural information processing systems , volume=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.753787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:eed9125b782857662326468b208437b5039903074382330ecc723432a774b7ae

Observation 26ef8c1e-1280-469b-8d4b-4c349aea037b · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.755955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:12057c4a946d20c85fd5807c77ba48747c49b92a767e013b1a6d0df32b25eb7f

Observation 0d87d880-78a0-4a57-8e8e-9750f59de282 · outbound

This paper cites Understanding diffusion objectives as the.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Understanding diffusion objectives as the

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.758285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:3295d210112516c02ba94cd6d2f9f3acc91c867ea3460f40f185da02d63c3f2b

Observation 411ac7fa-a3e5-430c-8ea6-16f88d4466dc · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.760460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:8c8ec4d5bf42ec9868f80f7179ec5482614f11e3ea6f4ffc41d0f34e97222693

Observation 82fbe785-5f99-4d63-9918-b609b8e65531 · outbound

This paper cites 2023 , organization=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching 2023 , organization=

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.762685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:3b404f11c8eabe3de178aac9bb99dfb3975239a2fc7443ad6aa67c1cf6167837

Observation 2f06da8a-d704-4014-94e6-3a8e4d9f3db4 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.764958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:0d1bb5b4de278cbef5e91e54812ca4b1f68a3e697aa0bb8b74b184f871461c49

Observation d194696d-76e5-4924-b7a8-606dc2911204 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.767403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:5d26e1082f32b45cab6528f050bf0b1fb1e31ddd92e5090967464578e68c6158

Observation 0bcea5b8-4d79-4284-a15f-dff3fff7690e · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.769592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:4147f157465a1f7d28a3ec893c9f929289c57cb5d607d51d3b265f3808075367

Observation 92ff2a18-647b-443b-815b-43bc75fab998 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.771813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:52d5db56049543cb9b15700f3e781ea7ae4ce75673ddc3d11effdfebdb6be645

Observation 020b22f3-fab8-4e07-a8f2-38449abeea6d · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.773970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:a4acd0aecd4a1a1d43914fa5c871020e440887e4cda8d0025548d184213fe47f

Observation 16c00806-af5b-4d14-b8ef-94440a3eac18 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.776463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:67ada409cad9902a0f636304eba3adbd3a0fa647a7fed124e668440bee09fcf7

Observation 26fcc227-5de7-4684-9e4b-5c7e426ab52e · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.779064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:94178a7fca3151b9b81936723193df2332d895ad85aeecfb82f6a5c4b199ab01

Observation a8ca92a8-3162-4f1d-b98f-8b6e6dd9887b · outbound

This paper cites Advances in neural information processing systems , volume=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Advances in neural information processing systems , volume=

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.781689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:2bb36c88ce2f1d8d7da72f1e98c094b4dec8f5f7e8fb2e87d6fefbd6e3d83261

Observation 6761d9de-d2c6-434a-beca-def5d7147d96 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.783760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:eb7407248a821104d0d859603219c89d09ac393b0d0ca343f3e7d7b54ce9b64e

Observation f2809e8c-20f6-40f2-9dc5-5c41d97b3318 · outbound

This paper cites IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching IEEE/ACM Transactions on Audio, Speech, and Language Processing , volume=

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.786425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:fa7fbf9bad4a5af18c4004497028416e8324c7ec949b263cb5d96f5b4156b726

Observation 83a51553-10c6-42cf-8c5d-aa8583ff4597 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.788918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:8b4e1112d44c7e8f90d22f8ea3f08b0f33d0eeaa563bfcf8e08ceff4085abf45

Observation a16b1765-db14-4c51-b69d-5dae4b513d03 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.572451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:b4c17c8ea8d4d77e167db9873f97c436638c9b43e8343c5233c35a59782fffa1

Observation ec2f71ab-6c37-47e2-b0a6-a1cc055b37aa · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 66

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.574606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:9f9880e063c0ccf5631322d850342803854ec4facea211cbb2ab9044d929988d

Observation 00f85ccd-5fa9-412d-9468-55b2ddff56ab · outbound

This paper cites Advances in neural information processing systems , volume=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Advances in neural information processing systems , volume=

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.576800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:81de8a2370001691577a2afbae0b3bc8b3f82208047d4f019e4ff77a5e0e4aec

Observation 8cddbb06-1171-4eae-a20b-53a99ecf5a04 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.579005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:e33655b6256a03cc71a11f3057b6e03bde156442457cc77c6f9098ad9590af29

Observation e5b8928c-0cf0-4d15-92ce-8aed9ad4c250 · outbound

This paper cites International Conference on Machine Learning , pages=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching International Conference on Machine Learning , pages=

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.581164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:085f48e1e8011a980a314acd034534c3c18a26e4abd9c1448af7e01917cf4fdd

Observation aa1bd310-0371-41ee-b7c9-6b4a5968e9d9 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.583533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:17a24226ec91ce1b7053236a695558ffdbe1e7bafa65e867965a8eae665771af

Observation 8c546d79-2c3b-48b3-bbff-1a5dc384d6b6 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.585504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:4584deeae1a4bbeae904c8a4a89dcebe5f55dde1534ad948b4cfaa34f0d76281

Observation d95f7b5e-f10c-4fc4-b773-84ef826e13d4 · outbound

This paper cites International Conference on Machine Learning , pages=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching International Conference on Machine Learning , pages=

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.587791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:5bcbe4d98a6b741747a8a755b969d6c9b33007b42d46157bb4c429febf598537

Observation ff238eba-802b-4d79-90ec-d0a0f4004ea4 · outbound

This paper cites Proceedings of the AAAI conference on artificial intelligence , volume=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Proceedings of the AAAI conference on artificial intelligence , volume=

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.590030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:b0bf5f8a05b372572319c89c8c9fc6c890aa39ea28c7aefe423560174152a551

Observation b4f434d5-f068-4806-98d4-9ecfd0cf3bd2 · outbound

This paper cites IEEE Transactions on Pattern Analysis and Machine Intelligence , year=.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching IEEE Transactions on Pattern Analysis and Machine Intelligence , year=

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T06:06:41.592109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:dcb5353db64446b462a515d9341598b9255f2516bbf2e0ef5469a4d927184c01

Observation 6000b29f-4584-435c-9497-184926710ea2 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:06:41.513885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:88358971ef7f6d49a058a914547b8e031de41d9e34efdca2a18b446e184a41e2

Observation 0193b1ef-cf8c-4438-bfde-1bb999c3fee4 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Common Voice: A Massively-Multilingual Speech Corpus

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.464496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:e83ef3cb4029a1437e2fcbeeba8498ba45f2b6a6ee2981929b76a5a2fa53a243

Observation a50f5152-1a8d-4454-8575-9085bb31d957 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.594217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:4e2d394424c29fe9fdc37e951cd746d609040199694cec1cee49507c96995265

Observation bd46321d-de84-432d-9afc-2b9df90a991f · outbound

This paper cites dMel: Speech Tokenization made Simple.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching dMel: Speech Tokenization made Simple

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.561300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:6bde974ab2361392152526ebba198e48e478d25e108125c10e20c28d2bbd78dd

Observation 8392d6dd-47e2-4862-9973-e4996aab7dcb · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 84

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.596322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:96495d6bebd5f761765b87edf7498055c555a0cd9ca41912c2334e27a0c80f17

Observation 73616e53-bf2f-45a7-b1ba-18f39d2029b0 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 85

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.598357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:d212f2c44bdac418f59d965b2242219f007aa6945692db023b6555da45c37391

Observation 251b1e65-964d-4550-9ac4-6742880289fb · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.535511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:039ad3f930b82b2ab86fad1f96cf0dbc9839e7c2b5e091975342f67c03888de8

Observation 990c6a6c-abb9-401e-b739-6751c5e2b797 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 87

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.600249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:71189a886d725ccc6a9417e8c60126202bd20d8dcc08277caf0856655e448527

Observation 6540a55b-bd22-4fc7-a56c-1406a6e6ec8e · outbound

This paper cites High Fidelity Neural Audio Compression.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching High Fidelity Neural Audio Compression

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:06:41.424218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:ba16a35838f2acdea803e8686fcf8e482b0f31da66cc2e959d454c9dda400028

Observation d0dfb163-bede-4b44-acf5-25899bfda549 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 89

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.602094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:5b3a47aef1ce8d0ae1879741a9d90a44c7f22ebc7fe6ffc847c0269c534b060e

Observation 0e2b1200-720a-4f80-9466-db9ffa8ffb45 · outbound

This paper cites VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.479977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:ea46274d522fbec2875093c04ecb76c58973fc5721e4e0a908b49e90efd77fd7

Observation 3593ad95-210e-44da-9a38-a0c5fac028af · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 91

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:06:41.487270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:b0c51e3388ea3e0eda1bc993b7246a749ce4d9fc0c85ac6eb6ad19a41df2c7bf

Observation f9b72dbf-3421-4be2-88ea-9ebd7141c765 · outbound

This paper cites E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.502207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:8d9a7395127aa4cc067751c7aee05fbbd6cbf1ed45009ae8756b5ecdc51b7001

Observation 76e26e0e-c924-4807-a742-e30414002df5 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 93

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.603949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:5414293265e4ab47dce6de2e1258f155323361775a2ccc92a565b3074f752015

Observation 55cffd1f-1a25-4931-86e6-90ce62e2b721 · outbound

This paper cites FLUX that Plays Music.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching FLUX that Plays Music

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.532023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:63fc34033d1b8d4258103ead1c51b24a61c839e8a00e920871383be1ce394764

Observation 39e0d087-6075-4aad-beaa-316ec414676f · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 95

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.605786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:c799c67f6ed92778367d556ffad2d431a2486efd767161153a6b1fc2c12fe4c9

Observation f19d96cf-cee9-4b9d-a175-f7cc473c622b · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.549207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:a3e2370b0081b4c93327c25de8421ee229186e92d8c2daf10c32b21698b9ee40

Observation d20dfc8c-8fda-41d8-81f3-2483d87ab40e · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.552278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:bc6abaaeae925c4e26f0face03c81ce0607c9d87e4a049b53e332c2c4d8a67ec

Observation d842a766-2845-4756-8cb9-753c6aa29881 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 98

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.607964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:013a2108e8ae9fc63568587bc5e1cc70bcbd4ce93356c9289c90e8c5ba626b3b

Observation 5dc1fe4d-63f6-412b-88ee-90cb657f3cde · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 99

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.609832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:22a9948cd6b85162f2a9ea85976bc669b59dc1d50402d7dae9ecae255e96bdd9

Observation 50e7f70c-2504-4599-a672-a57d5f920719 · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.450210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:6daf73f3ce8d29a0bb1229623d2d1e218f32e6cb50bb0480e81ffd3011cb159a

Observation 0847645e-1357-4457-be39-54962b7a6d8c · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 101

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.461429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:32486091eee53a565c453e7795fce75d3b2281b312d6f0c7cbaa1b12ca95a90f

Observation 90f4b4a3-c070-4b6e-a7af-7c7cb611bbc9 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 102

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.611821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:b6ce52e0b65fcd7982763e3220c8e3c2f5e5b6108da71691209c182aa2713fbc

Observation 566aa25a-8a79-4cf7-a0bc-13dd147b3e14 · outbound

This paper cites Classifier-Free Diffusion Guidance.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Classifier-Free Diffusion Guidance

Reference 103

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:06:41.475519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:46892b72c7cb283511191b3870366944cbbf4f2b24246762b00002769a9e6b47

Observation f32a933f-30a4-4a78-922c-264139a17116 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 104

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.613796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:84d5192ef133679c8100f2c270a1d7c942b825af5d69de9174e2d9175ef68624

Observation ebbb30ca-95a3-4d9e-a81a-06451c8c585b · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 105

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.616081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:643fac7edab8d41ef133a1d4dad564079c59a437fa5e8f39b2d5ba9d47b88a77

Observation 8d16dd5c-9eaf-459a-abd1-6dc409c68dfe · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.494637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:ff1aa3ceaf3d409aafa371f0c08ae5c2592d299b7616935c0b75510f9e3d05a9

Observation f73d03f8-4924-44a8-8e3b-917d82665098 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 107

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.618130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:14d7e3839ff73653e85e10f154be15ebb0e3b7aa2a69a51875c6cce19ab2c51f

Observation 62abbae3-ad0e-4124-8f4f-33644ca81839 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 108

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.687869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:cf0685fb0f9592ecece0f7ba2b81a0af2c51aff8eb8683eb5e31ae341f0fe5b2

Observation 5bf03bf3-f29d-4228-b0d0-20fa8b84e7f8 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 109

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.636734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:df1234ba4bb8aeedcd8cf06aa936d8312b9ed72407a0ac2f099414c480071cb3

Observation 81efa80b-8c0e-4e10-9ea8-69379bd77a1c · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 110

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.639483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:b1185be66a8094cfb974636a62de665e6c46f17570df0e47aff8a7292c428500

Observation 029a410c-f93a-49be-80d4-6883e245f025 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 111

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.642324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:c454cb31f9baa9ad68690a598986157e7ddc733f2178cae4fbe7537e3d9bd17b

Observation 61509622-b74e-468f-990e-15b303224801 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 112

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.644737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:a6d9c5e988c0cfef29d718e1120b77d3d93569cc7ff30c4f98fb4be32192a395

Observation 4d3017f8-b17b-4dea-b9e5-562add6d2a9a · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 113

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.647121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:4dc0a96127bb427a1e9c35e066a8d6fe6193fb1cd2f03b11e785ad0b602450da

Observation de99c861-4fa9-4af9-b871-0aa3d7d9c1e7 · outbound

This paper cites DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.413952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:fd0054dc702fee7bcb655ecc5b74c5eefcd4a9af30c64091defae859a950da65

Observation 2739dbb6-bbdd-4204-ada2-6c2aa237b80e · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.419455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:19955d1dda9e508c1e7610e063d18b5eee319ad159bdfcb23735a90e56145e82

Observation 0ed13fdd-c2d6-4ba0-a0fe-e40d3e846436 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 116

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.649322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:e02fc7d4e2bac578cdf5f2b24b919e5fb3ac3cdbe0f61c6470e8cd2b4b8fb1b9

Observation 125ddf1a-becc-4768-b27d-ad19755f2a1c · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Autoregressive Image Generation without Vector Quantization

Reference 117

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.430014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:05dcf0330077add91c78d0e754614203a4cbe6b5ce9829d5efdcd8dd51327fc0

Observation 331e86a3-f3da-4ae3-9738-c8e8b7bd7894 · outbound

This paper cites Flow Matching for Generative Modeling.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Flow Matching for Generative Modeling

Reference 118

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:06:41.435065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:f66b04efd8ec32f61600ec6c8e4fe028e44f666bb0456344ede5daf9bdee2b29

Observation a37e1d39-7169-43b7-b78e-807172d6bac5 · outbound

This paper cites Autoregressive Diffusion Transformer for Text-to-Speech Synthesis.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 119

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.440978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:2ff3d623aa2ae1d7fb98bdacabb42a552f067f9e3dfdecf7a4407d9dd619dec9

Observation 75cd1411-76c2-4917-a511-a2d538ac4d89 · outbound

This paper cites E1 TTS: Simple and Fast Non-Autoregressive TTS.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching E1 TTS: Simple and Fast Non-Autoregressive TTS

Reference 120

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.445688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:d6c9888170f7f1e6faa30349e6e6d258ac0c9af515057797b6f2d2260dde355f

Observation 0f0822c1-aaa1-42e3-9779-3c74a398ece2 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 121

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.651636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:3ef78caaa837c9d815dcdf8c34cf4978db3ebaa20a5f3a7231541ff02187d23d

Observation c01e52fb-c86c-41d3-9724-b39fa8213131 · outbound

This paper cites Decoupled Weight Decay Regularization.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Decoupled Weight Decay Regularization

Reference 122

Resolution
verified exact
local_arxiv, observed 2026-05-16T06:06:41.454756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:d7b8a80dd251a8c31487f0c0e75432f9aed5747b16644425c6b3d3ab951eed87

Observation 138ad91b-1082-4925-afe0-dcdaae9df868 · outbound

This paper cites WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 123

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.458352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:c7875025000c861797c2065171b8f40a9eeeb19703b3b7c5238dc801aa8071bd

Observation 75a063f3-3f23-413c-ab19-d24643d01d78 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 124

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.653865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:46afd6105488ae0ad9f44940a6fc638ab06623cc474f6a85485da4861b900f12

Observation 5c4f6778-c367-4528-9c16-244c161e9833 · outbound

This paper cites an unresolved cited work.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Unresolved cited work

Reference 125

Resolution
unresolved
raw_fallback, observed 2026-05-16T06:06:41.656334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:f90e96d1ff59fbdcc5c7582f1f321face57271df4df506abdfec64923d534e62

Observation e06e1dfd-3074-4fde-a7d2-ea6c240bb9b7 · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Autoregressive Speech Synthesis without Vector Quantization

Reference 126

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.468405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:eebc7466e66a42579085a3c479d0e72f6c05d7298ab3bf68638a6da04d4e931f

Observation 1686f71e-2a9a-4f76-b18f-d723e9cd13db · outbound

This paper cites NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching NDVQ: Robust Neural Audio Codec with Normal Distribution-Based Vector Quantization

Reference 127

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.471859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:326b602914e5930b79e3107c9308fdf11ff081cbbf95a1506ece1c3c74d2a1fc

Pith citing papers

Observation 6c87e634-1e3f-4c4c-abf6-a5802d28e46d · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:00cd3f9255430e70d85fe69dd45ba9b2ad4b5c3514e99c75e9766a23394b6b7d

Observation 6c6c408e-c345-4eb9-acb5-498095597e23 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:b1db482527852468a059dc732adc8e6fb8e8c898f3ec2da4f3ccd7dac02f5f57

Observation 79a70a04-4116-494e-9f57-3e4dbf87a9df · inbound

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching cites this paper.

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T04:32:03.653380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T04:29:41.285194Z digest=sha256:16576d85c993e3742b7b875be7f0d06d85cc51569ca57c6c812ac565043e993a

Observation a9f25310-eeaf-42b5-a2f9-e18080661d87 · inbound

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation cites this paper.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T06:04:29.398373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T06:04:29.398373Z digest=sha256:5b33a4e7849457a87ede81000aa3f1b64bbabf617d0a6e52daed4814b1385adf

Observation 68fb7631-fa48-4740-92ed-bf2752351d82 · inbound

REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers cites this paper.

REF-VC: Robust, Expressive and Fast Zero-Shot Voice Conversion with Diffusion Transformers F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T23:41:51.090161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:41:51.090161Z digest=sha256:402f2fc3c4f3497648a8ce08ddb707937d931e35e05b68aef3e57801e4b91504

Observation a296521a-1a36-4701-984e-94ed130ce88a · inbound

Transient Noise Removal via Diffusion-based Speech Inpainting cites this paper.

Transient Noise Removal via Diffusion-based Speech Inpainting F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T21:23:56.183003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:23:56.183003Z digest=sha256:8bc1c01ff05234da18fe9998f4b84a5fa7c8ff862509032671004e1d0a5dd5b2

Observation 8a8a0fbf-90ba-4799-b43e-99a25a9b2dc5 · inbound

ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs cites this paper.

ProMode: A Speech Prosody Model Conditioned on Acoustic and Textual Inputs F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T21:07:03.845459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:07:03.845459Z digest=sha256:de4661ade354399dc63d25d0d2766b420789aa8efc728374caf94a1e61a08158

Observation 745110e8-993f-4339-8a4b-51e03547cfbb · inbound

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis cites this paper.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.541441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.541441Z digest=sha256:abb75e71621cdf053769bb293dea40685e44d2ca449aa3fbd5001e2607a0403b

Observation c30e0ee2-aebe-45ec-8876-4811e67e03ed · inbound

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction cites this paper.

DRASP: A Dual-Resolution Attentive Statistics Pooling Framework for Automatic MOS Prediction F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:22:47.082941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:22:47.082941Z digest=sha256:97c0465749ace6dc1ec32272c2dafaf2039782bd4f8aff327fb563ed5d97930c

Observation cc062fb2-9d13-477c-897c-c469a86beab9 · inbound

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot cites this paper.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.309565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.309565Z digest=sha256:fe6b12824d56e6629041ffab5ee2e8d8bfbdf6a143146ff762140c3aae8cd5dc

Observation 4a78d0d3-ffe7-4e6b-ace3-6a0c7ed7d7a9 · inbound

Computational Narrative Understanding for Expressive Text-to-Speech cites this paper.

Computational Narrative Understanding for Expressive Text-to-Speech F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-18T19:31:47.040481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T19:30:53.250247Z digest=sha256:7fb8ac4bd25cca60c9d5eb1dc95bb2ae335f3854f473340640c36a1cb8af9e29

Observation 7a4f428d-ff2b-48bb-8435-8c428573856f · inbound

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching cites this paper.

DiFlow-TTS: Compact and Low-Latency Zero-Shot Text-to-Speech with Discrete Flow Matching F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T18:51:21.004855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T18:51:21.004855Z digest=sha256:3dfb6300c4a7396a3a07349252a12feb4ad011719a9f3a4456ee0ae465e82507

Observation ed3f9ca8-e5be-4851-8410-9a44505abc26 · inbound

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration cites this paper.

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T19:11:59.493517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:11:59.493517Z digest=sha256:1262986462abf37102ef9b8ac5e249656c99805c8d86690d6608aaafb8cd5edf

Observation 8abe8bc2-12fa-4844-b473-5600f0f41221 · inbound

Length-Aware Rotary Position Embedding for Text-Speech Alignment cites this paper.

Length-Aware Rotary Position Embedding for Text-Speech Alignment F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T17:10:04.032952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T17:10:04.032952Z digest=sha256:a26fa90256f83df0d329b896b8a6ddbe5bf8ef63f2b5b471d02dc2a2cc7a526b

Observation 4f640102-5e34-46e4-8537-08aaf241db06 · inbound

OLaPh: Optimal Language Phonemizer cites this paper.

OLaPh: Optimal Language Phonemizer F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-18T14:26:28.197858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T14:24:51.262622Z digest=sha256:b84209a62a6aa1d5aac9f3e0c48f7f711574f3272c24bea90705b36714422b2e

Observation 33ab5622-c8eb-4cf2-a821-eedc99c37674 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:30.853473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:30.853473Z digest=sha256:415b162fef5852adfde3b282d505f53768429880db9df6419deee6a976b33155

Observation 625e4c1e-9328-459d-a0f0-3e4a73d85bb5 · inbound

EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection cites this paper.

EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:02:23.389284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T05:01:57.044869Z digest=sha256:0ef1e33c62c7fa3aa5f2f7aede54b9cd884c05681725cceaca9aec9aff4921c7

Observation 6802849f-fdff-40e8-99d3-261260535888 · inbound

Qwen3-TTS Technical Report cites this paper.

Qwen3-TTS Technical Report F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:24:56.082402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T19:24:56.057631Z digest=sha256:3869c9cd1b64c1441cebb41b6cdcde7b8c5e20d450cce69fede04ade15f60500

Observation 4f30ab23-c136-4ca9-a2b3-42d2b2e93d6e · inbound

DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching cites this paper.

DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature Caching F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:30:44.408615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T07:28:08.873654Z digest=sha256:f79c4b261eaaaa43d70e727bb7f32b16b9f7cdc517e87af3937aaccd8f69e4db

Observation 28100a7f-b896-488a-9ba6-cfc2b263102e · inbound

Momentum Guidance: Plug-and-Play Guidance for Flow Models cites this paper.

Momentum Guidance: Plug-and-Play Guidance for Flow Models F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T21:26:58.384960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:26:58.384960Z digest=sha256:77d071cbd45507af4c733fa30048ae660c91a99479995df4447691c0dcd061c4

Observation 37f5b00f-d071-43bc-a879-d03417914af0 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:0aab0de7218f5817fd54c935e92ab4842a1bac841a2524c27d3a3aead8315094

Observation fa5d5a27-28a4-4ba2-85c8-85e6b23575d1 · inbound

Evaluating Generalization and Robustness in Russian Anti-Spoofing: The RuASD Initiative cites this paper.

Evaluating Generalization and Robustness in Russian Anti-Spoofing: The RuASD Initiative F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T02:27:05.776290Z digest=sha256:185675f58e4b3e79b608695dd5f90b550d36b7dd72169dcb477ac1b5f6ab4ea7

Observation bf142335-2d34-4897-b53f-351548ac5e5b · inbound

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation cites this paper.

CapTalk: Unified Voice Design for Single-Utterance and Dialogue Speech Generation F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:15:28.918204Z digest=sha256:5e92fc11ddb78b7b40ff2aed4d7c2ce380dc56d7a93b9d56d06c57585d2cbed3

Observation 95ab61d6-9068-438f-ac13-cf519b90c0db · inbound

ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing cites this paper.

ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T16:03:15.572657Z digest=sha256:e7a26721b75d6564d0a2652d7f7325ef66dca51832dc2921c3a8971db55a3a9a

Observation fddd77c5-97cf-4285-a811-33f2c59fbeef · inbound

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing cites this paper.

CoSyncDiT: Cognitive Synchronous Diffusion Transformer for Movie Dubbing F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:46:58.010112Z digest=sha256:210fc7c76cbd31d788566a15e078b69fadbb9cb19230475dba9b8d10e06d79f0

Observation 31716ffd-6c08-48dd-8ef0-3fc45e234fff · inbound

ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks cites this paper.

ProSDD: Learning Prosodic Representations for Speech Deepfake Detection against Expressive and Emotional Attacks F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T13:31:12.802897Z digest=sha256:f68eedcadb16b3a053761ab8ebd2f9b9f3898071149eb8bd7ce3bbfcd29c7788

Observation 44a8a20f-d1f7-4db4-8ac0-a524fcd1c024 · inbound

ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis cites this paper.

ATRIE: Adaptive Tuning for Robust Inference and Emotion in Persona-Driven Speech Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:11:12.460695Z digest=sha256:85d94906dc8cee58528676c501191965cb5108e1ed9d5ae470ddbe9f64d4c718

Observation 66f91db6-cd7d-4f8c-b658-9046eec7917c · inbound

MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control cites this paper.

MAGIC-TTS: Fine-Grained Controllable Speech Synthesis with Explicit Local Duration and Pause Control F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T13:49:09.386368Z digest=sha256:e93370156f0794cf788416d862c8811748570f184a09b7fc19205c54b5eb9528

Observation 5efca61d-fc14-483f-a8a5-b570aff7a329 · inbound

Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages cites this paper.

Preferences of a Voice-First Nation: Large-Scale Pairwise Evaluation and Preference Analysis for TTS in Indian Languages F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T22:07:47.847453Z digest=sha256:3f18400b38d761df7761e7b5649f85e2e27b11ac70038cd0f1bc57aae27e8a3d

Observation 7d49c06d-cd33-45e5-b390-cde1a62975b1 · inbound

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions cites this paper.

UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text Instructions F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T09:18:04.870414Z digest=sha256:9f9fb258528fe97f0c85666b367b90b6becc45099465634d47d0242e27e5e9eb

Observation 3f88102a-d3f8-4e2b-bea1-dc1833fd02e2 · inbound

Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost cites this paper.

Praxy Voice: Voice-Prompt Recovery + BUPS for Commercial-Class Indic TTS from a Frozen Non-Indic Base at Zero Commercial-Training-Data Cost F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-07T14:31:27.333704Z digest=sha256:f4332f1b6bfe01d5d9dad6562b6010c5392eec6c790b50cc47f0197294d82fca

Observation 26b8b0d4-5e92-495c-87f2-f791f2106e7d · inbound

Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment cites this paper.

Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T01:07:50.243903Z digest=sha256:6cd54648c1df6aafaeea441e0cfdd49c42338902a43b65bf2f2aded709c491b7

Observation 1ea04f1a-b303-4ce7-81a4-4590c918ac7c · inbound

Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation cites this paper.

Tibetan-TTS:Low-Resource Tibetan Speech Synthesis with Large Model Adaptation F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T02:59:39.588729Z digest=sha256:cbad1dade76e81daa070ca99a3f411a02c3f5d122335dec924201d764e3d43fa

Observation e09260d3-9114-4559-bf35-bcfada59cc96 · inbound

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech cites this paper.

Kinetic-Optimal Scheduling with Moment Correction for Metric-Induced Discrete Flow Matching in Zero-Shot Text-to-Speech F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T03:04:44.050889Z digest=sha256:b52520c91afd9986f02eb3cae02b40cb5213f1339426919a47ddbdc2cce9acea

Observation ffe31654-4602-4adc-97d5-69d55aef05a1 · inbound

Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling cites this paper.

Poly-SVC: Polyphony-Aware Singing Voice Conversion with Harmonic Modeling F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T04:09:48.737629Z digest=sha256:fdd5708631dc05d7e1edd55a95859b221e1373226695891bf4ca35df1290f10f

Observation 7db77c5a-aa51-45d0-ac6d-f6fc51618eb5 · inbound

Break-the-Beat! Controllable MIDI-to-Drum Audio Synthesis cites this paper.

Break-the-Beat! Controllable MIDI-to-Drum Audio Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.789827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T01:28:16.995989Z digest=sha256:51ef0608245e3ac0e27fd8733ad9fd3eaeb543cc479511c1b3bfeb3308e45843

Observation cc24fda1-0762-4c78-853a-83af54cecc58 · inbound

Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues cites this paper.

Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic Cues F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-05-19T18:37:42.801214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-19T18:36:34.351533Z digest=sha256:77a37fdbd812dc51d4bac863f30729d8874ea430cd97d6e60894212ed8ad10c6

Observation c1ba00f5-d7b0-4a96-86ff-54f2c60366c4 · inbound

Natural Yet Challenging to Detect: Robust In-the-Wild TTS through EMA and Dual-Scoring Prompt Selection -- Submission for WildSpoof 2026 TTS Track cites this paper.

Natural Yet Challenging to Detect: Robust In-the-Wild TTS through EMA and Dual-Scoring Prompt Selection -- Submission for WildSpoof 2026 TTS Track F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-25T02:20:14.259206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T02:19:26.775357Z digest=sha256:d60518274d74adab31ac53876e851b3cd2e656c5fee9f95837b00b9df60c927d

Observation c5733c39-a634-47a8-8395-b5968f033f06 · inbound

FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations cites this paper.

FC-TTS: Style and Timbre Control in Zero-Shot Text-to-Speech with Disentangled Speech Representations F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T12:24:39.709349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T12:20:43.383513Z digest=sha256:f1bb9560e80926eb75136e2aea54ecc19e1277c09893fe2700fd69f44f912429

Observation 04100367-437a-4eaf-a515-afbc032ae967 · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-01T20:26:13.156506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:c8455f4e2e580d78d5d60f12ae6bccd922c6a5f733642b086ae3d1b758078905

Observation fca80790-ff8d-46e6-bd37-84b3b2c5517b · inbound

Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection cites this paper.

Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-01T20:36:12.148206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-28T18:17:44.307688Z digest=sha256:f5e51e2ad969e8c7e39c7de577fc6786348b2c0dc8d129ffe9532944325ca850

Observation ca2cf036-900e-4586-a4fc-d868ea73ac25 · inbound

Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection cites this paper.

Local Diagnostics of Continuous Normalizing Flow for Out-of-Distribution Detection F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-02T22:47:25.335375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-02T22:43:42.208048Z digest=sha256:8f852e858d9eddcb66c9660249996c1656ff42750730176b927df27ccee83340

Observation baa741de-c517-4131-a48f-adffa0ecc70d · inbound

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech cites this paper.

GLASS: GRPO-Trained LoRA for Acoustic Style Steering in Zero-Shot Text-to-Speech F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T15:27:04.899928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-27T23:53:55.385445Z digest=sha256:cc71158364b8dc48d526f6c568fab8bf9e4fb76729e275e792a00353f02a3fa7

Observation 594969c4-8b0c-4670-b72f-172c0f77fe7b · inbound

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation cites this paper.

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:37:35.158085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T15:15:10.060770Z digest=sha256:5486205cbea6cb8031e4830914c9f6ef72a866b5f184c94e34c9267e71e5b8cd

Observation 0210a7e3-3fac-4255-bc50-7475f8cbaa51 · inbound

End-to-End Training for Discrete Token LLM based TTS System cites this paper.

End-to-End Training for Discrete Token LLM based TTS System F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-03T03:27:34.864723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T15:22:06.893507Z digest=sha256:e59aab5265f31e9b85c193dd8e830a445a81c2887ee0fff66d85cd55efa3450f

Observation bee5dda8-50fc-4493-a0e4-8f8a1d89b6bc · inbound

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors cites this paper.

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-04T02:49:24.873272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T19:04:25.542629Z digest=sha256:d514a297c2fc30c8545bc2fbec102d24f8f54a5078b37d8c7243979714812a7e

Observation 645a0c5e-e4ef-406a-aefd-704cbcf3ac60 · inbound

The Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection cites this paper.

The Watermark Shortcut: How Provenance Marking Sabotages Audio Deepfake Detection F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T12:29:51.994305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-26T06:45:08.471913Z digest=sha256:0bfbcd2312a1cadbadf65a591f563df0aa9c610768be9bf83b758b12952800e8

Observation e71f793e-d49b-4a8e-84e3-8902340d7684 · inbound

A Survey of Automated Presentation Coaching: Systems, Methods, and Open Challenges cites this paper.

A Survey of Automated Presentation Coaching: Systems, Methods, and Open Challenges F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:15:47.511667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-30T22:06:35.036905Z digest=sha256:ddae2a11135f0c5f1032e68b1f21be8129da4ee55ce062dfbdee747260506001

Observation f4333c2d-f6ed-498b-b44d-d4806be4ead3 · inbound

MeloDISinger: Melody-Aware & Duration-Preserving Singing Voice Editing with Audio Infilling cites this paper.

MeloDISinger: Melody-Aware & Duration-Preserving Singing Voice Editing with Audio Infilling F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-06-30T03:24:12.479645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T03:20:05.125635Z digest=sha256:29e4f095ebd2b627fd46d37f3ad3bc4717c2a96ecdfb4eb6723f1eff63f4947d

Observation 86df8d3d-97bb-4123-871d-8241dbcddc23 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 206

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T11:45:47.128085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:bd11f94c74052495c676deecdfd524d3bc7ff43c808d22d6cd46817ed5e5fa44

Observation d806db55-8911-4940-b939-33017a980de0 · inbound

SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings cites this paper.

SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-04T01:29:22.073256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-04T01:21:02.597103Z digest=sha256:014e576856f4d6664e80fce44b96afcd8ef15ff7f48637ceaebbbde588a7a0d4

Observation 894c7ba6-3a9b-4448-adf4-85d4262a94f7 · inbound

GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech cites this paper.

GRAFT: Grafted Reference Audio for Fine-grained Pronunciation in Zero-shot Text-to-Speech F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T08:16:37.711326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T08:16:37.711326Z digest=sha256:104a0112068f1d9de660b228472ccfaeeeb3e1de96f222a8293f7a2c7769ab1a

Observation 5d34f0f4-396a-4ccd-aeda-d7cc5a1d7f0a · inbound

DETECT-3B-Omni is Agnostic of Content and Demographics cites this paper.

DETECT-3B-Omni is Agnostic of Content and Demographics F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T02:38:17.163398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T02:38:17.163398Z digest=sha256:940b3cad6720af3a5d900cde3b43d436f2e2696be5decb80122798ff8aa65dd3

Observation 5d387887-6bbb-4190-8ccb-1624c7f6a54b · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 226

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.329310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:07e585dba675903e2fc8d4c97a5677e66ece1c24eed4d33dbc6f60f1c3645fdb

Observation 5d07cda7-e665-4a31-bcdb-218d6e89ce22 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 226

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:9041bbc5097e83098552ccaea2fdfa8a7aa107770af6918508abc6b09968b211

Observation b412697a-b251-4b4a-b94f-6f3556a123e7 · inbound

BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech cites this paper.

BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-08T18:15:21.418131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-07-08T18:09:22.207379Z digest=sha256:785a22f4b49348cfdc03319d57a5d000179d1c50120fb8fec5e78f7e63dec95d

Observation c71bf561-fdf8-43ae-9893-d5f3ed4cb1e2 · inbound

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models cites this paper.

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 73

Resolution
unresolved
no resolver link, observed 2026-07-13T05:10:26.667731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:10:26.667731Z digest=sha256:f29923e623fc2949301c3aeb6b40dc34d328f150eb24616fbd062d13ada973fd

Observation 5552ee31-eff9-4fb4-b7d4-e900b9788094 · inbound

Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer cites this paper.

Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T18:46:13.279565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:46:13.279565Z digest=sha256:0a6824beea6e5e738eef64b627e68a911b37166350dc3d8a0f9db13d531001bf

Observation 559dbbb9-104c-4d86-b497-679ab7fe43a6 · inbound

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis cites this paper.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:12.225012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:12.225012Z digest=sha256:b62f142d15e4445aff3e78db325f47c7b0728981e4a84f943aa981ebb28ecac3

Observation a17ef54d-5c2e-4f32-b7b6-52593258209a · inbound

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm cites this paper.

Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T23:35:21.520078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:35:21.520078Z digest=sha256:1ca40c9daccb304adfbbf2e59de105ab201f28ecce1eb3fe42a51d808d204e18

Observation 599d5572-5398-415a-9896-66d7c5a80a71 · inbound

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis cites this paper.

DLLM-TTS: Block Discrete Diffusion Language Model for Text-to-Speech Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T02:47:35.599102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T02:47:35.599102Z digest=sha256:2d74b2c5127269246a8e6457797f811e36d676a4160598151b1e5769c3f632f9

Observation 22b08551-8d3f-4c2b-b52d-d8db65fe5228 · inbound

Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry cites this paper.

Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T20:44:55.874935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:44:55.874935Z digest=sha256:1d230ac69d5e8e519a64e59a43e46314fe8d16d880762b398944341cc9088531

Observation 3ebdeb05-2734-4491-b463-0b8deb33a428 · inbound

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks cites this paper.

SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T16:29:18.581793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:29:18.581793Z digest=sha256:3743370fe34f278493afb7a42a91ac19bbf7797b67827975c75f4c140573e00e