Pith. sign in

Paper Citation Record · LEDGER

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style

As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2508.11187.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.11187 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:08:58.582356Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy39
  • unresolved5
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2987e7b9-c65c-420f-b679-ceba3ebe440a · outbound

This paper cites Retrieval and browsing of spoken content,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Retrieval and browsing of spoken content,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:08.591689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:53.581070Z digest=sha256:6a3cf87a91e3b8120378af7e1f9989fb34f1b24ee531f6656a287b9dd3b115af

Observation bf2f4c16-bcd1-4ed3-8686-cffeca2a4677 · outbound

This paper cites V oice-based information retrieval—how far are we from the text-based information retrieval?.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style V oice-based information retrieval—how far are we from the text-based information retrieval?

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:08.410347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:53.693568Z digest=sha256:5272211b1d1340d973024d7c1b3aa67d6a92c31f55cf71c50c656fd4d219ca36

Observation b7769f79-5289-4bb2-a1d8-24ac5c0951d7 · outbound

This paper cites Spoken content retrieval: A survey of techniques and technologies,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Spoken content retrieval: A survey of techniques and technologies,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:08.272085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:53.819901Z digest=sha256:0fdfc417da64e0d9b9e31612997134adc1ec7f55c0e842b85c3e9af3b421902c

Observation 0460dde0-cf7d-4deb-b65a-92845e325e4e · outbound

This paper cites Spoken content retrieval—beyond cascading speech recognition with text retrieval,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Spoken content retrieval—beyond cascading speech recognition with text retrieval,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:08.061091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:53.970416Z digest=sha256:343055548358066f3c4058128bbf34eefd0040c5efa791ba6e2b755bff32e84f

Observation 3a8108d8-5251-4262-924a-7044d05beca3 · outbound

This paper cites SpeechDPR: End-to-end spoken passage retrieval for open-domain spoken question answering,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style SpeechDPR: End-to-end spoken passage retrieval for open-domain spoken question answering,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:07.834944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:54.089386Z digest=sha256:a2de1d9d1b222c16e884af762d71f4f7bee2a4ae8ae4dacefc4d07f9077025c9

Observation 3447d916-4f3c-4c70-8c79-e20aedddd086 · outbound

This paper cites Retrieval augmented end-to-end spoken dialog models,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Retrieval augmented end-to-end spoken dialog models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:07.599886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:54.211609Z digest=sha256:ee3362afaae40a281dc3e808135b418cd03e46314a0ab7fccfbef35f68c0239a

Observation 88292e66-930c-407c-b267-94e6aee151af · outbound

This paper cites WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style WavRAG: Audio-Integrated Retrieval Augmented Generation for Spoken Dialogue Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T20:08:54.327856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:08:54.327856Z digest=sha256:9f191455f36de11b4724d9bd0338a03fcad56410f792b4eccbe3a628682a52ee

Observation bbbe176b-4fc8-4de3-acf0-0507ebdf67f6 · outbound

This paper cites Speech retrieval-augmented generation without automatic speech recognition,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Speech retrieval-augmented generation without automatic speech recognition,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:07.394761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:54.447022Z digest=sha256:efbc4890e7ce8d592f3327d1ce3304c71872bc74b144f3a7b528f88008a90f3a

Observation 048add9f-8558-4874-90a5-8277fe7b6ca7 · outbound

This paper cites Prompting audios using acoustic properties for emotion representation,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Prompting audios using acoustic properties for emotion representation,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:07.222536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:54.615420Z digest=sha256:530df4b9f7b030ddac44fbd20638ed437e83eb924e8fdc3f7d36666fbf0b3b40

Observation fd5d1968-4cc0-4b3e-b386-f081012e6276 · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:07.029219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:54.731774Z digest=sha256:bae9c9d2cb4464d27d0e81b214e2f58afe63ceabd48afc18c691a69c8b6743f4

Observation 9409ad26-5b5a-419b-a9d6-2d18f66845ab · outbound

This paper cites IEMOCAP: Interactive emotional dyadic motion capture database,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style IEMOCAP: Interactive emotional dyadic motion capture database,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:06.804722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:54.894348Z digest=sha256:0ca245e95dcd824d65006d5ffc0504a5ad1e58905cf9c58790a7aaf03a888504

Observation 2b00b25b-503a-462f-b924-f2987c8e4824 · outbound

This paper cites Emotional voice conversion: Theory, databases and ESD,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Emotional voice conversion: Theory, databases and ESD,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:06.618778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:55.011380Z digest=sha256:ddf1ccd2b4f7865c5d7a40355c23d2b8313c7e991c8c5b8dd2c40a4162c256c3

Observation a631dcd0-1d34-4f75-9465-1b2af2119817 · outbound

This paper cites Expresso: A benchmark and analysis of discrete expressive speech resynthesis,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Expresso: A benchmark and analysis of discrete expressive speech resynthesis,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:06.446662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:55.123871Z digest=sha256:e9d49634e56f928fdcb91de179516706ed512540944b0d15dbb669c715288495

Observation 383da48d-113c-467a-9679-9cf71a7c0c63 · outbound

This paper cites PromptTTS: Controllable text-to-speech with text descriptions,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style PromptTTS: Controllable text-to-speech with text descriptions,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:06.256855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:55.239997Z digest=sha256:9c428b858db8a5737d716494e52b6a32cfbfd3fa2f1fa8efbd44f43aecbeb256

Observation 4f6c3a39-8f78-4e9d-8f30-8295154507e9 · outbound

This paper cites PromptStyle: Controllable style transfer for text-to-speech with natural language descriptions,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style PromptStyle: Controllable style transfer for text-to-speech with natural language descriptions,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:06.074349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:55.336467Z digest=sha256:bb2bbfdb8f5187b6e2e1be4fbdb3cc83f2a0d56ea2a49a77b05d6fd327fa0cfb

Observation cefdac49-bc5f-4d28-b2f0-5f0bbdc2d1c4 · outbound

This paper cites PromptTTS 2: Describing and generating voices with text prompt,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style PromptTTS 2: Describing and generating voices with text prompt,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:05.843066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:55.462957Z digest=sha256:f5354079334901154e4f5e84813f30e58f74712b87c4e0cb4fec8f1583ae30a5

Observation 4ca6dce2-e074-4916-aa4f-9695a3d41165 · outbound

This paper cites DreamV oice: Text-guided voice conversion,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style DreamV oice: Text-guided voice conversion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:05.624428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:55.587642Z digest=sha256:f34d86c2a50bfcc644968b1773f41f5a46806fc1ee29c7960f6d59777e7c836d

Observation a5f17ddf-3b7c-43e4-be54-ad62714dc880 · outbound

This paper cites StyleCap: Automatic speaking-style captioning from speech based on speech and language self-supervised learning models,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style StyleCap: Automatic speaking-style captioning from speech based on speech and language self-supervised learning models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:05.438828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:55.706895Z digest=sha256:403c968d46d8fc7b9de2acaa5c2e35a6f5f1ea3ade5caadd0986198df00c6661

Observation d516af7a-56b6-4212-a7c5-7033bc9a7978 · outbound

This paper cites LibriTTS-P: A corpus with speaking style and speaker identity prompts for text-to-speech and style captioning,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style LibriTTS-P: A corpus with speaking style and speaker identity prompts for text-to-speech and style captioning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:05.193490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:55.823325Z digest=sha256:e2f92196ba1f77a94edd0fb41008cb1a856620dde5af86ad6e18b4579ced729c

Observation 96f0f9f0-92d8-460f-97fd-6d22f5e94e2b · outbound

This paper cites V ocabulary independent spoken term detection,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style V ocabulary independent spoken term detection,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:04.909259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:56.015717Z digest=sha256:ae7a12c447462acda9f211d2156912eee4fbb6a190b77affe69c2c65892f3d94

Observation 7c8bf58c-9105-4f0e-84ba-6ff34ea07034 · outbound

This paper cites Statistical lattice-based spoken document retrieval,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Statistical lattice-based spoken document retrieval,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:04.619531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:56.143258Z digest=sha256:b26d083fead5500893b63289b267532b84e9f74b2cf91540d1eb58d703de43fe

Observation e7af4a0a-085a-48da-b017-70bc70675d86 · outbound

This paper cites Improved semantic retrieval of spoken content by document/query expansion with random walk over acoustic similarity graphs,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Improved semantic retrieval of spoken content by document/query expansion with random walk over acoustic similarity graphs,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:04.334449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:56.260715Z digest=sha256:af2a1013459f239c285bc68dcd02966e88175ac1e7efdb3a6bf9d5e8a7b7742b

Observation 35a12892-4b63-4ab0-a6fd-87b1bbd05521 · outbound

This paper cites Evaluating ASR output for information retrieval,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Evaluating ASR output for information retrieval,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:04.030441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:56.414580Z digest=sha256:5dc68487e9deedc05961e66b7e98d2ef63cb7a05da5141ffce3ac16044c558f2

Observation 89d11070-717c-47cf-8bf1-3df3ad898e04 · outbound

This paper cites Investigating the global semantic impact of speech recognition error on spoken content collections,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Investigating the global semantic impact of speech recognition error on spoken content collections,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:03.622881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:56.573972Z digest=sha256:e5c9b1960d1301b13ee4d56338b2e54d2f1b8a2a18844af3241329a11fb82644

Observation 166dafa7-2cb9-4188-9026-d506fbb459bc · outbound

This paper cites Speech-centric information processing: An optimization-oriented approach,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Speech-centric information processing: An optimization-oriented approach,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:03.133742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:56.654325Z digest=sha256:7fe2abab6304f60db16c9c3b055fe270268fc59bca561b15ff4dd484a412329b

Observation e1876a5c-0e8c-4fef-b17a-322ba56511c9 · outbound

This paper cites Unsupervised spoken-term detection with spoken queries using segment-based dynamic time warping.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Unsupervised spoken-term detection with spoken queries using segment-based dynamic time warping

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:02.778218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:56.754424Z digest=sha256:962c6356e6ef552e36ebf560dc7f3bda2b00dc90677725c088aa74964c3ef327

Observation a1b05aaa-b43b-41c9-b51d-705d123fa730 · outbound

This paper cites Memory efficient subsequence dtw for query-by-example spoken term detection,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Memory efficient subsequence dtw for query-by-example spoken term detection,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:02.467412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:56.873814Z digest=sha256:ff189681c98fede325e0b5dbc681a64c938cbb29c75419d5dc7082fc52c4ecb6

Observation 3cddcd3e-154d-49d2-bdde-aaa3697cb39e · outbound

This paper cites The spoken web search task at MediaEval 2012,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style The spoken web search task at MediaEval 2012,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:02.140897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:57.012996Z digest=sha256:09a4fa7e458829663aa147abad21a38ae473b4896071a217840b7687268be3cb

Observation 29f2ff46-011e-4f9a-8824-b0d76f1b53e0 · outbound

This paper cites RECAP: Retrieval-augmented audio captioning,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style RECAP: Retrieval-augmented audio captioning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:01.880751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:57.134441Z digest=sha256:d6d758d74ae25cdeef9221540beae51b4161f7639826ab7f1aa64a89da51fae2

Observation 710f025a-c68c-43a0-a5c5-7e05f52846b8 · outbound

This paper cites Beyond speaker identity: Text guided target speech extraction,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Beyond speaker identity: Text guided target speech extraction,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:01.586330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:57.182972Z digest=sha256:5e7f2c55fcb120af020a76406ae5fcfdb60a6d9adbc50041be60f5f69a67dbdd

Observation 4ce8b136-65d5-40bd-b6db-cb39af5900cd · outbound

This paper cites WavLM: Large-scale self-supervised pre- training for full stack speech processing,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style WavLM: Large-scale self-supervised pre- training for full stack speech processing,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T20:08:57.318770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:08:57.318770Z digest=sha256:314ca4613155e98550a94b0ff5def03db895ba88251ab99a5125ab45040fd9cc

Observation 02b9ac51-317e-4829-b155-f89905685ddb · outbound

This paper cites emo- tion2vec: Self-supervised pre-training for speech emotion representation,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style emo- tion2vec: Self-supervised pre-training for speech emotion representation,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:01.170207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:57.429737Z digest=sha256:5f909738ec13a7dfe6d3dc7907a59a42bef1ba5388dc0f795b9008e397f98a31

Observation 448eae5e-275a-4316-9d96-bf266f4d046b · outbound

This paper cites BERT: Pre- training of deep bidirectional transformers for language understanding,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style BERT: Pre- training of deep bidirectional transformers for language understanding,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:00.963076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:57.528474Z digest=sha256:9fbbc1e2020f74d904a0147b73f3a9f5b423dd9ddd5862ee79f64899b9e14e93

Observation 79accd8c-d82c-4429-9aa6-6ce3e18e6b2b · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T20:08:57.647214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:08:57.647214Z digest=sha256:393356a9f6d7a69d08071fabb8f7356973c1ca85a5d566e70f8f5e86751239f7

Observation d53a25fb-05a3-405c-8373-612ac569b576 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:00.773425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:57.795428Z digest=sha256:769bc2a6f51a48064a104f2107aa7a5b175bdec53e534d8285b4233f7e5aa3b0

Observation f681b8d2-59e9-4d9b-9092-8f3d21cf3520 · outbound

This paper cites The flan collection: Designing data and methods for effective instruction tuning,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style The flan collection: Designing data and methods for effective instruction tuning,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:00.530926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:57.923525Z digest=sha256:ecc809189e4246d40988a4a5e26ac12b670593c43ef687ec8e8474594253b949

Observation 9b112b2f-79a3-4cd5-a356-4d8360b9f105 · outbound

This paper cites Sentence-T5: Scalable sentence encoders from pre-trained text-to-text models,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Sentence-T5: Scalable sentence encoders from pre-trained text-to-text models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:09:00.271532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:58.045198Z digest=sha256:885b8474f32cb6bbf85af4d307c0d762a8b32503671fe010b97526f829ee1b27

Observation 83195699-1641-47d1-af58-b3e32cbbb861 · outbound

This paper cites Domain-adversarial training of neural networks,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Domain-adversarial training of neural networks,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:08:59.969229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:58.052707Z digest=sha256:48c513b35256e90aa2cacff5a0b0a6c9da836cf3af7d0ca82e8d56b8666d9790

Observation 24279dca-a0f0-444d-b756-51de2a8d3365 · outbound

This paper cites Unsupervised domain adaptation by backpropagation,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Unsupervised domain adaptation by backpropagation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:08:59.689390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:58.056120Z digest=sha256:6f78db39dda480890c35a33e523169a219f5f80ca8e3a1766a02b239c255f170

Observation b802835a-261b-4e40-bf7e-a3b5a91fb24e · outbound

This paper cites GPT-4o System Card.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style GPT-4o System Card

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T20:08:58.063279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:08:58.063279Z digest=sha256:4668fceda94313b9d9f043f34667e30d7fe30c53f56f9f27793a80dd69725a6a

Observation 96bdecf8-9edd-4b16-a788-e8cf996cd90e · outbound

This paper cites Decoupled weight decay regularization,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Decoupled weight decay regularization,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T20:08:58.165179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:08:58.165179Z digest=sha256:af5ab47659b2b905320f34d8c41dfc410114910626e61c892f1686a401d80a8e

Observation f5870a2b-bf04-496f-b0cd-2d8a5fdf736e · outbound

This paper cites PromptTTS++: Controlling speaker identity in prompt-based text-to-speech using natural language descriptions,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style PromptTTS++: Controlling speaker identity in prompt-based text-to-speech using natural language descriptions,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:08:59.461612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:58.281102Z digest=sha256:ddd3b2acc6d3472606d8ef218576d7096b3089c5b6de50c560a04ea78c74e841

Observation c3b5de4c-bd7f-4b68-b677-f8ce08f3a5e3 · outbound

This paper cites PromotiCon: Prompt- based emotion controllable text-to-speech via prompt generation and matching,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style PromotiCon: Prompt- based emotion controllable text-to-speech via prompt generation and matching,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:08:59.165146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:58.359548Z digest=sha256:e6a6f99583a25175c965d94cf5ecda6361e9371f6d01b77a5d0436a6f2c91131

Observation f6bc0f0e-80fd-42db-9cf4-d6fef5e76115 · outbound

This paper cites Visualizing data using t-SNE,.

Expressive Speech Retrieval using Natural Language Descriptions of Speaking Style Visualizing data using t-SNE,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T20:08:58.841982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T20:08:58.582356Z digest=sha256:2b3520d638cded9f8f33e148cff73598017e88135421ee4d78442fcfd7a0a809

Pith citing papers

No inbound Pith citation observations are available.