Pith. sign in

Paper Citation Record · LEDGER

Text-based Audio Retrieval by Learning from Similarities between Audio Captions

As of 18 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2412.01356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.01356 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:29:30.849620Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact3
  • verified fuzzy23
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d54ccc7c-78ef-43de-891d-e2760bac62ad · outbound

This paper cites Language-Based Audio Retrieval Task in DCASE 2022 Challenge,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Language-Based Audio Retrieval Task in DCASE 2022 Challenge,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.171487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.756641Z digest=sha256:c01af1dd47368b3af5e2a69ab3d94ba57a8a4e87eb3f3163ee4109a86d094e22

Observation f62bbf35-64f5-40d8-8267-bee3aae09633 · outbound

This paper cites Improving Natural-Language-Based Audio Retrieval with Transfer Learning and Audio & Text Augmentations,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Improving Natural-Language-Based Audio Retrieval with Transfer Learning and Audio & Text Augmentations,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.161685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.760639Z digest=sha256:ffc77f9ba318735cff7698eba9c4b038600e77319ea9b9f8c71e0c2845650474

Observation d3d6a855-df62-41aa-a998-ef08512a331f · outbound

This paper cites Matching Text and Audio Embeddings: Exploring Transfer-Learning Strategies for Language-Based Audio Retrieval,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Matching Text and Audio Embeddings: Exploring Transfer-Learning Strategies for Language-Based Audio Retrieval,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.151265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.764107Z digest=sha256:17cbed5b4eba4cc944fa63281a12b1553fb7f5ad00d28d51a28411e17336dd34

Observation 932fcd97-c9a4-4002-baf7-e022f3b1eb62 · outbound

This paper cites Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Large-Scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.141738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.767800Z digest=sha256:3b613457d1888d07fe8621eedaa349bf76f66bd28e214ef637934a4ece31a191

Observation dd3dbe2f-db55-4696-9997-d2104e6f41f1 · outbound

This paper cites WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.132169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.772078Z digest=sha256:e0403f0d9e1294aa72a5572723414aa13aa499c9a6c896decd2c6134e5e7901b

Observation cffba290-5e42-4089-9ee6-0a04c523689f · outbound

This paper cites Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Auto-ACD: A Large-scale Dataset for Audio-Language Representation Learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.121899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.776273Z digest=sha256:4de3e2e39934332811cc34231a32f4c8e884f08f3b06d894f61bd2e2d52611be

Observation 8fb02410-cb4e-4d2d-954d-174335f52f97 · outbound

This paper cites Advancing Natural-Language Based Audio Retrieval with Passt and Large Audio-Caption Data Sets,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Advancing Natural-Language Based Audio Retrieval with Passt and Large Audio-Caption Data Sets,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.111192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.779789Z digest=sha256:b310c2b83213faa1b48a88c5ad3d5ae79611db3e9eae99ec43e70e1700bd05a2

Observation b2fab24d-c5ba-4752-a01a-a134a45c827c · outbound

This paper cites Language-based Audio Retrieval in DCASE 2023 Challenge,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Language-based Audio Retrieval in DCASE 2023 Challenge,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.101147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.783281Z digest=sha256:b5e8238f0955e5eb6d5057934a4be8c100b23b5f87c91cb78355481318769380

Observation 4f6547c7-2fa3-4bb2-bc93-b781b1605063 · outbound

This paper cites AudioCaps: Generating Cap- tions for Audios in The Wild,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions AudioCaps: Generating Cap- tions for Audios in The Wild,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.090549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.786604Z digest=sha256:75813aeeed2ebb23e986d3623dcfa85c68056ebe7742aa023107efd458c00bfd

Observation 2a84f74e-62bf-48c7-a417-791813cf2f8a · outbound

This paper cites Clotho: an Audio Captioning Dataset,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Clotho: an Audio Captioning Dataset,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.080742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.789720Z digest=sha256:94a191770414b719b75ba1461140cd6a070f3114043c5d07a657137fab456235

Observation 73cdbc5b-87b8-47e1-88fd-b13e44179738 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.070655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.793581Z digest=sha256:8eeed1c095d31c20297507739223b075da39cf53c07d3f81a90aa29bb34a5f39

Observation 293b2e9f-5f5c-4ff6-a492-e3cd7cd04e0a · outbound

This paper cites Graded Relevance,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Graded Relevance,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.060980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.797234Z digest=sha256:dd935df247760421eacd812441b5972a693088ccbe7896af557856f54a89be4f

Observation f75ef85d-f33b-4c95-ab54-88d4ceac6072 · outbound

This paper cites On the effect of relevance scales in crowdsourcing relevance assessments for Information Retrieval evaluation,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions On the effect of relevance scales in crowdsourcing relevance assessments for Information Retrieval evaluation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.050644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.800644Z digest=sha256:bae07ecd2e89b60c57a207e6701b1a2d687b7ff63eb2ed6bd0bc4c8ae31cdc26

Observation 5416fdc1-b405-4654-90bb-878c841f2856 · outbound

This paper cites Crowdsourcing and Evaluating Text-Based Audio Retrieval Relevances,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Crowdsourcing and Evaluating Text-Based Audio Retrieval Relevances,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.040285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.804013Z digest=sha256:ec64d2e1c7f8e63c97a1a78b8b110ea38dea29936a636b778ad0e7a270fa84e7

Observation b8bb1445-9b22-41b0-9c59-ca365e320b1c · outbound

This paper cites Integrating Continuous and Binary Relevances in Audio-Text Relevance Learning.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Integrating Continuous and Binary Relevances in Audio-Text Relevance Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:29:30.931577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.807338Z digest=sha256:c2cfb8a8fdd99a91f4f2a48074e52f37c333cc13adad9d395e3658378adf0d2e

Observation 3e037dc8-4f0e-47e6-a502-fa7717086dbd · outbound

This paper cites Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Estimated Audio-Caption Correspondences Improve Language-Based Audio Retrieval

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:29:30.917928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.810891Z digest=sha256:fa4e78c6cb7e41c28ef327545f96239cd18ebf44f9fc0410f14c02dd6ea89411

Observation dd0aa89d-c1d4-4730-9ad4-80e70c453f0b · outbound

This paper cites Learning to Rank: From Pairwise Approach to Listwise Approach,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Learning to Rank: From Pairwise Approach to Listwise Approach,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.030190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.814229Z digest=sha256:8f0c4d7357cf9052783ab07af35ce9b672aee3e799e26552830309a89f78fa23

Observation f718f14e-6ca7-4604-a238-c06c5766a496 · outbound

This paper cites Image-Text Retrieval with Binary and Continuous Label Supervision.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Image-Text Retrieval with Binary and Continuous Label Supervision

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-12T04:29:30.903867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.817138Z digest=sha256:fad6975ee67e82f4d6a530be9e9c38975970e988d049fab99ff0af6a135e99c1

Observation 6ac29d83-9077-4ee8-be3d-b052e91e8680 · outbound

This paper cites Integrating Listwise Ranking into Pairwise-based Image-Text Retrieval,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Integrating Listwise Ranking into Pairwise-based Image-Text Retrieval,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.020252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.820672Z digest=sha256:c82f5ef40fdc764a404407e411d46a246a2d405b708aeadb72b100391b52acde

Observation 9d1228eb-5555-42f0-bd00-00002d6f57de · outbound

This paper cites Freesound Technical Demo,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Freesound Technical Demo,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:31.010085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.823645Z digest=sha256:814544fcbd0c5ce3c73e8cfe018817183d54ebb1d054aa4a8f068709d1b5c287

Observation 63af5136-c1ae-4ed2-81f3-692e514dc221 · outbound

This paper cites BBC Sound Effects,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions BBC Sound Effects,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:30.998687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.826930Z digest=sha256:f29e0ae34b9d581b2e0d9ef6091e31e43f3d9c4350a8cf4a5b7e4445762fb37e

Observation 4cf472e2-8dfe-4c29-8ac7-a9f3dafcd981 · outbound

This paper cites SoundBible,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions SoundBible,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:30.987322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.830028Z digest=sha256:8d7e331fc53ba347ef62988b87f4e48264b009d4375337f1c4fca792b9215895

Observation 0c1be653-1ff1-4e4f-a461-af16978f09c4 · outbound

This paper cites The Benefit of Temporally-Strong Labels in Audio Event Classification,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions The Benefit of Temporally-Strong Labels in Audio Event Classification,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:30.976781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.833197Z digest=sha256:935cf8a03dced0b3888fe3104ba9383bc8b567677495d6dd9930429ee5ab0085

Observation 3921b18b-826e-4cc7-bdf9-bd3c10dfc271 · outbound

This paper cites Language-based Audio Retrieval in DCASE 2024 Challenge,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Language-based Audio Retrieval in DCASE 2024 Challenge,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:30.965574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.836301Z digest=sha256:7a67e0f3171b4353dd11d05e418798178fba7fa33c58c91143676b756168d6cf

Observation a1ff3e74-046f-4aea-a084-2ef2fc42d95a · outbound

This paper cites Efficient Training of Audio Transformers with Patchout,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Efficient Training of Audio Transformers with Patchout,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:30.954811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.839396Z digest=sha256:0ce5664e8e1297ed446d3cbcd4e045bd2b1bfdd075d5bf73c624a9862fca474e

Observation 98763c99-d3e5-469e-a47e-4cb60febaaf5 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T04:29:30.842632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:29:30.842632Z digest=sha256:6e311f1375fba1a262abc855ee69211a295ef1678cfa16b5cbe8f0f5fbf93f62

Observation 2e1ff558-e9ed-4098-aea2-61ba95c9401a · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions Representation Learning with Contrastive Predictive Coding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:29:30.845996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:29:30.845996Z digest=sha256:9eed67a9d505ee94f25cfe17ce9384ad5ba594072c026c26b471903fd09ca1ef

Observation c6a3a6c6-771b-40d9-8e12-631211197e55 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts,.

Text-based Audio Retrieval by Learning from Similarities between Audio Captions SGDR: Stochastic Gradient Descent with Warm Restarts,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T04:29:30.943386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-12T04:29:30.849620Z digest=sha256:59705170d6b6413a7671566bda2d0aa8874a68b48458d2203b69508b1deef8af

Pith citing papers

No inbound Pith citation observations are available.