Pith. sign in

Paper Citation Record · LEDGER

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification

As of 20 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2501.00398.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00398 v2

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:57:23.680640Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved12
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f3ff5e62-e4b9-4575-90fd-b83780ebab0c · outbound

This paper cites ”A study of instrument-wise onset detection in Beijing opera percussion ensembles.” 2014 ieee international conference on acoustics, speech and signal processing (icassp).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”A study of instrument-wise onset detection in Beijing opera percussion ensembles.” 2014 ieee international conference on acoustics, speech and signal processing (icassp)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:24.113914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.558174Z digest=sha256:45664993a1a3922cf88148ab4eef4314f63acbd23bdcceb328aafa92a80812db

Observation 59f6d812-2d95-4724-ae1c-b37be4f21720 · outbound

This paper cites an unresolved cited work.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:57:24.103039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.570304Z digest=sha256:305537ffb6166c22de982a9521acb142a4842701527018c15a3baff47e9685db

Observation 14f174d0-0744-49b9-b135-a30925ddfed9 · outbound

This paper cites 2013 IEEE International Con- ference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification 2013 IEEE International Con- ference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:24.092131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.573873Z digest=sha256:116e725920046677aab878e25e194356f2b6a848560d438b8a96ba1573f8b54e

Observation a96608b0-10bc-4007-8bd2-3e01a3d2533f · outbound

This paper cites ”Musical genre classification of audio signals.” IEEE Transactions on speech and audio processing 10.5 (2002): 293-302.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Musical genre classification of audio signals.” IEEE Transactions on speech and audio processing 10.5 (2002): 293-302

Reference 6

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:57:24.071044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.584436Z digest=sha256:7406d4fa80630ccc00820c0164c9a4e00e3088835e32b5118a08a9aad5eb1180

Observation c2462b01-279e-4de9-81ed-5d95ec954053 · outbound

This paper cites ”Neural audio synthesis of musical notes with wavenet autoencoders.” International Conference on Machine Learning.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Neural audio synthesis of musical notes with wavenet autoencoders.” International Conference on Machine Learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:24.081712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.581063Z digest=sha256:8b4dc161c3442c05590f1ffb89f709c67b2641fbfd0ca9d0c858d27823cabc4a

Observation 8d60bf92-6d2b-452f-8029-b58b6fddaba6 · outbound

This paper cites an unresolved cited work.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:57:24.049036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.592778Z digest=sha256:031c98a3da2a42f7367004f2f80f999155f6667399b070e8d500d5e6fdd7c265

Observation 8f179eb2-69c5-4504-8b95-cf699af64461 · outbound

This paper cites ”CochlScene: Acquisition of acoustic scene data using crowdsourcing.” 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”CochlScene: Acquisition of acoustic scene data using crowdsourcing.” 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:24.059954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.588616Z digest=sha256:40e25ea3be5fcfd187d2dc5b08a8e5911729f7ac41c22be4be0402379f9b8548

Observation de48d8d0-0e8e-4ccd-99c4-1409c0a4ad5d · outbound

This paper cites ”A dataset and taxonomy for urban sound research.” Proceedings of the 22nd ACM international conference on Multimedia.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”A dataset and taxonomy for urban sound research.” Proceedings of the 22nd ACM international conference on Multimedia

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:24.025204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.600842Z digest=sha256:1e367a8b32ed15a13f836e05c781673118823a99b583f1f951781d126085a819

Observation cdc34476-76ff-463f-b945-672c9ca00a26 · outbound

This paper cites an unresolved cited work.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:57:24.036958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.596860Z digest=sha256:6d1549a444a8c63ff6b27739b8fc0b482062ecd3a0d99d219fd662e44c3f92ac

Observation 4fe66c3b-1ff4-49f9-bd88-d220b363e587 · outbound

This paper cites ”V ocalsound: A dataset for improving human vocal sounds recognition.” ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”V ocalsound: A dataset for improving human vocal sounds recognition.” ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:24.002755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.608456Z digest=sha256:c39dfbb24b3d2bc2ff60df572ee35831eeec76d98cf75f659027112646bf67b1

Observation 7028d418-1b2a-4788-a39e-eb680ab0e745 · outbound

This paper cites an unresolved cited work.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:57:24.014200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.604749Z digest=sha256:e0f8445f51372793821dbe904c20853bf0352be6898f13800805f88d38d7d41d

Observation 80d9d2d7-ec0e-475e-95fc-0309d3c49a78 · outbound

This paper cites A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.615953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.615953Z digest=sha256:fd30573592254b63351614e68c2423b2e8de1ee9f4c28cfddfa189912d60a746

Observation f62f1d8d-0862-4ace-b748-8f52fc79df2b · outbound

This paper cites GPT-4 Technical Report.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification GPT-4 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.612246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.612246Z digest=sha256:7b52a449bb1c6114fff1e8b2e4d85526d4113f0e2dd129dfef2cbc50126c52c3

Observation 225d9c59-5dd9-45cd-b148-00b952c19f18 · outbound

This paper cites ”Learning transferable visual models from natural language supervision.” International conference on machine learning.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Learning transferable visual models from natural language supervision.” International conference on machine learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.980319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.623829Z digest=sha256:8eff5304210a41ce3d2369115db530cf79e2258c0cda2d2537bd99ef32866b6c

Observation 7717f753-d9ef-4d6c-9588-99460779f6c2 · outbound

This paper cites ”Learning to prompt for vision-language models.” International Journal of Computer Vision 130.9 (2022): 2337-2348.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Learning to prompt for vision-language models.” International Journal of Computer Vision 130.9 (2022): 2337-2348

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.991427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.620230Z digest=sha256:098f7c0f34050f1db12017d0808a17146edb5f85fa9df687a98d55a0a638a31c

Observation 6c73568e-5fa0-45fb-9416-4ab96c6eb591 · outbound

This paper cites ”Audioclip: Extending clip to image, text and au- dio.” ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Audioclip: Extending clip to image, text and au- dio.” ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.957421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.631243Z digest=sha256:8cff401897a99cd8941b4019867834d1209aa45250084dd20743a8c32559a22c

Observation 6bf72149-9438-413e-9fe5-a7cad5bc9d97 · outbound

This paper cites ”Wav2clip: Learning robust audio representa- tions from clip.” ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Wav2clip: Learning robust audio representa- tions from clip.” ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.968920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.627338Z digest=sha256:c8aa62187bd7bd6b5df407a6f4f7479942bd7427b30e5e834706602343f0f65f

Observation fc8281e9-2d77-4a14-82f7-a70ae672c784 · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.638165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.638165Z digest=sha256:c08723ba099a6808c33878a2a619cc8b1b7956065bd9bf04132ab0cc8b53d5f1

Observation 854e330a-e655-4613-b919-27553f9e5fe0 · outbound

This paper cites ”Clap learning audio concepts from natural language supervision.” ICASSP 2023-2023 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Clap learning audio concepts from natural language supervision.” ICASSP 2023-2023 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.945615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.634776Z digest=sha256:6b6b83a976a05da15fb3fa743a8eec1f05ecc2bbc93b366d0300568b2d7a6fe2

Observation 5bdcc50c-219e-4801-b8fe-d411431c881b · outbound

This paper cites ”Maple: Multi-modal prompt learn- ing.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Maple: Multi-modal prompt learn- ing.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.931765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.646340Z digest=sha256:1d9ceb35b20ef03818cd51214dd78e3172344a6cd2c17c2b67096a9914852fd6

Observation ab12a8f3-5809-46b3-a42c-439e1f7797b8 · outbound

This paper cites CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.642312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.642312Z digest=sha256:d3acb1f06cbda1128ea3bef406fd296193b6e9753825baca84e6b00d597f1cd8

Observation 74e3b013-9c6d-4c11-b848-c853725a3259 · outbound

This paper cites ”Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery.” Advances in Neural Information Processing Systems 36 (2024).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery.” Advances in Neural Information Processing Systems 36 (2024)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.905785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.653503Z digest=sha256:39ffe7f530247aa039772a49d15a99ec19bdbd62745d83f076a407d4e1ff2947

Observation 98dbaab5-52fb-4ae4-a9f7-b14c5fc28df3 · outbound

This paper cites ”A survey of audio classification using deep learning.” IEEE Access (2023).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”A survey of audio classification using deep learning.” IEEE Access (2023)

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.919081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.649755Z digest=sha256:27adbb5ec63eb442aa1deadabfa57252b5a5ddb5bcf86e8e5e308a98a524da44

Observation 606ae98f-7685-4978-8fdc-e0c19507394f · outbound

This paper cites Adapting Language-Audio Models as Few-Shot Audio Learners.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification Adapting Language-Audio Models as Few-Shot Audio Learners

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.661121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.661121Z digest=sha256:7527151dd3efdbc1c5fc5a746f4f453237ec439608710c1f3dc1b33c84d0bb85

Observation 3c9f63b6-f367-4d26-971c-175c8f69ad61 · outbound

This paper cites ”Audio-Free Prompt Tuning for Language-Audio Models.” ICASSP 2024-2024 IEEE In- ternational Conference on Acoustics, Speech and Signal Processing (ICASSP).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Audio-Free Prompt Tuning for Language-Audio Models.” ICASSP 2024-2024 IEEE In- ternational Conference on Acoustics, Speech and Signal Processing (ICASSP)

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.892140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.657173Z digest=sha256:b47bf5d32fae23ba78e69f9c244992c4e55e69444d036201ddc3a3491d5be3fd

Observation 4f502568-90fe-46f1-ab8c-ac0e1c6c6272 · outbound

This paper cites A Survey of Multimodal Large Language Model from A Data-centric Perspective.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification A Survey of Multimodal Large Language Model from A Data-centric Perspective

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.668961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.668961Z digest=sha256:8f48f587f30486eac1686c3628f0a56d8d8f1c5aa28ee0f1debff4565d8025f9

Observation c5175c41-0b58-4570-a3a1-fc17d45c4611 · outbound

This paper cites ”Multimodal large language models: A survey.” 2023 IEEE International Conference on Big Data (BigData).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Multimodal large language models: A survey.” 2023 IEEE International Conference on Big Data (BigData)

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.879355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.665175Z digest=sha256:8d33b0fe1d7ce0c942cedaba18d7b2f908a850af8ae956d8628d5fd6553245a9

Observation 806f129c-24c2-4a50-bb8f-8c3a2313824c · outbound

This paper cites Grounding Multimodal Large Language Models in Actions.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification Grounding Multimodal Large Language Models in Actions

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:57:23.729374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T22:57:23.676860Z digest=sha256:16c5bec61143efc3fad5b6cd56b579e32d3e3c2881be069ac62b3caae393dd68

Observation 3fe66244-c858-49a3-90f3-a4f1b055d70c · outbound

This paper cites ”Efficient multimodal large language models: A survey.” arXiv preprint arXiv:2405.10739 (2024).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Efficient multimodal large language models: A survey.” arXiv preprint arXiv:2405.10739 (2024)

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.673281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.673281Z digest=sha256:73da3f814c02f6d5dcf611e0950d2b16de19938be0582aa2aa4d6970df196337

Observation 8e3ec007-f0cb-4712-9512-b4af9f4661a5 · outbound

This paper cites A Review of Multi-Modal Large Language and Vision Models.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification A Review of Multi-Modal Large Language and Vision Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.680640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.680640Z digest=sha256:7367a2130e42c8ad7225f1ca1983962e24ac1986b574e6f26236218d77fb494e

Pith citing papers

No inbound Pith citation observations are available.