Pith. sign in

Paper Citation Record · LEDGER

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification

As of 14 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2501.00398.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.00398 v2

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:57:23.680640Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved12
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f3ff5e62-e4b9-4575-90fd-b83780ebab0c · outbound

This paper cites ”A study of instrument-wise onset detection in Beijing opera percussion ensembles.” 2014 ieee international conference on acoustics, speech and signal processing (icassp).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”A study of instrument-wise onset detection in Beijing opera percussion ensembles.” 2014 ieee international conference on acoustics, speech and signal processing (icassp)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:24.113914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.558174Z digest=sha256:05b02341c5161fda680a0d1ddeb6e544b7491567a881b0ed2f21e606370c038b

Observation 59f6d812-2d95-4724-ae1c-b37be4f21720 · outbound

This paper cites an unresolved cited work.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:57:24.103039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.570304Z digest=sha256:770cefc8df1ae0a4bf86473a7bebc929dc9a2d4c83aebdea6a37a4326699509e

Observation 14f174d0-0744-49b9-b135-a30925ddfed9 · outbound

This paper cites 2013 IEEE International Con- ference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification 2013 IEEE International Con- ference on Acoustics, Speech and Signal Processing, Vancouver, BC, Canada

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:24.092131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.573873Z digest=sha256:295ea23db03109e74967ef571e13cef764fe0350683a34e0a962555f7d8f6a39

Observation a96608b0-10bc-4007-8bd2-3e01a3d2533f · outbound

This paper cites ”Musical genre classification of audio signals.” IEEE Transactions on speech and audio processing 10.5 (2002): 293-302.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Musical genre classification of audio signals.” IEEE Transactions on speech and audio processing 10.5 (2002): 293-302

Reference 6

Resolution
malformed identifier
raw_fallback, observed 2026-08-10T22:57:24.071044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.584436Z digest=sha256:9d4340d62046570c4e1a93f539cef881be974047ab7911a727e247e8c23cfc5e

Observation c2462b01-279e-4de9-81ed-5d95ec954053 · outbound

This paper cites ”Neural audio synthesis of musical notes with wavenet autoencoders.” International Conference on Machine Learning.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Neural audio synthesis of musical notes with wavenet autoencoders.” International Conference on Machine Learning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:24.081712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.581063Z digest=sha256:ca38f4c2ed420f32cd117fde50e88aeddcf308f7228bbe1b1dbf67511c1af6ef

Observation 8d60bf92-6d2b-452f-8029-b58b6fddaba6 · outbound

This paper cites an unresolved cited work.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:57:24.049036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.592778Z digest=sha256:f939504aa27f38deb105ccf8f4d5a6cad9bdc697ad54492280ec488d562f652d

Observation 8f179eb2-69c5-4504-8b95-cf699af64461 · outbound

This paper cites ”CochlScene: Acquisition of acoustic scene data using crowdsourcing.” 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”CochlScene: Acquisition of acoustic scene data using crowdsourcing.” 2022 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:24.059954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.588616Z digest=sha256:35fabf4d1069326751d274730f0b999c5de1f7b6510fb355865622bbdbc7e2c1

Observation de48d8d0-0e8e-4ccd-99c4-1409c0a4ad5d · outbound

This paper cites ”A dataset and taxonomy for urban sound research.” Proceedings of the 22nd ACM international conference on Multimedia.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”A dataset and taxonomy for urban sound research.” Proceedings of the 22nd ACM international conference on Multimedia

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:24.025204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.600842Z digest=sha256:119872800796f00dc4d705094b68338aeb0b89892e8ab795c00785617ca474e6

Observation cdc34476-76ff-463f-b945-672c9ca00a26 · outbound

This paper cites an unresolved cited work.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:57:24.036958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.596860Z digest=sha256:95bb195a443ea0d10852013a73d6a9332c3a1f89c709162805d082da66022a5a

Observation 4fe66c3b-1ff4-49f9-bd88-d220b363e587 · outbound

This paper cites ”V ocalsound: A dataset for improving human vocal sounds recognition.” ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”V ocalsound: A dataset for improving human vocal sounds recognition.” ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:24.002755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.608456Z digest=sha256:b5a7fe8eacd88729d75b39db18bae81e98e8c838c991f0e29af22627f8625218

Observation 7028d418-1b2a-4788-a39e-eb680ab0e745 · outbound

This paper cites an unresolved cited work.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T22:57:24.014200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.604749Z digest=sha256:65aee4c7c96c3224deddd3bb20e32d079d93a85bcb824baf310e8dae9cc4e45a

Observation 80d9d2d7-ec0e-475e-95fc-0309d3c49a78 · outbound

This paper cites A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification A Good Prompt Is Worth Millions of Parameters: Low-resource Prompt-based Learning for Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.615953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.615953Z digest=sha256:394c6989f32f90a75d729bb52d5f07dda83fd1a48d415701d9d0c32f3891e4ff

Observation f62f1d8d-0862-4ace-b748-8f52fc79df2b · outbound

This paper cites GPT-4 Technical Report.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification GPT-4 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.612246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.612246Z digest=sha256:a5d90603201e1a46faab72ea4e0c82a3e91b89c99cc64b9b321045ff08f96ee7

Observation 225d9c59-5dd9-45cd-b148-00b952c19f18 · outbound

This paper cites ”Learning transferable visual models from natural language supervision.” International conference on machine learning.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Learning transferable visual models from natural language supervision.” International conference on machine learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.980319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.623829Z digest=sha256:09e8be308c2ffd5570e678d5f6497dc1ad36677feae0d2881dadadb531c57bd9

Observation 7717f753-d9ef-4d6c-9588-99460779f6c2 · outbound

This paper cites ”Learning to prompt for vision-language models.” International Journal of Computer Vision 130.9 (2022): 2337-2348.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Learning to prompt for vision-language models.” International Journal of Computer Vision 130.9 (2022): 2337-2348

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.991427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.620230Z digest=sha256:b6b4b42ad769fa57c69c0fcafb6b7775ad31092e0a1bacfc592216870f9a0a61

Observation 6c73568e-5fa0-45fb-9416-4ab96c6eb591 · outbound

This paper cites ”Audioclip: Extending clip to image, text and au- dio.” ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Audioclip: Extending clip to image, text and au- dio.” ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.957421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.631243Z digest=sha256:548c6403672b279b962b0b6916aa3bca2938c62c4979daeb80ba634b5b18ddce

Observation 6bf72149-9438-413e-9fe5-a7cad5bc9d97 · outbound

This paper cites ”Wav2clip: Learning robust audio representa- tions from clip.” ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Wav2clip: Learning robust audio representa- tions from clip.” ICASSP 2022-2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.968920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.627338Z digest=sha256:d7f94cd738821059d29dd832d6925d19191d404981f6d12e62637b93242cb6d5

Observation fc8281e9-2d77-4a14-82f7-a70ae672c784 · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.638165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.638165Z digest=sha256:bf26b502fae1abec946f87e5bcc969c0d227e2714a6cf73f50ba2d9ccdd2e5ae

Observation 854e330a-e655-4613-b919-27553f9e5fe0 · outbound

This paper cites ”Clap learning audio concepts from natural language supervision.” ICASSP 2023-2023 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Clap learning audio concepts from natural language supervision.” ICASSP 2023-2023 IEEE International Confer- ence on Acoustics, Speech and Signal Processing (ICASSP)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.945615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.634776Z digest=sha256:5de3825202a93678905caa3bcf029f78575e80882a754cbbcce8fabb5b6ea696

Observation 5bdcc50c-219e-4801-b8fe-d411431c881b · outbound

This paper cites ”Maple: Multi-modal prompt learn- ing.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Maple: Multi-modal prompt learn- ing.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.931765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.646340Z digest=sha256:e7f147af07f1dbe3805f76db1dde69aeefcc0595a38c945f8b9bc39ee0a56efe

Observation ab12a8f3-5809-46b3-a42c-439e1f7797b8 · outbound

This paper cites CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification CompA: Addressing the Gap in Compositional Reasoning in Audio-Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.642312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.642312Z digest=sha256:b484d6bce25c567a9ef49aa19b63a9b472ce4e2578a4f7b6bc4f496a335bd7fc

Observation 74e3b013-9c6d-4c11-b848-c853725a3259 · outbound

This paper cites ”Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery.” Advances in Neural Information Processing Systems 36 (2024).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Hard prompts made easy: Gradient-based discrete optimization for prompt tuning and discovery.” Advances in Neural Information Processing Systems 36 (2024)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.905785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.653503Z digest=sha256:b086ccb44026e0bf9ecbd313be02ca7ef7ca45857b04df19b82ae872d330690c

Observation 98dbaab5-52fb-4ae4-a9f7-b14c5fc28df3 · outbound

This paper cites ”A survey of audio classification using deep learning.” IEEE Access (2023).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”A survey of audio classification using deep learning.” IEEE Access (2023)

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.919081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.649755Z digest=sha256:6452b65f10c66270ef1e6cf378a739a5c78381301b3c35964ede1244a478a15a

Observation 606ae98f-7685-4978-8fdc-e0c19507394f · outbound

This paper cites Adapting Language-Audio Models as Few-Shot Audio Learners.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification Adapting Language-Audio Models as Few-Shot Audio Learners

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.661121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.661121Z digest=sha256:d652be8407b13faa39cf592b97fd06f4027146626e4df813fdb7f15286ac4d75

Observation 3c9f63b6-f367-4d26-971c-175c8f69ad61 · outbound

This paper cites ”Audio-Free Prompt Tuning for Language-Audio Models.” ICASSP 2024-2024 IEEE In- ternational Conference on Acoustics, Speech and Signal Processing (ICASSP).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Audio-Free Prompt Tuning for Language-Audio Models.” ICASSP 2024-2024 IEEE In- ternational Conference on Acoustics, Speech and Signal Processing (ICASSP)

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.892140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.657173Z digest=sha256:3266b38c44bcce650eb75dfbfbb8040bdc363973aec2dc7442006e6434ce9170

Observation 4f502568-90fe-46f1-ab8c-ac0e1c6c6272 · outbound

This paper cites A Survey of Multimodal Large Language Model from A Data-centric Perspective.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification A Survey of Multimodal Large Language Model from A Data-centric Perspective

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.668961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.668961Z digest=sha256:7c920619fcdd7b563ffda6ce3ba13c863e8cbc6acc6720bc99a8612f104c6f49

Observation c5175c41-0b58-4570-a3a1-fc17d45c4611 · outbound

This paper cites ”Multimodal large language models: A survey.” 2023 IEEE International Conference on Big Data (BigData).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Multimodal large language models: A survey.” 2023 IEEE International Conference on Big Data (BigData)

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:57:23.879355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.665175Z digest=sha256:9d6445d8eba96090c226f2e3a44838747d412158a0d8e95226e5b1b671dea667

Observation 806f129c-24c2-4a50-bb8f-8c3a2313824c · outbound

This paper cites Grounding Multimodal Large Language Models in Actions.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification Grounding Multimodal Large Language Models in Actions

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-10T22:57:23.729374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T22:57:23.676860Z digest=sha256:98e24554f76850695b2315e78a9c1615400d85df8ba3b9b86393443db62befd6

Observation 3fe66244-c858-49a3-90f3-a4f1b055d70c · outbound

This paper cites ”Efficient multimodal large language models: A survey.” arXiv preprint arXiv:2405.10739 (2024).

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification ”Efficient multimodal large language models: A survey.” arXiv preprint arXiv:2405.10739 (2024)

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.673281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.673281Z digest=sha256:ed9b780252e63d19272460f81a524a92770edc51e2b22fa863055d4b4a82b4e3

Observation 8e3ec007-f0cb-4712-9512-b4af9f4661a5 · outbound

This paper cites A Review of Multi-Modal Large Language and Vision Models.

TSPE: Task-Specific Prompt Ensemble for Improved Zero-Shot Audio Classification A Review of Multi-Modal Large Language and Vision Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T22:57:23.680640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:57:23.680640Z digest=sha256:5e64f9c77abb9ec31b7e4a1147754b2ec58d7779d19f3ab56e020d99f02b3ddb

Pith citing papers

No inbound Pith citation observations are available.