Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T18:02:17.907076Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 1 inbound Pith citation observation for arXiv:2411.12058.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T18:02:17.907076Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T05:12:48.603860Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T05:12:48.684609Z
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 84a94cad-1bac-40a1-a253-4d03a51f6d88 · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6d603cf-8af9-4c41-b8bd-d6a9861370f4 · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers The claude 3 model family: Opus, sonnet, haiku
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9ecb54c4-d082-4cb3-8c54-6e78759b9407 · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb34f316-b7c8-493c-a508-18ed37a2ec27 · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers Convolutional recurrent neural networks for music classification, 2016
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 09169264-1425-4151-87d9-911d618b1ced · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers Pengi: An audio language model for audio tasks, 2024
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fa67ffe7-38db-4906-b790-6ae62503d303 · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers A survey on in-context learning, 2024
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe6ed44f-6557-40e5-b82e-fd172199735e · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers Clotho: an audio captioning dataset
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a5127ce9-90dd-4d0a-9bdb-246047bd707d · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers A survey of vision-language pre-trained models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bfb2794c-6194-45a8-84c0-2282716afd6d · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd2d327e-83b2-438e-9962-52dbbeeb07f7 · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers Gama: A large audio-language model with advanced audio understanding and complex reasoning abilities, 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 56ecca0c-d76a-4ebc-9fcc-1a6ac6b0c67a · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers Ast: Audio spectrogram transformer, 2021
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8c941e31-7bbe-41eb-a9a8-1bcbd1041893 · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 57caf24e-b710-47b1-bec7-d3cdef154cbe · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers Effectiveness assessment of recent large vision-language models, 2024
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5de494b3-ae2a-4ad2-b90e-dc859b8169b5 · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers Esc: Dataset for environmental sound classification
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80d5344d-843b-4a97-938b-4092e8bf42ed · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers End-to-end learning for music audio tagging at scale
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b5c505f-5176-40b8-b079-3bd78ea8b4d0 · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers Salmonn: Towards generic hearing abilities for large language models, 2024
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7bd0f064-9bac-4d7e-a2d3-bbef2dc266fd · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers Gemini: A Family of Highly Capable Multimodal Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 276b9982-83a7-4438-9ddf-cad58e3b63f3 · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi, 2024
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed514dd4-678b-4298-8268-e0147a4ee7f6 · outbound
Vision Language Models Are Few-Shot Audio Spectrogram Classifiers "" 2 { 3
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 23b033b1-833f-4248-8e28-7bdae6be5854 · inbound
Knowledge-Augmented Vision Language Models for Underwater Bioacoustic Spectrogram Analysis Vision Language Models Are Few-Shot Audio Spectrogram Classifiers
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.