Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:14:57.714512Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2505.05335.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:14:57.714512Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T10:23:37.367700Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T21:55:00.563579Z
20 of 20 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 967a5c6c-c212-49f0-a4a7-7c49e9fabb8d · outbound
FLAM: Frame-Wise Language-Audio Modeling keyword, tag
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c90fbfa0-e57c-493d-b1fc-ad9a8689d44d · outbound
FLAM: Frame-Wise Language-Audio Modeling Clap learning audio concepts from natural language su- pervision
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f8aa9a3-8b8e-4f56-a02f-6fcddb605a64 · outbound
FLAM: Frame-Wise Language-Audio Modeling P., Fonseca, E., Jansen, A., Liu, C., Moore, R
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dc47e202-dfa9-4f5a-b9c3-80092f6f0dd9 · outbound
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b80bfbc1-7662-42ea-b0ca-ea5e83ff1bab · outbound
FLAM: Frame-Wise Language-Audio Modeling D., Kim, B., Lee, H., and Kim, G
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e68316b9-b23e-4c42-b5ba-7ba129a8b19d · outbound
FLAM: Frame-Wise Language-Audio Modeling RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4db5f27a-8455-499a-a68d-67a314c8fd55 · outbound
FLAM: Frame-Wise Language-Audio Modeling Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81bd6d2f-1260-43bc-aad7-f341b9e331be · outbound
FLAM: Frame-Wise Language-Audio Modeling Sound event detection in synthetic domestic environments
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9a847350-0a3b-47f8-a3f7-77897add420e · outbound
FLAM: Frame-Wise Language-Audio Modeling P., and Salamon, J
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12f7623d-c9a0-4803-b172-a3bc76e3cf64 · outbound
FLAM: Frame-Wise Language-Audio Modeling Large-scale contrastive language- audio pretraining with feature fusion and keyword-to- caption augmentation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 732d6036-e9da-45dc-aad2-a9ff31ac24fd · outbound
FLAM: Frame-Wise Language-Audio Modeling Towards Weakly Supervised Text-to-Audio Grounding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 80215732-3d7b-4d1e-b4dd-04592a0a027c · outbound
FLAM: Frame-Wise Language-Audio Modeling T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7354fde-6d19-4af0-828f-a21dbcc5c5b4 · outbound
FLAM: Frame-Wise Language-Audio Modeling Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 79de48b6-c08d-4bde-a116-e7b818a4ac28 · outbound
FLAM: Frame-Wise Language-Audio Modeling Clotho: An audio 9 FLAM: Frame-Wise Language-Audio Modeling captioning dataset
Reference 2018
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 13c7ae50-4fd3-4ded-87a9-f763dcb78a89 · outbound
FLAM: Frame-Wise Language-Audio Modeling Mean teacher convolution system for dcase 2018 task
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 606fde2d-a84f-49f3-8823-1c3621ec4a6c · outbound
FLAM: Frame-Wise Language-Audio Modeling Threshold independent evaluation of sound event detection scores
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 998f6195-1d0b-4116-9d24-ae86f8738651 · outbound
FLAM: Frame-Wise Language-Audio Modeling Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 78374fd8-036b-4c55-bc37-08dacc6bbb98 · outbound
FLAM: Frame-Wise Language-Audio Modeling DCASE 2024 Task 4: Sound Event Detection with Heterogeneous Data and Missing Labels
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c97f3cde-cbda-41f2-8853-7725e4642df8 · outbound
FLAM: Frame-Wise Language-Audio Modeling Representation Learning with Contrastive Predictive Coding
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f0b86cf-312a-40b0-8da0-d88b1a84335d · outbound
FLAM: Frame-Wise Language-Audio Modeling BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5ee9494-cf1b-4d4a-9f89-20a716f4fc31 · inbound
Melody-Lyrics Matching with Contrastive Alignment Loss FLAM: Frame-Wise Language-Audio Modeling
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3f7fd21-81b5-464d-82dc-1d6b272b7750 · inbound
Auditory Intelligence: Understanding the World Through Sound FLAM: Frame-Wise Language-Audio Modeling
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.