Pith. sign in

Paper Citation Record · LEDGER

FLAM: Frame-Wise Language-Audio Modeling

As of 16 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 2 inbound Pith citation observations for arXiv:2505.05335.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05335 v2

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:14:57.714512Z

measured 22 of 22 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:23:37.367700Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T21:55:00.563579Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 967a5c6c-c212-49f0-a4a7-7c49e9fabb8d · outbound

This paper cites keyword, tag.

FLAM: Frame-Wise Language-Audio Modeling keyword, tag

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.840610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:14:57.709975Z digest=sha256:1d5737c52e248f5c770e54e31b48fc83162946d9f36d0342bcca6e289d9a2b7b

Observation c90fbfa0-e57c-493d-b1fc-ad9a8689d44d · outbound

This paper cites Clap learning audio concepts from natural language su- pervision.

FLAM: Frame-Wise Language-Audio Modeling Clap learning audio concepts from natural language su- pervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.659853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.659853Z digest=sha256:e598ee102395a70661509c1e82452e5b4c53a6f3e00d64856f9162db63128479

Observation 2f8aa9a3-8b8e-4f56-a02f-6fcddb605a64 · outbound

This paper cites P., Fonseca, E., Jansen, A., Liu, C., Moore, R.

FLAM: Frame-Wise Language-Audio Modeling P., Fonseca, E., Jansen, A., Liu, C., Moore, R

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.902565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:14:57.663216Z digest=sha256:fb1a26a61ff7f064694af1f00a7010b1c0aad410dacc5eea501067784a3d3c52

Observation dc47e202-dfa9-4f5a-b9c3-80092f6f0dd9 · outbound

This paper cites Mixtral of Experts.

FLAM: Frame-Wise Language-Audio Modeling Mixtral of Experts

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.670465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.670465Z digest=sha256:d3856da36a10ec8f5e53219d04f9bd5c317b0f9753dbae68a20418e4fc1efe25

Observation b80bfbc1-7662-42ea-b0ca-ea5e83ff1bab · outbound

This paper cites D., Kim, B., Lee, H., and Kim, G.

FLAM: Frame-Wise Language-Audio Modeling D., Kim, B., Lee, H., and Kim, G

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.674591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.674591Z digest=sha256:db616b00d67f2619f0e0c046b9d5c6cbb2a9e9d3a181169fcde938a5788629f8

Observation e68316b9-b23e-4c42-b5ba-7ba129a8b19d · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

FLAM: Frame-Wise Language-Audio Modeling RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.677909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.677909Z digest=sha256:78dcdd73eb742667a074e98afc52093a25a39cfbc79035e34c4c69586df9db18

Observation 4db5f27a-8455-499a-a68d-67a314c8fd55 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

FLAM: Frame-Wise Language-Audio Modeling Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.681336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.681336Z digest=sha256:7b0216af1476f38df557d10425536ab89c0525cb208322c7296218f51c718dfd

Observation 81bd6d2f-1260-43bc-aad7-f341b9e331be · outbound

This paper cites Sound event detection in synthetic domestic environments.

FLAM: Frame-Wise Language-Audio Modeling Sound event detection in synthetic domestic environments

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.871860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:14:57.685805Z digest=sha256:0e8087592c9f44c15ad9ef4b21c9f6553c0bd73a286b5aea4a24d1326c3fc1ca

Observation 9a847350-0a3b-47f8-a3f7-77897add420e · outbound

This paper cites P., and Salamon, J.

FLAM: Frame-Wise Language-Audio Modeling P., and Salamon, J

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.692841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.692841Z digest=sha256:bd6c44e646a6e7cc3c3a17e29b338f5359ad44ff73e6ca6c1b2ff4ddd98d2db5

Observation 12f7623d-c9a0-4803-b172-a3bc76e3cf64 · outbound

This paper cites Large-scale contrastive language- audio pretraining with feature fusion and keyword-to- caption augmentation.

FLAM: Frame-Wise Language-Audio Modeling Large-scale contrastive language- audio pretraining with feature fusion and keyword-to- caption augmentation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.853356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:14:57.697076Z digest=sha256:cbf3cce8198d4e70a015b2862d1ac024f506325aa192e81782a7e582f1d4f877

Observation 732d6036-e9da-45dc-aad2-a9ff31ac24fd · outbound

This paper cites Towards Weakly Supervised Text-to-Audio Grounding.

FLAM: Frame-Wise Language-Audio Modeling Towards Weakly Supervised Text-to-Audio Grounding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.700621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.700621Z digest=sha256:6e08cd06ca918ebe8bf4cf9ba77ec78df674dbdad8a94acbca20ed79a1e63e4e

Observation 80215732-3d7b-4d1e-b4dd-04592a0a027c · outbound

This paper cites T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining.

FLAM: Frame-Wise Language-Audio Modeling T-CLAP: Temporal-Enhanced Contrastive Language-Audio Pretraining

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.704987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.704987Z digest=sha256:75f1997487824176452f276231f37bcaacad9def63fffbbc5392c84884d5e278

Observation c7354fde-6d19-4af0-828f-a21dbcc5c5b4 · outbound

This paper cites an unresolved cited work.

FLAM: Frame-Wise Language-Audio Modeling Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:14:57.828514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:14:57.714512Z digest=sha256:302d82e6c411bd58321163252ac622414fe16eea8c0529ca2ed521f6808c5d1c

Observation 79de48b6-c08d-4bde-a116-e7b818a4ac28 · outbound

This paper cites Clotho: An audio 9 FLAM: Frame-Wise Language-Audio Modeling captioning dataset.

FLAM: Frame-Wise Language-Audio Modeling Clotho: An audio 9 FLAM: Frame-Wise Language-Audio Modeling captioning dataset

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.931363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:14:57.651200Z digest=sha256:da905d9f7e9472906e64b434e86d533d5788a7441ca44e4d9505c04adce51680

Observation 13c7ae50-4fd3-4ded-87a9-f763dcb78a89 · outbound

This paper cites Mean teacher convolution system for dcase 2018 task.

FLAM: Frame-Wise Language-Audio Modeling Mean teacher convolution system for dcase 2018 task

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.890929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:14:57.667433Z digest=sha256:7fdda33c33d20b3170bdef6a2c9c9a9f18045102ced14b9d591f57f0e201b3dc

Observation 606fde2d-a84f-49f3-8823-1c3621ec4a6c · outbound

This paper cites Threshold independent evaluation of sound event detection scores.

FLAM: Frame-Wise Language-Audio Modeling Threshold independent evaluation of sound event detection scores

Reference 2020

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.920000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:14:57.656329Z digest=sha256:a60d08a4d3cb639783ec0f36f5ef951910672e685b8af4c485e8a2de8003c266

Observation 998f6195-1d0b-4116-9d24-ae86f8738651 · outbound

This paper cites Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection.

FLAM: Frame-Wise Language-Audio Modeling Hts-at: A hierarchical token-semantic audio transformer for sound classification and detection

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:14:57.943055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-15T23:14:57.638635Z digest=sha256:9b67252f1f94e9f91187446d3628ae9eca563ca1182939ffcd456e2f7ba466ea

Observation 78374fd8-036b-4c55-bc37-08dacc6bbb98 · outbound

This paper cites DCASE 2024 Task 4: Sound Event Detection with Heterogeneous Data and Missing Labels.

FLAM: Frame-Wise Language-Audio Modeling DCASE 2024 Task 4: Sound Event Detection with Heterogeneous Data and Missing Labels

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.642924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.642924Z digest=sha256:0b90a6b886378ac4f5e666db8d6f50eb85a88db8b49f302ea5840facc9da0649

Observation c97f3cde-cbda-41f2-8853-7725e4642df8 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

FLAM: Frame-Wise Language-Audio Modeling Representation Learning with Contrastive Predictive Coding

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.689004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.689004Z digest=sha256:82578f10c6de4ed1f413e2ef743d30137144df4b26d9dd415b3fa616771ef97f

Observation 0f0b86cf-312a-40b0-8da0-d88b1a84335d · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

FLAM: Frame-Wise Language-Audio Modeling BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T23:14:57.647008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:14:57.647008Z digest=sha256:9470e33b928dfdad32d311db96943a928599de212a8839df8df62e97c97405a5

Pith citing papers

Observation f5ee9494-cf1b-4d4a-9f89-20a716f4fc31 · inbound

Melody-Lyrics Matching with Contrastive Alignment Loss cites this paper.

Melody-Lyrics Matching with Contrastive Alignment Loss FLAM: Frame-Wise Language-Audio Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T10:23:37.367700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:23:37.367700Z digest=sha256:f9b34835e279570233f18ab018bb1ea71a5f1de7bc2b355ec17ef9077081d3af

Observation f3f7fd21-81b5-464d-82dc-1d6b272b7750 · inbound

Auditory Intelligence: Understanding the World Through Sound cites this paper.

Auditory Intelligence: Understanding the World Through Sound FLAM: Frame-Wise Language-Audio Modeling

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:55:00.654739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-05T21:55:00.164281Z digest=sha256:d4bbebb0a5f216822708fb799134c8d68d66f78a5c68e256a84fb0f4f9d11058