Pith. sign in

Paper Citation Record · LEDGER

Language-based Audio Retrieval with Co-Attention Networks

As of 11 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 0 inbound Pith citation observations for arXiv:2412.20914.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20914 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:11:53.338754Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9aa033aa-f697-4706-9212-d2e9f86151ef · outbound

This paper cites Language-based Audio Retrieval Task in DCASE 2022 Challenge.

Language-based Audio Retrieval with Co-Attention Networks Language-based Audio Retrieval Task in DCASE 2022 Challenge

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-10T23:11:53.478857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.207613Z digest=sha256:d7983ae64059eedd71ec2697fd1749e37d3f7b3d3975bfd3cf33eb9696a7def0

Observation 4da52f8e-65e7-4c62-8056-e2f46ac56379 · outbound

This paper cites Language- based audio retrieval with gpt-augmented captions and self-attended audio clips,.

Language-based Audio Retrieval with Co-Attention Networks Language- based audio retrieval with gpt-augmented captions and self-attended audio clips,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.805965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.213405Z digest=sha256:c47e56b361b4e7feff25fc6be9b7683995a4152b5ffbd25319e2e6eae6702aa1

Observation 0e1d112d-5f07-49bc-956c-61183aea2df7 · outbound

This paper cites Dynamic modality interaction modeling for image-text retrieval,.

Language-based Audio Retrieval with Co-Attention Networks Dynamic modality interaction modeling for image-text retrieval,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.791604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.218440Z digest=sha256:e40148b9b623bea84f73516b5a40241cd86aad9b2fb45dec1d6f82b821f867c5

Observation 01607172-21c6-4127-a020-e504e1cf9779 · outbound

This paper cites Look, listen, and attend: Co-attention network for self-supervised audio-visual representa- tion learning,.

Language-based Audio Retrieval with Co-Attention Networks Look, listen, and attend: Co-attention network for self-supervised audio-visual representa- tion learning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.776723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.223460Z digest=sha256:acd69345c5d396cdd23af040b9edc6ca12b54c071e862543d2d6492f80b827b2

Observation 680ed02a-0ab9-4f34-bf58-121edbe1a70e · outbound

This paper cites Audio-text retrieval in context,.

Language-based Audio Retrieval with Co-Attention Networks Audio-text retrieval in context,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.761242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.228603Z digest=sha256:cf0f6566ad470338150aea12f8307130024e286034f45771d1a83482c9ede709

Observation 56935820-ef78-46a3-9b31-ff43499ce979 · outbound

This paper cites Improving text-audio retrieval by text- aware attention pooling and prior matrix revised loss,.

Language-based Audio Retrieval with Co-Attention Networks Improving text-audio retrieval by text- aware attention pooling and prior matrix revised loss,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.746653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.233420Z digest=sha256:370532e357044e65739861965d153492fde52c5749634b128c424f5141a5aed8

Observation 5df9c68c-5113-4144-b9e8-65cea9a02dc4 · outbound

This paper cites A ResNet-Based CLIP Text-to-Audio Retrieval System for DCASE Challenge 2022 Task 6B,.

Language-based Audio Retrieval with Co-Attention Networks A ResNet-Based CLIP Text-to-Audio Retrieval System for DCASE Challenge 2022 Task 6B,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.731345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.238887Z digest=sha256:9e8cb35eb4ab569458dfc6f775af4954d0d14c5df4b5bb9f63d1029bd8d711f1

Observation c1957c6f-84a6-4e42-9fde-da08ffb3f409 · outbound

This paper cites Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering,.

Language-based Audio Retrieval with Co-Attention Networks Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.715522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.243606Z digest=sha256:443d9cce8f45e62057ccfffb13340a5026e76aa02391e1b200321f1081456f32

Observation dd9b1ac8-f38d-49ec-ace9-3cf8cf4551c1 · outbound

This paper cites Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss.

Language-based Audio Retrieval with Co-Attention Networks Improving Video-Text Retrieval by Multi-Stream Corpus Alignment and Dual Softmax Loss

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:11:53.248081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:11:53.248081Z digest=sha256:b534380816ccc9675968431c1f0f06529db218f9facc4280f620fc90d98d8c46

Observation a3bb2161-e228-4d1e-b0d2-d8bc28753664 · outbound

This paper cites Audio-text retrieval in context,.

Language-based Audio Retrieval with Co-Attention Networks Audio-text retrieval in context,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.700118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.253195Z digest=sha256:7eeea9fcf12782a68982394e4f4548cac63aae8d69830e85c20ef8edd353c446

Observation 0868d13b-e32d-4ced-bbda-565045ea2662 · outbound

This paper cites Audio retrieval with natural language queries: A benchmark study,.

Language-based Audio Retrieval with Co-Attention Networks Audio retrieval with natural language queries: A benchmark study,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.685697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.257714Z digest=sha256:77addaa5f879bcaacbd3847a1b8229efdb4b9a5d229a15db9ebc596a115f3ce7

Observation 82e821b3-cc4b-4e6c-8b3c-f72bf39af2c8 · outbound

This paper cites Attentive Pooling Networks.

Language-based Audio Retrieval with Co-Attention Networks Attentive Pooling Networks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:11:53.262353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:11:53.262353Z digest=sha256:47a634160b50deda9b5920ae3a2f7ea2826a3c3f14cf7ba9f61d57cf5f71d888

Observation c27f7d74-445b-45c4-b2f6-b0c9f2c63b0c · outbound

This paper cites Multi-pointer co-attention networks for recommendation,.

Language-based Audio Retrieval with Co-Attention Networks Multi-pointer co-attention networks for recommendation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.671184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.267464Z digest=sha256:86a731933ba9bd2e374ea19bef24cae56d932445f9489d87a8dcb76c6ea25eda

Observation 5e94300f-3af7-499a-abee-bf95cf38df05 · outbound

This paper cites Query by example of audio signals using euclidean distance between gaussian mixture models,.

Language-based Audio Retrieval with Co-Attention Networks Query by example of audio signals using euclidean distance between gaussian mixture models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.654247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.272748Z digest=sha256:27247ff8606aec9953c4271b8436a8a806aa30a691d22b07f0c9fcbb218d9897

Observation 0bc3020f-ef14-4fcf-89c8-209b58750fb9 · outbound

This paper cites Semantic-audio retrieval,.

Language-based Audio Retrieval with Co-Attention Networks Semantic-audio retrieval,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.639148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.277373Z digest=sha256:58c6baf02ee604a5b76e837bc4a247265bc989134ffef765ca4e476b6fba55b0

Observation a7f1cc83-dc32-4ac3-b925-1857c0d3c73f · outbound

This paper cites Music information retrieval using social tags and audio,.

Language-based Audio Retrieval with Co-Attention Networks Music information retrieval using social tags and audio,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.623933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.282220Z digest=sha256:04ac1ed4810288b7461e68cb1d808b9290ae7d45236898b3326c3481093138f3

Observation 9de22329-1f1d-4bdf-95d2-6ec5904131f1 · outbound

This paper cites Deep visual-semantic alignments for generating image descriptions,.

Language-based Audio Retrieval with Co-Attention Networks Deep visual-semantic alignments for generating image descriptions,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.608077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.286769Z digest=sha256:1f5e039516580307d837f0947fdb07f6626054bbd0bfd48ae5cfb9edd0c56cfe

Observation 5b0165e1-e89b-434c-a70e-54939682360d · outbound

This paper cites Cross modal audio search and retrieval with joint embeddings based on text and audio,.

Language-based Audio Retrieval with Co-Attention Networks Cross modal audio search and retrieval with joint embeddings based on text and audio,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.593365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.291376Z digest=sha256:313cfb5a26ac21aefb7de3cf1c4f99adbb6eb611ebf6111d3a16e9964c680f9b

Observation cfcbfd94-a3d3-4c24-855f-109a54f2ec06 · outbound

This paper cites Clap learning audio concepts from natural language supervision,.

Language-based Audio Retrieval with Co-Attention Networks Clap learning audio concepts from natural language supervision,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.578328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.296009Z digest=sha256:6437a6d81effb6cf7c4a553a5a19bbc470411ae8c33da56528d3c6138767e826

Observation 9c28ddac-6366-4f65-b40f-ea4f8d17769c · outbound

This paper cites Attention Is All You Need.

Language-based Audio Retrieval with Co-Attention Networks Attention Is All You Need

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:11:53.300501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:11:53.300501Z digest=sha256:93252f460dbc02694496b0c677187814ced1a9bff1ba885fbf296683b9dbc5b2

Observation dd487599-81a5-488c-9df2-821922247c59 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Language-based Audio Retrieval with Co-Attention Networks An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:11:53.305661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:11:53.305661Z digest=sha256:a09c39216f32f2218e0dfa42c0ade542427e325ee98d08817afd64c8891e972e

Observation 5d9f756f-3af9-4b6a-9aa0-0af87cade58d · outbound

This paper cites Dynamic Coattention Networks For Question Answering.

Language-based Audio Retrieval with Co-Attention Networks Dynamic Coattention Networks For Question Answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T23:11:53.310693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:11:53.310693Z digest=sha256:9d99b75303e9ea02dfbda2b31cdcbc368e1f4e60b8c3df6d5fd3346a3a58ba2e

Observation 4aa8024a-b191-4e11-bc4e-c9706e32de90 · outbound

This paper cites Co-attention network with label embedding for text classification,.

Language-based Audio Retrieval with Co-Attention Networks Co-attention network with label embedding for text classification,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.560744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.315763Z digest=sha256:9dc3eae77e001d6ae802824bf73ca5b97f41311e751c3b85f713d232af475cd7

Observation 4161afbd-e238-4fd1-9995-471a30c11b46 · outbound

This paper cites Attentive interactive neural networks for answer selection in community question answering,.

Language-based Audio Retrieval with Co-Attention Networks Attentive interactive neural networks for answer selection in community question answering,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.544650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.320357Z digest=sha256:f52178261945021e693a36110df5fbf67d8d1186ee67eecbf86d74d3e6445b47

Observation ed42d913-afb8-423e-b632-4a4206628043 · outbound

This paper cites Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,.

Language-based Audio Retrieval with Co-Attention Networks Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.527825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.324944Z digest=sha256:91307f0a4cc8ab6865a05a8cfaf284245be69248f22495487b6d9d5a138abdea

Observation ab6ba47b-87d3-4c06-ad22-2fa9a56c3afc · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Language-based Audio Retrieval with Co-Attention Networks RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:11:53.329420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:11:53.329420Z digest=sha256:b7d0432096e13eb02d3ab886e78b9e9b0d03987150ad60a3a0e7c159c786afc1

Observation 0ab219a1-5345-4849-83c7-4875f6f135fe · outbound

This paper cites Clotho: an audio captioning dataset,.

Language-based Audio Retrieval with Co-Attention Networks Clotho: an audio captioning dataset,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.511284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.334182Z digest=sha256:2199d2911cbe438a1c85ff8a25b7f634ed946a7bb3b9e7d2365ea5eeadd0683a

Observation 4b35fb4c-967b-42d0-9c96-eda303bf2ed6 · outbound

This paper cites AudioCaps: Generating captions for audios in the wild,.

Language-based Audio Retrieval with Co-Attention Networks AudioCaps: Generating captions for audios in the wild,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:11:53.495327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T23:11:53.338754Z digest=sha256:83cf1586c3bdcb600ae832e38d0bfdf91db0a293c5e3a2ff54f069e876bd69ef

Pith citing papers

No inbound Pith citation observations are available.