Pith. sign in

Paper Citation Record · LEDGER

Audio Captioning Transformer

As of 24 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 4 inbound Pith citation observations for arXiv:2107.09817.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2107.09817 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:03:38.077680Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:20:18.135780Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4df7bf77-24f9-40a4-9b94-593b97e65ecd · inbound

Piano Transcription by Hierarchical Language Modeling with Pretrained Roll-based Encoders cites this paper.

Piano Transcription by Hierarchical Language Modeling with Pretrained Roll-based Encoders Audio Captioning Transformer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T22:03:38.077680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:03:38.077680Z digest=sha256:160bf57d2bcbcb56c3076fe298794d47e5b13445294b13213058260527068122

Observation 40e40877-d0cb-406e-b5cd-c70b24a183e4 · inbound

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey cites this paper.

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Audio Captioning Transformer

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-10T14:36:19.608414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:36:19.608414Z digest=sha256:0bc6689b93ee873004eeec732026b05d9511b88a91f26198bdf2e8375dc17ffb

Observation b8e6e2e4-5745-481d-abaa-bceab90c8304 · inbound

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning cites this paper.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Audio Captioning Transformer

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:20:18.267155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T13:20:15.714734Z digest=sha256:45ff42cd863706fefcb2f3e966c40463aed89272c36d43dea1fd4296bcccbb75

Observation 69b1d4a8-4ef6-45f9-9aa4-545aa54c80d8 · inbound

LGFNet: A CTC-Guided Local-Global Fusion Framework for Single-Channel Sleep Staging cites this paper.

LGFNet: A CTC-Guided Local-Global Fusion Framework for Single-Channel Sleep Staging Audio Captioning Transformer

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T03:12:06.514554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T03:12:06.514554Z digest=sha256:e942f508958f5ad8e29a091c495a29b6566f0f1f14e8d1980ed2b68a92651f0c