Pith. sign in

Paper Citation Record · LEDGER

AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 2 inbound Pith citation observations for arXiv:2407.07801.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.07801 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 2 of 2 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:25:31.044527Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T21:52:09.528067Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d77a8053-30e2-4f0f-a42a-4f421eb3052c · inbound

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model cites this paper.

Hearing from Silence: Reasoning Audio Descriptions from Silent Videos via Vision-Language Model AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:25:31.044527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:25:31.044527Z digest=sha256:56bb6278054679918e5d2e79cfc08d4e7c3bba0e753a187563145f57284a07b4

Observation e045ada3-9866-492a-9455-47c4ad783cc4 · inbound

Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation cites this paper.

Mettle: Meta-Token Learning for Memory-Efficient Audio-Visual Adaptation AVCap: Leveraging Audio-Visual Features as Text Tokens for Captioning

Reference 2021

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:52:09.610064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:52:05.884792Z digest=sha256:37d07db72706dd401031eaab52bbbce7d9416511a018c5b99e8200cf98700dd4