Pith. sign in

Paper Citation Record · LEDGER

Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

As of 24 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2407.04416.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.04416 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T10:46:23.511855Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T00:04:22.536675Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9ee79af6-7f1e-4d93-86f8-67acf816466b · inbound

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models cites this paper.

AudioSetCaps: An Enriched Audio-Caption Dataset using Automated Generation Pipeline with Large Audio and Language Models Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T10:46:23.511855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:46:23.511855Z digest=sha256:66091d5b6ed58e8e25def7c9491f01c36236afde76dd8248e01d293118450e65

Observation c81f82fa-c6b0-4067-9217-595cf16890a9 · inbound

ETTA: Elucidating the Design Space of Text-to-Audio Models cites this paper.

ETTA: Elucidating the Design Space of Text-to-Audio Models Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T00:45:19.911892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:45:19.911892Z digest=sha256:7b56e728b12b5e66cfe0a6186e8444e8235c05f03c11de19683c4bdc212d7b56

Observation 00306dbf-1ccc-40bd-9615-44556ec9a28a · inbound

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey cites this paper.

Audio-Language Models for Audio-Centric Tasks: A Systematic Survey Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 135

Resolution
unresolved
no resolver link, observed 2026-08-10T14:36:19.799975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:36:19.799975Z digest=sha256:61c5f14bb7cc9b8d90f3ad3c4cef9d12eccc54078d93bbafee4119c359d6ce2e

Observation f491f5a3-11f0-4614-94de-1a71b8fca4f4 · inbound

Sounding that Object: Interactive Object-Aware Image to Audio Generation cites this paper.

Sounding that Object: Interactive Object-Aware Image to Audio Generation Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:53:17.755251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:53:17.755251Z digest=sha256:62ca0203e53071c1575480048f6b84ceed94a2d1aa61503e85f4ff8cebf59005

Observation 400238fd-6039-465d-a1a4-3597ab50cd9d · inbound

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos cites this paper.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:46.490467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:46.490467Z digest=sha256:5c0c0d8059e54b3a33380d779c2de6736f5b13ead2ce35e04a8e0b28ce0096cd

Observation 42eee5fb-72fb-4fed-b827-dc747763a45f · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 67

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.538025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:d4544fcc11d52b148b84d2e09858ab41e47c71fec34161add7afc5904cb10b02

Observation 9b7c5262-0aff-4d9f-bd34-95b37b4fdf84 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Sound-VECaps: Improving Audio Generation with Visual Enhanced Captions

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:d6f05bf10203d22319e8a2aa284d911cf5c30b3213f88e05fe2f57d84a299485