Pith. sign in

Paper Citation Record · LEDGER

A Survey on Audio Synthesis and Audio-Visual Multimodal Processing

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 3 inbound Pith citation observations for arXiv:2108.00443.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2108.00443 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 3 of 3 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:03:21.327224Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-04T21:25:27.812803Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3c7d422b-5187-4053-8ea8-18ee6b274945 · inbound

Learning from Silence and Noise for Visual Sound Source Localization cites this paper.

Learning from Silence and Noise for Visual Sound Source Localization A Survey on Audio Synthesis and Audio-Visual Multimodal Processing

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T14:03:21.327224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:03:21.327224Z digest=sha256:07863fc1a24fa66f7fbe3ebd80b390ef19789136d675e8f9d5a3e334e0ee8e85

Observation b956306a-aaff-4ab1-bd9c-34721c77bd74 · inbound

Effectively obtaining acoustic, visual and textual data from videos cites this paper.

Effectively obtaining acoustic, visual and textual data from videos A Survey on Audio Synthesis and Audio-Visual Multimodal Processing

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-05T05:01:36.985927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:01:36.985927Z digest=sha256:d9dd6652c9021260b89613f5f78bd02f19398ddfa9074b77a6cdfc42e68267a6

Observation a0fc0fbc-155f-4321-8f35-c0266598a0db · inbound

Testing chatbots on the creation of encoders for audio conditioned image generation cites this paper.

Testing chatbots on the creation of encoders for audio conditioned image generation A Survey on Audio Synthesis and Audio-Visual Multimodal Processing

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-08-04T21:25:27.823348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-04T21:25:27.206894Z digest=sha256:a3bb59200bb99e994322f8d8d74f013a0ef2441285d89e4561a01d44e3486e86