Pith. sign in

Paper Citation Record · LEDGER

Contrastive Audio-Visual Masked Autoencoder

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2210.07839.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2210.07839 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:20:16.505610Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T05:56:39.834913Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4b5530ac-5d7b-4943-9ce5-9a3b7aae252e · inbound

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning cites this paper.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Contrastive Audio-Visual Masked Autoencoder

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.505610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.505610Z digest=sha256:32692fded7292a41c761d13ccb689f5610254875752f2af54c37d98d8e7fbaf2

Observation b8a4800c-84d1-4295-8a4d-20dd0bc26f7b · inbound

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation cites this paper.

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation Contrastive Audio-Visual Masked Autoencoder

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:39:53.588345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:39:53.588345Z digest=sha256:96edba483d36b6723264d81372f8c31c3e696dbc0166feb16605473c5a52860c

Observation 8696b57a-3e26-4303-8f96-eccab980b3f1 · inbound

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition cites this paper.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Contrastive Audio-Visual Masked Autoencoder

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.507305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.507305Z digest=sha256:f8866ee2d2499637d947e21ce4e258691584c6e4b887897f5bcb8eafcf3a3e00

Observation abba1354-4e34-4433-88c0-93456f3e50af · inbound

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis cites this paper.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Contrastive Audio-Visual Masked Autoencoder

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.781335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.781335Z digest=sha256:947d9d11a89fd3ac54eb2c6fa8682529f18bc967c16c59d256e827f025f67cc4

Observation 90dd38b2-8ab7-4064-bc64-7caa43145675 · inbound

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos cites this paper.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Contrastive Audio-Visual Masked Autoencoder

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:45.037290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:45.037290Z digest=sha256:79c427183e81ef470f748e844c6b9dee664a9d234a65a5411e4b095d49218a61

Observation 0b1f7c7a-b55a-4d1e-b348-b32c898633d5 · inbound

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation cites this paper.

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Contrastive Audio-Visual Masked Autoencoder

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:09.947756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:09.947756Z digest=sha256:961d2f478e1ce0d98d322daacd9ba55473a2d8746e9ddfe654c1996c1047e684

Observation 85d2e9d4-b713-46dd-8943-ff6ff105517f · inbound

Audio-Visual Continual Test-Time Adaptation without Forgetting cites this paper.

Audio-Visual Continual Test-Time Adaptation without Forgetting Contrastive Audio-Visual Masked Autoencoder

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:23.442235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:23.442235Z digest=sha256:97c5aa1444cb02ab3cbd0ca75d6d694c32f0569927a96c919bbd7eccbff7060d

Observation a8c91a3a-8934-4f4c-83ae-7ef929538a5f · inbound

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling cites this paper.

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling Contrastive Audio-Visual Masked Autoencoder

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.072872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T09:22:10.483263Z digest=sha256:b612b66bc3731b60b375b6166fd0db1dffa11e6ebf5037b33f1071f645583890

Observation ab8cf168-448e-4b94-97af-331555ecf050 · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Contrastive Audio-Visual Masked Autoencoder

Reference 106

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T05:56:39.836610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:8eb309a3903a9f7e9cc1ec6e513593e1ab8a5fd2f536ca8c068568295d6341cc