Pith. sign in

Paper Citation Record · LEDGER

Contrastive Audio-Visual Masked Autoencoder

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2210.07839.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2210.07839 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:27:56.939307Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T05:56:39.834913Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation dd674aa9-a0f8-43e6-a43c-906b034b9baa · inbound

The Sound of Water: Inferring Physical Properties from Pouring Liquids cites this paper.

The Sound of Water: Inferring Physical Properties from Pouring Liquids Contrastive Audio-Visual Masked Autoencoder

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T18:52:43.055811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:52:43.055811Z digest=sha256:79780e70f8edf749a58d16cca1a423e87aaf83020710db7bd4422529ed7fbd95

Observation 2cb8e19a-25b3-495a-91d9-2549a660af15 · inbound

KDC-MAE: Knowledge Distilled Contrastive Mask Auto-Encoder cites this paper.

KDC-MAE: Knowledge Distilled Contrastive Mask Auto-Encoder Contrastive Audio-Visual Masked Autoencoder

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T17:49:24.415073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:49:24.415073Z digest=sha256:3699f35114983e4d18716d0ca99c057942d04efdce81811a0c0c2f8217614866

Observation 395dd764-c639-47b6-b6af-fccc845ea6ce · inbound

A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning cites this paper.

A Survey of Recent Advances and Challenges in Deep Audio-Visual Correlation Learning Contrastive Audio-Visual Masked Autoencoder

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:04:00.385262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:04:00.385262Z digest=sha256:b6b8caadc3373b1b9851e9d4cf59c8f1a3d288104356646c83bdb56dacb92ea1

Observation a983b572-ac08-48bc-a575-5ef48fcf5c18 · inbound

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization cites this paper.

TACO: Training-free Sound Prompted Segmentation via Semantically Constrained Audio-visual CO-factorization Contrastive Audio-Visual Masked Autoencoder

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T04:23:41.440713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:23:41.440713Z digest=sha256:9a75e1048f5bfafa87fc6dc88c2a8775e9706a8d313305289f9ed88468a507ad

Observation a89a43ff-8a4f-459a-8418-c6a3b65d07b2 · inbound

JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts cites this paper.

JoVALE: Detecting Human Actions in Video Using Audiovisual and Language Contexts Contrastive Audio-Visual Masked Autoencoder

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:55:40.498345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:55:40.498345Z digest=sha256:903e2aa2017146cc6d0e1575afd591adb187cd77303d6947ed16d471a7c6f756

Observation 4b5530ac-5d7b-4943-9ce5-9a3b7aae252e · inbound

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning cites this paper.

Mitigating Audiovisual Mismatch in Visual-Guide Audio Captioning Contrastive Audio-Visual Masked Autoencoder

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:16.505610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:16.505610Z digest=sha256:cedb6d06c19aaa66c9dfca90f11654d2e7669d75717f67a27b074986fcde5034

Observation ac127a94-68f0-4b79-b5e2-1c199391e695 · inbound

Let Your Video Listen to Your Music! cites this paper.

Let Your Video Listen to Your Music! Contrastive Audio-Visual Masked Autoencoder

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T18:47:36.875662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:47:36.875662Z digest=sha256:38fdc2b3b78301b00dcf31b491b764586a64ab9e68925e220972b04812d77342

Observation b8a4800c-84d1-4295-8a4d-20dd0bc26f7b · inbound

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation cites this paper.

Hear-Your-Click: Interactive Object-Specific Video-to-Audio Generation Contrastive Audio-Visual Masked Autoencoder

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:39:53.588345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:39:53.588345Z digest=sha256:cd8813800e10f842ccc5391f869164265e2574aeb4af0433732e84dbd03db1d6

Observation 8696b57a-3e26-4303-8f96-eccab980b3f1 · inbound

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition cites this paper.

Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Contrastive Audio-Visual Masked Autoencoder

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:03:03.507305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:03:03.507305Z digest=sha256:89aee1a91cd3a8d66c1a97777663234012763efe48f02662ed3ab81c385e54e1

Observation abba1354-4e34-4433-88c0-93456f3e50af · inbound

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis cites this paper.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Contrastive Audio-Visual Masked Autoencoder

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.781335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.781335Z digest=sha256:8d63cbabf9467497b7bd91b55235d30cb8512cae11e19c530c33377dbcda81bf

Observation 90dd38b2-8ab7-4064-bc64-7caa43145675 · inbound

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos cites this paper.

Language-Guided Contrastive Audio-Visual Masked Autoencoder with Automatically Generated Audio-Visual-Text Triplets from Videos Contrastive Audio-Visual Masked Autoencoder

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:02:45.037290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:02:45.037290Z digest=sha256:a295d2d7df44ae7d34e8254ec6c4803518bb453bd14a0a6c59088fd888a17a7d

Observation 0b1f7c7a-b55a-4d1e-b348-b32c898633d5 · inbound

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation cites this paper.

Revisiting Audio-language Pretraining for Learning General-purpose Audio Representation Contrastive Audio-Visual Masked Autoencoder

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T21:06:09.947756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:06:09.947756Z digest=sha256:599d3432b0b1044cbbedc2f69a93b5c1bf110ee6580c2536e1288ab8e6638c0a

Observation 85d2e9d4-b713-46dd-8943-ff6ff105517f · inbound

Audio-Visual Continual Test-Time Adaptation without Forgetting cites this paper.

Audio-Visual Continual Test-Time Adaptation without Forgetting Contrastive Audio-Visual Masked Autoencoder

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T22:08:23.442235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:08:23.442235Z digest=sha256:46b205321085ef8dd678c3c66cffc90e27687d876d4a10b55927731fad6e85e9

Observation a8c91a3a-8934-4f4c-83ae-7ef929538a5f · inbound

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling cites this paper.

ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling Contrastive Audio-Visual Masked Autoencoder

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.072872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T09:22:10.483263Z digest=sha256:138267e79fc5c5d9780c0e094c479d310076feb7a335f4462eb2a078f9fb2b58

Observation ab8cf168-448e-4b94-97af-331555ecf050 · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning Contrastive Audio-Visual Masked Autoencoder

Reference 106

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T05:56:39.836610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:e8339c5d9a6d30d1f4f1b0d89f4e03c3c4632f5777d0d638cd46964efd194bda

Observation 0f64479d-064e-458f-b908-30950fa85215 · inbound

FATE: Frame-Level Audio-Visual Temporal Embedding cites this paper.

FATE: Frame-Level Audio-Visual Temporal Embedding Contrastive Audio-Visual Masked Autoencoder

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-15T15:15:06.491259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:15:06.491259Z digest=sha256:89ed3645ae17a67f9312735d1a5f86dad9a2091965cffe33444c8345909e82c0

Observation 9a3c7be1-1c5c-46b4-a022-e8d7256a9cfa · inbound

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion cites this paper.

HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion Contrastive Audio-Visual Masked Autoencoder

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T00:27:56.939307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:27:56.939307Z digest=sha256:96a7007bea394cdef144aa8e9385f71e54bfca7dbbbb7f3785094d0fa2ff59bb