Pith. sign in

Paper Citation Record · LEDGER

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment

As of 16 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2505.01237.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.01237 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:26:32.991891Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ab5ba190-b74b-4850-853f-0225be2800ec · outbound

This paper cites Self-supervised learning of audio-visual objects from video.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Self-supervised learning of audio-visual objects from video

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:32.775790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:32.775790Z digest=sha256:39baea83c071511967c16d4fdcc01f0479b296d6fee33ca616a255c6d3f005d0

Observation 73b9163a-48bc-4536-91e9-ba49dd15249f · outbound

This paper cites Look, listen and learn.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Look, listen and learn

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.820348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.780800Z digest=sha256:342c5b0629a5f349d6caae6bfe62037bcb561a68142816c081e6652b1076145e

Observation 8721d22e-4783-44c8-99e1-4be0d12cd5fd · outbound

This paper cites Objects that sound.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Objects that sound

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.805349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.785314Z digest=sha256:49f8f3783c4f6f634d4faf83c9c22d53c504b2184da355b31692a2ecd2464d2f

Observation abb8d35f-e874-40ce-b211-b62f120a2b8e · outbound

This paper cites Sound- net: Learning sound representations from unlabeled video.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Sound- net: Learning sound representations from unlabeled video

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.790489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.789973Z digest=sha256:2ac74d2632f631ed3e7e0065729d6f1ef5634e0ecae26b7bdb9b8fbf0b9cfe83

Observation 59c34e70-af9a-468a-94aa-9aac795b5112 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Emerg- ing properties in self-supervised vision transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:32.794959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:32.794959Z digest=sha256:4f892a7c65d46218b007dfb84e5f526664ea22377b8e1932e26536d6c4e84041

Observation 9420bd3c-3674-43c4-8681-a06a2d76b89d · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Vggsound: A large-scale audio-visual dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.766400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.799505Z digest=sha256:7e17ffce1b21c3b12b8902a193b046202982c20b0e57f0797abe151f8cbf0c95

Observation f19c7b18-3466-41f9-b614-78e182dda387 · outbound

This paper cites Localiz- ing visual sounds the hard way.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Localiz- ing visual sounds the hard way

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.751761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.804317Z digest=sha256:8ced95f9fc9ea89df7af2070c4d1c6a6a2d7b196f49a57de72169a5f0a05319f

Observation aba398b6-c9ce-4d75-a004-151350fab2e0 · outbound

This paper cites Distilling audio-visual knowledge by com- positional contrastive learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Distilling audio-visual knowledge by com- positional contrastive learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.738412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.809033Z digest=sha256:990562bfea7b03b484406c6d40679361000bedd359473561b30f0f3dc90f10ee

Observation 389a656d-d290-4307-9cbf-a00ce46c72ff · outbound

This paper cites Look, listen, and attend: Co-attention network for self-supervised audio-visual representation learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Look, listen, and attend: Co-attention network for self-supervised audio-visual representation learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.724307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.813514Z digest=sha256:66d57d08e0c31d6d601703ed7b0b99db870ca484199b9f07773915862946a090

Observation cc1b77d1-1f77-4738-9521-0773c9fdb621 · outbound

This paper cites Vision transformers need registers.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Vision transformers need registers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.711043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.817897Z digest=sha256:8ddb0c1453647784cea415d9cdecddf79e1e23501879090ca4d982cd915ddbf0

Observation 29d9ea60-eb9f-47fc-998d-7f82934b04ed · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audio set: An ontology and human- labeled dataset for audio events

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.697065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.822446Z digest=sha256:13651b8d8dea9ac55beec719350ce8116333cd8883d5dfd9828863d1e695bc16

Observation 3cd7ed33-9754-4237-a3c0-202be15b27a7 · outbound

This paper cites Audiovisual masked autoencoders.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audiovisual masked autoencoders

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.683003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.826992Z digest=sha256:6bb81c7f606eeddbaceb6b72a841976868bb835df05d7457eb5bb25a66843b1d

Observation fa9b2a8a-99e5-4572-96a7-4ba46a39fa6e · outbound

This paper cites Imagebind: One embedding space to bind them all.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Imagebind: One embedding space to bind them all

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.669127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.831640Z digest=sha256:f07a94a57bb5ae98066215f638625722c8b6496bc305e7a8d3fdbea8cf62acac

Observation 6c0f4c33-db18-466c-aa5d-712e26f1a4f8 · outbound

This paper cites Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.654970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.835914Z digest=sha256:cd434d3ce64f56e34c9ccff4b171e9042164c01c5763dcbd80d9d469f80958eb

Observation 7003157a-d225-473d-adb4-ac16868a791a · outbound

This paper cites Cross- mae: Cross-modality masked autoencoders for region-aware audio-visual pre-training.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Cross- mae: Cross-modality masked autoencoders for region-aware audio-visual pre-training

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.639798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.840157Z digest=sha256:32960ea9972001ce4541e741d6c3dfcd1fc6aefc588661e45300a5b574034eb9

Observation 7485129c-84de-4caa-ae10-4b5e257db1a9 · outbound

This paper cites Separating the” chirp” from the” chat”: Self-supervised visual grounding of sound and language.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Separating the” chirp” from the” chat”: Self-supervised visual grounding of sound and language

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.626057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.844720Z digest=sha256:5c69c111c0db56a4db18bfb01af98dee36ff15371e53d64cd87bef2323cfe86b

Observation 1e006760-16ec-420d-bc6b-b7abd2a55357 · outbound

This paper cites Jointly dis- covering visual objects and spoken words from raw sensory input.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Jointly dis- covering visual objects and spoken words from raw sensory input

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.611946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.848959Z digest=sha256:71cf77a3db8a23fe9bd6b41c791e6dc2fb427c9bf4cf823341fdb3baea57fe4d

Observation 78ea52a8-0f76-417b-8399-76e4059b0a34 · outbound

This paper cites Multi- modal attention for fusion of audio and spatiotemporal fea- tures for video description.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Multi- modal attention for fusion of audio and spatiotemporal fea- tures for video description

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.596598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.853357Z digest=sha256:3385521a531038e7e6c8ce2ad889d18e9de2c8938468d83da084ca7f727690b6

Observation 30f027fb-4e56-4c69-b832-00867cb4bbb8 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.581567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.857551Z digest=sha256:42fced57ae80be1612c4a32ddf054fd2e06d9d4c71d14de469fdec828c92be50

Observation a20842a8-ea8c-4cea-b304-763c02564db4 · outbound

This paper cites Mavil: Masked audio-video learners.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Mavil: Masked audio-video learners

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.457498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.862296Z digest=sha256:bc2981900f38e6ce037189261058f805016d6e3d20a925fe6c2f8a206719dc5a

Observation 52136e88-5727-4d12-b8b6-e0e58f7aa1f8 · outbound

This paper cites EquiA V: Leveraging Equivari- ance for Audio-Visual Contrastive Learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment EquiA V: Leveraging Equivari- ance for Audio-Visual Contrastive Learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.442031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.866833Z digest=sha256:9d27539561e7a3833a36495a3ee97056817c162b006e2244b5eeb3e22e17901c

Observation a6e05e86-f8c8-4a25-98b3-0ccfdbc1768a · outbound

This paper cites Coopera- tive learning of audio and video models from self-supervised synchronization.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Coopera- tive learning of audio and video models from self-supervised synchronization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.426938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.871032Z digest=sha256:ee23744a06b5c038c38fb24288c034117dcbcb50192740779eea0624616e7edc

Observation 4a87a967-a4e1-4ec5-8d90-9d8ce76aca57 · outbound

This paper cites Cross-attentional audio-visual fusion for weakly- supervised action localization.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Cross-attentional audio-visual fusion for weakly- supervised action localization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:32.875357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:32.875357Z digest=sha256:a38b6e7734d1e1f172727b01159967fed4fecf5903601206e1ae217a256c1377

Observation 97cfe34f-ac02-4013-af24-f9410fbd6eeb · outbound

This paper cites Siamese vision transform- ers are scalable audio-visual learners.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Siamese vision transform- ers are scalable audio-visual learners

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.403313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.879753Z digest=sha256:92a66bde12ec8042e30fbb7058ddac9f5fffe0558a6c43eec8ed19b2afa62485

Observation 19cb27be-d1c7-44cd-bc3d-7b3d14033aa8 · outbound

This paper cites Vision transformers are parameter-efficient audio- visual learners.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Vision transformers are parameter-efficient audio- visual learners

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.389148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.883993Z digest=sha256:2ef11095c1084d7ebda3356845fd1b39b104240cc861935801185e3c7412c4a5

Observation 4aa27521-bd89-44d0-8dde-73aff2b7ec49 · outbound

This paper cites Active contrastive learning of audio-visual video representa- tions.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Active contrastive learning of audio-visual video representa- tions

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.374522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.888136Z digest=sha256:ffbe9095b875b8f60f49d93bc8e64203610530f3ae7754253ba995529c8d7e66

Observation c48d9fe1-9393-437e-a92b-2beeafada6ea · outbound

This paper cites Active contrastive learning of audio-visual video representa- tions.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Active contrastive learning of audio-visual video representa- tions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.359075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.892420Z digest=sha256:c594094cd9e43456b5eba0f1359f3a363b651e833ba33f3e6ae78929d27bc62c

Observation 261f2fbd-a995-4dd1-8c0e-d4548bedf6c8 · outbound

This paper cites Robust audio-visual instance discrimination.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Robust audio-visual instance discrimination

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.344760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.896858Z digest=sha256:af47b32ebb2cdcc230832f60dd27f5e3ca59dd8bf774146a0a497c6cedb5c7f3

Observation c4844cf3-ab58-4ec3-88bc-72a32a372b69 · outbound

This paper cites Audio- visual instance discrimination with cross-modal agreement.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audio- visual instance discrimination with cross-modal agreement

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.331778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.901142Z digest=sha256:bb88415831dace3e8261fc3952016998554c9df93a7bc8a8b2ca79524244db77

Observation ded765ae-f5e6-4a40-aecb-402998a79a28 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audio-visual scene analysis with self-supervised multisensory features

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.317636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.905131Z digest=sha256:45f7d4995604f2b7a6cbfeb1f29c9e8f5364fc19cf7713898bad644547c6bf46

Observation 7e8392aa-f18e-4acf-95a7-682aaddfa66f · outbound

This paper cites Ambient sound provides supervision for visual learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Ambient sound provides supervision for visual learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.303731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.909268Z digest=sha256:b9a1d9242b67089626009075360c262c8b82d077cfd293331636fb49ae013bed

Observation 1c192f66-1479-475c-a088-ee4d811c364f · outbound

This paper cites On compositions of transformations in contrastive self-supervised learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment On compositions of transformations in contrastive self-supervised learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.289207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.913150Z digest=sha256:a3c3d992c773b2ab48cd5899ea77115ede1055843542f7ebaeca9072db726e63

Observation 741a9754-e2dc-4f0f-8cf8-db0f6c36e4ed · outbound

This paper cites Broaden your views for self-supervised video learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Broaden your views for self-supervised video learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.274828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.916743Z digest=sha256:f56a480b60aaf34fbd70ebb610cb8a87de532965ffcdfe9a778260950de49ad5

Observation 6a6e6f64-3bc0-4896-80db-b4cb6659f2fc · outbound

This paper cites Avlnet: Learning audio-visual language representa- tions from instructional videos.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Avlnet: Learning audio-visual language representa- tions from instructional videos

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.260253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.920815Z digest=sha256:864dadcb6e92b820685ba35acd3e58c5fad3742975df65c49fb867a6c0430ab6

Observation edc60b02-4d74-4097-b79c-e28efaf0a2bf · outbound

This paper cites Self-supervised audio- visual representation learning with relaxed cross-modal syn- chronicity.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Self-supervised audio- visual representation learning with relaxed cross-modal syn- chronicity

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.245489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.925091Z digest=sha256:09097eeb7007ebcc75ad043c85976effbfb6f17108a1cda56cb72b0c9bdb1db3

Observation 6ff6cfa9-1b25-4b6d-bf08-532115f77f59 · outbound

This paper cites Event-specific audio-visual fusion layers: A simple and new perspective on video understanding.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Event-specific audio-visual fusion layers: A simple and new perspective on video understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.230436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.929279Z digest=sha256:dbbed7a956aaceddcc791720633b41a356b885d1b2b0d7a652b9667c247a6e2e

Observation e4892cd6-6be3-40fc-ad05-52d2911f5a50 · outbound

This paper cites From vision to au- dio and beyond: A unified model for audio-visual representa- tion and generation.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment From vision to au- dio and beyond: A unified model for audio-visual representa- tion and generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.214522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.933357Z digest=sha256:5caf362cb435627bf6b6e884ffe6c7c94d5d2ad18f7915edb975ba4aa72d8c35

Observation 047fe13b-cd2f-4c29-82c8-3751fc36d5ed · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Learning audio-visual source localization via false negative aware contrastive learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.199227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.937457Z digest=sha256:213438d3a8221bd7eae6d06a6c13bdd5413fea60fbc89373e4642bbc97474441

Observation fc5418b9-2d54-4061-b30e-9488c3298d4a · outbound

This paper cites Multimodal Self-Supervised Learning of General Audio Representations.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Multimodal Self-Supervised Learning of General Audio Representations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:32.941470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:32.941470Z digest=sha256:50be37528e513c12d6db2cc80bac0af5b7dd16ec0c34d98154ac13e7fcec83c7

Observation 89288667-2be6-4606-8168-8889985cf1d2 · outbound

This paper cites Temporal cue guided video highlight detection with low-rank audio-visual fusion.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Temporal cue guided video highlight detection with low-rank audio-visual fusion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.184876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.946526Z digest=sha256:51c96fd375098e75ca8f00939aa007f7d5e210a1f08075cef9dc708e3f45a929

Observation 8620266b-ec0c-4d5e-9ab1-4257529d94c2 · outbound

This paper cites Con- trastive learning of global and local video representations.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Con- trastive learning of global and local video representations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.171128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.950818Z digest=sha256:77bb04452223768cf627facf6706d3cfabf41d680c92dc76cb91ffdae0c99037

Observation a0311b32-5c5d-4fc8-abad-76a0f5d08865 · outbound

This paper cites The sound of pixels.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment The sound of pixels

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.156494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.955193Z digest=sha256:64a9d1257ff68068d1ad097177b40e83882e4be91360333e51eebeb7d911dc42

Observation aa27f171-004a-4ceb-90f0-7c343d55ca27 · outbound

This paper cites Scene parsing through ade20k dataset.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Scene parsing through ade20k dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:32.959437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:32.959437Z digest=sha256:adac61a43aba40fece85f6dd3a0b36cd6ffb996b634f03a9859ebaf69830d307

Observation 56ae1a59-bbaf-4163-b650-9a12be9fcba5 · outbound

This paper cites Languagebind: Extending video-language pretraining to n- modality by language-based semantic alignment.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Languagebind: Extending video-language pretraining to n- modality by language-based semantic alignment

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.133027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.964120Z digest=sha256:6d0b9182055ba79fc574978de00349d9c9623784e26d92aea2425e97a0a01964

Observation 12498e67-28ca-4db0-b812-1f954bcbcfff · outbound

This paper cites an unresolved cited work.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:26:33.119294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.968662Z digest=sha256:ad6bd133a31140a2554c036c63e986233c039d562780f70e02f2a024eb0c9164

Observation 9bb38671-8523-400e-a3c8-de9b7e29c955 · outbound

This paper cites an unresolved cited work.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:26:33.104530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.973334Z digest=sha256:ac7a8a55ba61359fcec722276b227d3fbb7b349513a4c08ecffa7374b133de0c

Observation da8d6112-834b-4bc6-9fe2-b732aade6310 · outbound

This paper cites di- agonal mean.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment di- agonal mean

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.089873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.977885Z digest=sha256:584cc79f723bf7d38ee811286a03d1cc1659ec56b994436b02ffa0ebf542c1bb

Observation dbef4d5f-5e60-45ce-8820-907cc749079f · outbound

This paper cites Table 12 shows the performance comparison between register to- kens, patch tokens, and the global token.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Table 12 shows the performance comparison between register to- kens, patch tokens, and the global token

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.075447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.982956Z digest=sha256:eff3860a46684c1c36524b12c4f496f5232fe56e40f84efd0e60d60e1eb4d2bd

Observation e08250c1-bba3-4508-8115-66e6321d4225 · outbound

This paper cites writing on blackboard with chalk.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment writing on blackboard with chalk

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.060343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.987325Z digest=sha256:f640d9058e4a441ae5fec4c9a02e50ef9aeaf0a5c7249f7a2a8584019699a6d3

Observation 521aa4e2-2093-47a7-aa0d-aa4e7c0f997c · outbound

This paper cites For this experiment, we manually annotate the occurrence of the classes throughout the video.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment For this experiment, we manually annotate the occurrence of the classes throughout the video

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.045181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.991891Z digest=sha256:f6050e796e1733660872cf1463d3f1bd5e57406d3d59a6c31e7273575f51782c

Pith citing papers

No inbound Pith citation observations are available.