Pith. sign in

Paper Citation Record · LEDGER

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment

As of 16 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 0 inbound Pith citation observations for arXiv:2505.01237.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.01237 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T04:26:32.991891Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

50 of 50 outbound references displayed

  • verified exact0
  • verified fuzzy43
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ab5ba190-b74b-4850-853f-0225be2800ec · outbound

This paper cites Self-supervised learning of audio-visual objects from video.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Self-supervised learning of audio-visual objects from video

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:32.775790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:32.775790Z digest=sha256:2e93b7440190ca721094b81679ff5f04205b351d2a6ac124cd6e718d363f208a

Observation 73b9163a-48bc-4536-91e9-ba49dd15249f · outbound

This paper cites Look, listen and learn.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Look, listen and learn

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.820348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.780800Z digest=sha256:1ca07fbd0ee536f1b1d031bc54e5d8090a9ffb67e70fc9772ca5d9a0f5a00fb2

Observation 8721d22e-4783-44c8-99e1-4be0d12cd5fd · outbound

This paper cites Objects that sound.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Objects that sound

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.805349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.785314Z digest=sha256:4feeb88e6f899dbeca340ce30a4c388e7a19156a0d1bbe3c8ad1d998cff2415b

Observation abb8d35f-e874-40ce-b211-b62f120a2b8e · outbound

This paper cites Sound- net: Learning sound representations from unlabeled video.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Sound- net: Learning sound representations from unlabeled video

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.790489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.789973Z digest=sha256:4499745e763f5102d18aacb4a0b64d162104a6e5001db396eaa8e54a71309b61

Observation 59c34e70-af9a-468a-94aa-9aac795b5112 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Emerg- ing properties in self-supervised vision transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:32.794959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:32.794959Z digest=sha256:91ecb8decc3dc95dcc41ae142c962fc5450ee6b7c7296feb65354a6891b46027

Observation 9420bd3c-3674-43c4-8681-a06a2d76b89d · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Vggsound: A large-scale audio-visual dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.766400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.799505Z digest=sha256:6510fc63622948f29b08061301e0f11dd6d2062da72202b9c2808719ff5b7d2d

Observation f19c7b18-3466-41f9-b614-78e182dda387 · outbound

This paper cites Localiz- ing visual sounds the hard way.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Localiz- ing visual sounds the hard way

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.751761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.804317Z digest=sha256:31cb9640c6583c9f340d82506d2e280b8296270bec70504f9bbc0349c2d0cee8

Observation aba398b6-c9ce-4d75-a004-151350fab2e0 · outbound

This paper cites Distilling audio-visual knowledge by com- positional contrastive learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Distilling audio-visual knowledge by com- positional contrastive learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.738412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.809033Z digest=sha256:68b5beaa78875c58786bfff464e23c17bb0f81ba652d201491ca75444bd2d096

Observation 389a656d-d290-4307-9cbf-a00ce46c72ff · outbound

This paper cites Look, listen, and attend: Co-attention network for self-supervised audio-visual representation learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Look, listen, and attend: Co-attention network for self-supervised audio-visual representation learning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.724307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.813514Z digest=sha256:eb25ff644f83b7bb3e32754f2c4e455fb8b755a96faa679d31f77185f1f7d6e1

Observation cc1b77d1-1f77-4738-9521-0773c9fdb621 · outbound

This paper cites Vision transformers need registers.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Vision transformers need registers

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.711043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.817897Z digest=sha256:0c3fe56f2bc8931130781c1a80d9f7e3f7cd1f601927e4926ff8611a70afc268

Observation 29d9ea60-eb9f-47fc-998d-7f82934b04ed · outbound

This paper cites Audio set: An ontology and human- labeled dataset for audio events.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audio set: An ontology and human- labeled dataset for audio events

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.697065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.822446Z digest=sha256:08d345e4fc669dd46865e6b1a5a303d1514eb9e4252224a593816d1a998c2c4e

Observation 3cd7ed33-9754-4237-a3c0-202be15b27a7 · outbound

This paper cites Audiovisual masked autoencoders.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audiovisual masked autoencoders

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.683003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.826992Z digest=sha256:e5698dd8fe0f7fa1fc33f76de1ec0daa7421d01ea05788b878138f035198f60a

Observation fa9b2a8a-99e5-4572-96a7-4ba46a39fa6e · outbound

This paper cites Imagebind: One embedding space to bind them all.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Imagebind: One embedding space to bind them all

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.669127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.831640Z digest=sha256:a3ef633483e875b878be2d4a94ab4c3400ed070caff39e81cf35c43501399c5e

Observation 6c0f4c33-db18-466c-aa5d-712e26f1a4f8 · outbound

This paper cites Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Liu, David Harwath, Leonid Karlinsky, Hilde Kuehne, and James R

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.654970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.835914Z digest=sha256:ebdbcb4174aa800bfece5d1498db99ef2b2d55c29cc72ba487580110aa9deefc

Observation 7003157a-d225-473d-adb4-ac16868a791a · outbound

This paper cites Cross- mae: Cross-modality masked autoencoders for region-aware audio-visual pre-training.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Cross- mae: Cross-modality masked autoencoders for region-aware audio-visual pre-training

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.639798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.840157Z digest=sha256:fa4d6adcf65b7ef3d03ed0b7be534c0c12250cc135a0076e3a0027fa002343b2

Observation 7485129c-84de-4caa-ae10-4b5e257db1a9 · outbound

This paper cites Separating the” chirp” from the” chat”: Self-supervised visual grounding of sound and language.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Separating the” chirp” from the” chat”: Self-supervised visual grounding of sound and language

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.626057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.844720Z digest=sha256:0c695ec933fcb3cb82d6f864ef09469ff65c2adde5ae0dd340d9ec608b1088a8

Observation 1e006760-16ec-420d-bc6b-b7abd2a55357 · outbound

This paper cites Jointly dis- covering visual objects and spoken words from raw sensory input.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Jointly dis- covering visual objects and spoken words from raw sensory input

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.611946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.848959Z digest=sha256:0287bae4d0eb242dfa97d580f43a915b93368b493a2250b87d6362b28c66555f

Observation 78ea52a8-0f76-417b-8399-76e4059b0a34 · outbound

This paper cites Multi- modal attention for fusion of audio and spatiotemporal fea- tures for video description.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Multi- modal attention for fusion of audio and spatiotemporal fea- tures for video description

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.596598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.853357Z digest=sha256:b355b51ce977ebb65ae6217ba7b67d509e8d472b20489c1a5465a5eacf99aa5a

Observation 30f027fb-4e56-4c69-b832-00867cb4bbb8 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.581567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.857551Z digest=sha256:23eb4377c909fb4148327b416219e9d5681223b98475a2711d7e602d50b2a17f

Observation a20842a8-ea8c-4cea-b304-763c02564db4 · outbound

This paper cites Mavil: Masked audio-video learners.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Mavil: Masked audio-video learners

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.457498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.862296Z digest=sha256:e2e77f9ca235e228c3d4dacf08d3f11f22405df8af5202a73fcc3a3fb9671910

Observation 52136e88-5727-4d12-b8b6-e0e58f7aa1f8 · outbound

This paper cites EquiA V: Leveraging Equivari- ance for Audio-Visual Contrastive Learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment EquiA V: Leveraging Equivari- ance for Audio-Visual Contrastive Learning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.442031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.866833Z digest=sha256:59b7da48cfca133a529105a7ff9f0b082104a1208e58093f8f53516b417968f3

Observation a6e05e86-f8c8-4a25-98b3-0ccfdbc1768a · outbound

This paper cites Coopera- tive learning of audio and video models from self-supervised synchronization.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Coopera- tive learning of audio and video models from self-supervised synchronization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.426938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.871032Z digest=sha256:211f43a48a2fcc723501b84d7e6859a0befbe0bb6f3921a7c5fb7d40aa5e67f8

Observation 4a87a967-a4e1-4ec5-8d90-9d8ce76aca57 · outbound

This paper cites Cross-attentional audio-visual fusion for weakly- supervised action localization.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Cross-attentional audio-visual fusion for weakly- supervised action localization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:32.875357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:32.875357Z digest=sha256:9db7438397470ff555aa2b82087ddff262fc0bec29d19216bc33d95612721876

Observation 97cfe34f-ac02-4013-af24-f9410fbd6eeb · outbound

This paper cites Siamese vision transform- ers are scalable audio-visual learners.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Siamese vision transform- ers are scalable audio-visual learners

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.403313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.879753Z digest=sha256:df269bb66ed33b98af31ad654731e17f94c37f1b0f554df9e26e69f4b4b71416

Observation 19cb27be-d1c7-44cd-bc3d-7b3d14033aa8 · outbound

This paper cites Vision transformers are parameter-efficient audio- visual learners.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Vision transformers are parameter-efficient audio- visual learners

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.389148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.883993Z digest=sha256:4b21e59ac587d13eee5bf7e3d3da30a78be6f8cc1e880ef8f2af23a81711a287

Observation 4aa27521-bd89-44d0-8dde-73aff2b7ec49 · outbound

This paper cites Active contrastive learning of audio-visual video representa- tions.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Active contrastive learning of audio-visual video representa- tions

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.374522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.888136Z digest=sha256:aea726f2c044ad670543fdaa3cc3b4265d4080c593a9d7edb2145a727553429a

Observation c48d9fe1-9393-437e-a92b-2beeafada6ea · outbound

This paper cites Active contrastive learning of audio-visual video representa- tions.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Active contrastive learning of audio-visual video representa- tions

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.359075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.892420Z digest=sha256:b2b63b8a1b3da16cd7f93276ca9fd8155b434895bd97662d2ff50ac695fc4c44

Observation 261f2fbd-a995-4dd1-8c0e-d4548bedf6c8 · outbound

This paper cites Robust audio-visual instance discrimination.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Robust audio-visual instance discrimination

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.344760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.896858Z digest=sha256:6218978b41cd1b77283b93746e78552a2914fad25381daf0b2b9824d9c3f140b

Observation c4844cf3-ab58-4ec3-88bc-72a32a372b69 · outbound

This paper cites Audio- visual instance discrimination with cross-modal agreement.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audio- visual instance discrimination with cross-modal agreement

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.331778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.901142Z digest=sha256:619e95831a8f0d38cf228b3c2a884a298c86b700e2ec64b5929d4b3a5feb8594

Observation ded765ae-f5e6-4a40-aecb-402998a79a28 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Audio-visual scene analysis with self-supervised multisensory features

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.317636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.905131Z digest=sha256:944faab01577f70af9de01154084f98faf7800243d285c252df4cdf223b89ca6

Observation 7e8392aa-f18e-4acf-95a7-682aaddfa66f · outbound

This paper cites Ambient sound provides supervision for visual learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Ambient sound provides supervision for visual learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.303731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.909268Z digest=sha256:dd7e8158e248c58216c8e7901305e63e6f15e683fc082f636610460fc4520951

Observation 1c192f66-1479-475c-a088-ee4d811c364f · outbound

This paper cites On compositions of transformations in contrastive self-supervised learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment On compositions of transformations in contrastive self-supervised learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.289207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.913150Z digest=sha256:c6459d918331b161c8073e748a258b2080639570294d56707524e9822624c699

Observation 741a9754-e2dc-4f0f-8cf8-db0f6c36e4ed · outbound

This paper cites Broaden your views for self-supervised video learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Broaden your views for self-supervised video learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.274828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.916743Z digest=sha256:874bf6d14fc9ab3fbacd71b48499083314143918def17623f8823d9aa6a7886f

Observation 6a6e6f64-3bc0-4896-80db-b4cb6659f2fc · outbound

This paper cites Avlnet: Learning audio-visual language representa- tions from instructional videos.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Avlnet: Learning audio-visual language representa- tions from instructional videos

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.260253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.920815Z digest=sha256:3c67d741b34eb172d42e2b9d8d0c75c7534180b201db24677284b11e7d587fb0

Observation edc60b02-4d74-4097-b79c-e28efaf0a2bf · outbound

This paper cites Self-supervised audio- visual representation learning with relaxed cross-modal syn- chronicity.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Self-supervised audio- visual representation learning with relaxed cross-modal syn- chronicity

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.245489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.925091Z digest=sha256:698392bf976990c6172d2c84c6d096217992de81146bb81a13a985ee1429ec7b

Observation 6ff6cfa9-1b25-4b6d-bf08-532115f77f59 · outbound

This paper cites Event-specific audio-visual fusion layers: A simple and new perspective on video understanding.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Event-specific audio-visual fusion layers: A simple and new perspective on video understanding

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.230436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.929279Z digest=sha256:e592ccf5c023e35ab151bf80ab82062c9820b20aa89ed8c8b851df7097b685df

Observation e4892cd6-6be3-40fc-ad05-52d2911f5a50 · outbound

This paper cites From vision to au- dio and beyond: A unified model for audio-visual representa- tion and generation.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment From vision to au- dio and beyond: A unified model for audio-visual representa- tion and generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.214522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.933357Z digest=sha256:1ab7cb431ae1d9c9bcc819cf244ce0fa638852b4d45c9fb48778350cd027b407

Observation 047fe13b-cd2f-4c29-82c8-3751fc36d5ed · outbound

This paper cites Learning audio-visual source localization via false negative aware contrastive learning.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Learning audio-visual source localization via false negative aware contrastive learning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.199227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.937457Z digest=sha256:395d036c6a3f4e545f428b0955e63a024f710c69cb3a44c30304f3ae40d944e0

Observation fc5418b9-2d54-4061-b30e-9488c3298d4a · outbound

This paper cites Multimodal Self-Supervised Learning of General Audio Representations.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Multimodal Self-Supervised Learning of General Audio Representations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:32.941470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:32.941470Z digest=sha256:3e3006df82b0d1ed2de240dca8a49ab3e6a1a4b169563eacc41c57a320cd26db

Observation 89288667-2be6-4606-8168-8889985cf1d2 · outbound

This paper cites Temporal cue guided video highlight detection with low-rank audio-visual fusion.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Temporal cue guided video highlight detection with low-rank audio-visual fusion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.184876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.946526Z digest=sha256:c697beb5be59ac7aa618339fea8c41285799613c826c87dfbbe853e5279861d9

Observation 8620266b-ec0c-4d5e-9ab1-4257529d94c2 · outbound

This paper cites Con- trastive learning of global and local video representations.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Con- trastive learning of global and local video representations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.171128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.950818Z digest=sha256:3a1dc36563adf593a4dcfa95847a55f9fe5fc9eee306b0ddc550038952da49d5

Observation a0311b32-5c5d-4fc8-abad-76a0f5d08865 · outbound

This paper cites The sound of pixels.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment The sound of pixels

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.156494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.955193Z digest=sha256:c7b303ea75ee7b19d064036deca2dffc3cdd5aadf97d278c8a84e3282b9fa929

Observation aa27f171-004a-4ceb-90f0-7c343d55ca27 · outbound

This paper cites Scene parsing through ade20k dataset.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Scene parsing through ade20k dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T04:26:32.959437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:26:32.959437Z digest=sha256:50ee4e45c09fa5d91ffb8567d71701ae5fa4a4d6be74a01291c60673e23f0935

Observation 56ae1a59-bbaf-4163-b650-9a12be9fcba5 · outbound

This paper cites Languagebind: Extending video-language pretraining to n- modality by language-based semantic alignment.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Languagebind: Extending video-language pretraining to n- modality by language-based semantic alignment

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.133027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.964120Z digest=sha256:d5f8e15aadd805141659e7f6dcc55c26d5763bd35aff5f5727aa179984ef563b

Observation 12498e67-28ca-4db0-b812-1f954bcbcfff · outbound

This paper cites an unresolved cited work.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:26:33.119294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.968662Z digest=sha256:ffb2755bffd9334dec1ac9909834b2b7ca47d41a8eb9911f95644c105d4b86c2

Observation 9bb38671-8523-400e-a3c8-de9b7e29c955 · outbound

This paper cites an unresolved cited work.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-16T04:26:33.104530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.973334Z digest=sha256:e8938a9681b21d41e069d8ee44e5d7664f3d47b8d6a343ade8e245c733c87a0c

Observation da8d6112-834b-4bc6-9fe2-b732aade6310 · outbound

This paper cites di- agonal mean.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment di- agonal mean

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.089873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.977885Z digest=sha256:473032a807be44dc31915deefeb1b4521dc455549ee6a0ffeab8ca47bc314077

Observation dbef4d5f-5e60-45ce-8820-907cc749079f · outbound

This paper cites Table 12 shows the performance comparison between register to- kens, patch tokens, and the global token.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment Table 12 shows the performance comparison between register to- kens, patch tokens, and the global token

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.075447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.982956Z digest=sha256:0039c4a23e1155d9f7e0d5a1a284e2f4876920ee0424dc473af3cfcf0c62c0d3

Observation e08250c1-bba3-4508-8115-66e6321d4225 · outbound

This paper cites writing on blackboard with chalk.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment writing on blackboard with chalk

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.060343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.987325Z digest=sha256:d1f172d81f557ba3b1cdf1a5ac3db2265299d454a7efa9e992b88ccaa01cfe03

Observation 521aa4e2-2093-47a7-aa0d-aa4e7c0f997c · outbound

This paper cites For this experiment, we manually annotate the occurrence of the classes throughout the video.

CAV-MAE Sync: Improving Contrastive Audio-Visual Mask Autoencoders via Fine-Grained Alignment For this experiment, we manually annotate the occurrence of the classes throughout the video

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T04:26:33.045181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-16T04:26:32.991891Z digest=sha256:6ba9000a36c4ff65a0a2bffaa89a2917be0f1a58212abbfd7d5dd7f4efa696c2

Pith citing papers

No inbound Pith citation observations are available.