Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:14:31.426091Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2412.20872.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:14:31.426091Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T01:07:11.142298Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-16T01:07:11.189974Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 027ffa93-b264-4561-9501-c1dbc7f1e54b · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Audio- visual event localization in unconstrained videos,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation cb88e4eb-532b-4cd9-bbb8-27a098845744 · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Dual attention matching for audio-visual event localization,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a80ea30b-4a9b-47e5-bc61-38e74334137d · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Learning to answer questions in dynamic audio-visual scenarios,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 291f353f-4670-4556-b5a5-2814a3299113 · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Listen to look: Action recognition by previewing audio,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 084c3cf5-50bc-4850-9bab-73e7b19f63e7 · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Unified multisensory per- ception: Weakly-supervised audio-visual video parsing,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 78b33c2e-ab8e-4b1a-9698-0c5370b29680 · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing MM-Pyramid: Multimodal pyramid attentional network for audio-visual event localization and video parsing,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 772b055e-a842-4dae-8880-c4f23906bcb5 · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Cross-modal Prompts: Adapting Large Pre-trained Models for Audio-Visual Downstream Tasks
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3678b30a-fd30-4b7a-8caf-757610b41371 · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Modality-independent teachers meet weakly-supervised audio-visual event parser
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b00ba2d6-0458-437d-baf4-b8376e7bfa00 · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Label-anticipated Event Disentanglement for Audio-Visual Video Parsing,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1f1ea22e-4445-4d3c-80e6-f73bb5246b79 · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Cbam: Convo- lutional block attention module,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 02f180ef-7605-4869-b5be-14b205afe4d5 · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Towards Efficient Audio-Visual Learners via Empowering Pre-trained Vi- sion Transformers with Cross-Modal Adaptation,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f874c14f-f6c7-419f-a57f-5c741bf31e3a · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Learning transferable visual models from natural language supervision,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c57e30cc-86a3-408f-84d0-533adb9a56ae · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f9db92f0-8937-4978-a6c2-e1ab5c9da90a · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Balanced multimodal learning via on-the-fly gradient modulation,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 86bc2f74-c6ed-4c96-823a-881dafcffd81 · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing What makes training multi-modal classification networks hard?
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d9fe477f-bdc7-40ed-a1da-392e5ae1f55b · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Scal- ing multimodal pre-training via cross-modality gradient harmonization,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c8b9b600-26d9-4085-b466-68c64a65c0dd · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Text-IF: Leveraging Semantic Text Guidance for Degradation- Aware and Interactive Image Fusion,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5dcd6f78-f15d-456f-ac84-61bcdc1c61ba · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Language- driven All-in-one Adverse Weather Removal,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 261f4f31-11f2-498b-ac28-e0c59cc642ac · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Multi-modal grouping network for weakly-supervised audio-visual video parsing,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c695cf29-0b0d-4c24-b307-03af97c209ed · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Joint-modal label denoising for weakly-supervised audio-visual video parsing,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d4577133-bff7-4684-a3a1-cdca17f11bec · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Collecting cross-modal presence-absence evidence for weakly-supervised audio- visual event perception,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation efe3f719-9e97-4b0f-867f-25133ee6c53b · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing CoLeaF: A Contrastive-Collaborative Learning Framework for Weakly Supervised Audio-Visual Video Parsing
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 854a2959-0109-4ff4-bb39-9fcbbd030c5a · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing CM-PIE: Cross-modal perception for interactive-enhanced audio-visual video parsing,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bd2c0d7f-8b4b-4b52-a42f-fa13dda274b7 · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing V ALOR: Vision- Audio-Language Omni-Perception Pretraining Model and Dataset,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d7cf971f-fddf-4cb3-84f1-1c7240fbf6a6 · outbound
LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing Multi-grained representa- tion learning for cross-modal retrieval,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5ece17e2-c5a1-4539-b62f-01df159f5a7a · inbound
TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing LINK: Adaptive Modality Interaction for Audio-Visual Video Parsing
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.