Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:03:03.579463Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2507.09323.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T18:03:03.579463Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 81fd07c6-b149-4500-99ae-b80d39c05c51 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Vatt: Transformers for multimodal self-supervised learning from raw video, audio and text
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cefb466b-5949-46af-bcea-ebb65e9331cc · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b0797321-987a-43db-877a-5e0851227593 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Multimodal machine learning: A survey and tax- onomy
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8ee51abf-9a5d-485d-9cd4-b7c4d669e9d0 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Vggsound: A large-scale audio-visual dataset
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7537496e-dbc4-4ce6-b0a8-25fb935740eb · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition A simple framework for contrastive learning of visual representations
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87a38546-9c25-4da7-b423-bcd7692e4e77 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Scaling egocentric vision: The epic-kitchens dataset
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08276ac1-58f2-4dcf-be21-fe5fceb7038d · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition An image is worth 16x16 words: Transformers for image recognition at scale
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cda91782-4f55-477d-a412-4ffe5229e882 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Large Scale Audiovisual Learning of Sounds with Weakly Labeled Data
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation db469e9c-eab5-497f-84ec-e9d938026c04 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Audio set: An ontology and human- labeled dataset for audio events
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13d93ebc-8506-4e7e-ac6c-8d0f6ba74705 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Audiovisual masked autoencoders
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8eeb24fa-c3e4-4bf2-bce0-6e3ca5a37666 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition AST: Audio Spectrogram Transformer
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75677a28-1a87-4e6a-8b4c-a88b97687b0e · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Uavm: Towards unifying audio and visual models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8696b57a-3e26-4303-8f96-eccab980b3f1 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Contrastive Audio-Visual Masked Autoencoder
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43fe541e-1486-4185-b0fb-1a857cf9e9d8 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Deep residual learning for image recognition
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34a4a523-6847-4011-a690-2d8c5743b512 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Hubert: Self-supervised speech representation learning by masked prediction of hidden units
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 66f5017d-ab57-4e8a-9d3d-ba3f78cce6a2 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Perceiver io: A general architecture for structured inputs & outputs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9fd1149d-b653-4743-b5b1-1c9cda6cd8de · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Supervised contrastive learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3dfc1e15-bec4-4adf-9842-a7c36837af65 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Kipf and Max Welling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 91f8c0f3-0761-4cd8-88b6-c75018c4447e · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Swin transformer: Hierarchical vision transformer using shifted windows
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8884ea43-29e0-40e3-97c3-353b84b113d4 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Attention bottlenecks for multimodal fusion
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8e499b7-5b22-433e-a139-d600351e428f · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition OmniNet: A unified architecture for multi-modal multi-task learning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cbb55465-42fc-4aad-af02-d8caa9f7e5b2 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Robust speech recognition via large-scale weak supervision
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 45308ca1-9e25-49cf-8ec7-039260699ba7 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Multimodal fusion for audio-image and video action recognition
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 93092bbe-2d4a-4d8b-9715-0b47c7270d60 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69247b40-9d48-4fe2-82f4-5a69517c00d3 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Omnivec: Learn- ing robust representations with cross modal sharing
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43b5864b-2380-435f-9838-5310b5699aad · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition One-peace: Ex- ploring one general representation model toward unlimited modalities
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7db7b7c4-35da-46b2-9112-372c202ae259 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition What makes train- ing multi-modal classification networks hard? In Proceed- ings of the IEEE/CVF conference on computer vision and pattern recognition, pages 12695–12705, 2020
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1ee6d336-c67d-4573-95c9-601d530a2871 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Multi-stream multi-class fusion of deep net- works for video classification
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 430b57eb-fd29-4b5e-a891-bc128515017a · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Multimodal learning with transformers: A survey
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 386e133a-4d35-425b-b3ee-13d6a532b2b8 · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Peters, and Yejin Choi
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 43cce9bc-5253-4dba-a12c-ba2f1337c6fd · outbound
Dynamic Inter-Class Confusion-Aware Encoder for Audio-Visual Fusion in Human Activity Recognition Multimodal representation learning: Advances, trends and challenges
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
No inbound Pith citation observations are available.