Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:58:06.771420Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 3 inbound Pith citation observations for arXiv:2507.01384.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:58:06.771420Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T08:40:53.734483Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T03:06:19.435608Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation f1524d88-db8c-4eec-8bb6-9e54b618e48e · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing YOLOv4: Optimal Speed and Accuracy of Object Detection
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19f18f19-6993-46cf-b4d7-94b87ebe26a6 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Cm-pie: Cross-modal perception for interactive-enhanced audio-visual video pars- ing
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cf45796f-a654-4725-807c-735e07bcc76c · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Joint-modal label denois- ing for weakly-supervised audio-visual video parsing
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 560c4621-b056-414b-a676-215b7a3ed103 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Autoaugment: Learning augmentation strategies from data
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2c8e5958-95e2-4ba1-bbc2-991294eacc14 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Randaugment: Practical automated data augmen- tation with a reduced search space
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 452e0cba-d5f6-4472-bb85-350eb5a0da76 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce5ceb2c-df9c-4c00-9e7b-463252d0f3dd · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aeec829-a26a-4c11-8b4a-a7640cf62ed5 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Cross-modal prompts: Adapting large pre- trained models for audio-visual downstream tasks
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 226553c4-125e-47df-b314-b453fcafac38 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Revisit weakly-supervised audio-visual video parsing from the lan- guage perspective
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73015c3d-b739-4224-aefc-aaec94e1dfbb · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Col- lecting cross-modal presence-absence evidence for weakly- supervised audio-visual event perception
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ced513ac-8747-43e7-a380-f0e0a33944b9 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a739b27-4492-43ed-a0f6-645d7774e92b · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Explaining and Harnessing Adversarial Examples
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 151aea69-dc35-4d0a-a214-421f773f98c9 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bec3576d-f24e-49b7-bed3-54c3068c74da · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Efficiently Modeling Long Sequences with Structured State Spaces
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9d2b1b8-0d15-4049-857e-bb6fea2968ad · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Combining recurrent, convolutional, and continuous-time models with linear state space layers
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b33f58f-cfac-4ef9-93c0-2674aa8c290d · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Liquid Structural State-Space Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2e70f55-3791-4d6a-afba-899058713594 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing MambaVision: A Hybrid Mamba-Transformer Vision Backbone
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2dc4534-61f6-4fe7-bdc7-f9fd3e26e966 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Deep residual learning for image recognition
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27fe4404-ecc1-4c00-ae4b-1174922e344c · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89cb7c1d-36e9-4b82-abe0-ebb6b891b01d · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Cnn archi- tectures for large-scale audio classification
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e53fcd3-3a9c-4f42-96a6-e519130c7bcc · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing LocalMamba: Visual State Space Model with Windowed Selective Scan
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9301dcee-7cf8-4901-b4b6-acc4378b16f5 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Learning tem- porally invariant and localizable features via data augmenta- tion for video recognition
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 48ceaf10-62c6-4f91-9ae6-f2f7892359f4 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Modality-independent teachers meet weakly-supervised audio-visual event parser
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0cc45d7e-b710-426a-8604-a17bb499f518 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Coupled Mamba: Enhanced Multi-modal Fusion with Coupled State Space Model
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e794d33-65de-4e8c-96bf-bd59dadc2849 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Jamba: A Hybrid Transformer-Mamba Language Model
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef5608eb-864b-42fd-8792-7251e2f8cf24 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Vmamba: Visual state space model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65adcd42-ccc8-4f29-a9ab-0883e487c2f5 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Swin transformer: Hierarchical vision transformer using shifted windows
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6c89b0a-ea2d-470f-b872-c34fa8f808a7 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 38faa884-73fb-4956-a11e-51faa5c0e065 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Multi-modal grouping network for weakly-supervised audio-visual video parsing
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c56607c-b636-4ffe-b403-bee1d4f7e3a3 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Learning transferable visual models from natural language supervi- sion
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 655afca1-a5dc-4906-a1e6-e8d47a33df33 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Coleaf: A contrastive-collaborative learning framework for weakly supervised audio-visual video pars- ing
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 655ab876-e5e7-409e-ad01-f5f0d5185dc8 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Simplified State Space Layers for Sequence Modeling
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94fc1c35-2f04-4555-a62d-7c0d41389f16 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Audio-visual event localization in unconstrained videos
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4ee575d-c746-4e32-8eb6-9f0ad5b9ba5c · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 562c4318-668e-41cf-bfb4-410212aa1dc0 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Link: Adaptive modality interaction for audio-visual video parsing
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0605fbe7-85df-4b97-94d1-f91383469d89 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Cbam: Convolutional block attention module
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 73efefe9-3b44-43bc-b36b-68bdd364fe6b · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Exploring heterogeneous clues for weakly-supervised audio-visual video parsing
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9882aed9-e2eb-40e2-b052-e07aff7e68c0 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Dual attention matching for audio-visual event localization
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1d827604-f9d1-4ab5-a83f-76e170833c64 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ed0db64-e4dd-4aca-90dd-7496901d27f4 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing FusionMamba: Dynamic Feature Enhancement for Multimodal Image Fusion with Mamba
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d365f679-fdf1-4cea-b324-81e40585f0ab · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Language- driven all-in-one adverse weather removal
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60fd1cc5-0115-4de8-a785-3cadecebb228 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing DepMamba: Progressive Fusion Mamba for Multimodal Depression Detection
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5fd2380-7cf8-4973-af4e-a6605d2aa624 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Text-if: Leveraging semantic text guidance for degradation-aware and interactive image fusion
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d7ce16a-c48d-4525-aaaf-f072feda7598 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Mm-pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8b6998f6-0317-4d24-b6c9-93447fb321da · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing MambaOut: Do We Really Need Mamba for Vision?
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0a227d4-5167-4fd2-b780-5e17d58baadf · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Cutmix: Regu- larization strategy to train strong classifiers with localizable features
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9786d213-3270-4def-ba54-3f9774615a9f · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing VideoMix: Rethinking Data Augmentation for Video Classification
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ede876a5-9423-4eb5-b7a3-9d0cc4aba866 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Don't Judge by the Look: Towards Motion Coherent Video Representation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0ed3644-c223-4aed-be6c-570b4f284f05 · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Label-anticipated Event Disentanglement for Audio-Visual Video Parsing
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e9efd3a-ee14-4b9d-96a0-8c80449c38de · outbound
MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbc685ee-a8fd-4b6e-bbc9-03e16cefc264 · inbound
EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23fc593e-51da-478f-b991-062c5c2beefc · inbound
UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20d44590-0dd1-45c2-81d0-6b1c111972b9 · inbound
UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.