Pith. sign in

Paper Citation Record · LEDGER

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing

As of 7 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 3 inbound Pith citation observations for arXiv:2507.01384.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.01384 v2

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:58:06.771420Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T08:40:53.734483Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T03:06:19.435608Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact4
  • verified fuzzy21
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f1524d88-db8c-4eec-8bb6-9e54b618e48e · outbound

This paper cites YOLOv4: Optimal Speed and Accuracy of Object Detection.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing YOLOv4: Optimal Speed and Accuracy of Object Detection

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:01.906010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:01.906010Z digest=sha256:005e381a08c2e53d6ff216dea1b0e43803e90df3870ace4e819800694f9d4a5b

Observation 19f18f19-6993-46cf-b4d7-94b87ebe26a6 · outbound

This paper cites Cm-pie: Cross-modal perception for interactive-enhanced audio-visual video pars- ing.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Cm-pie: Cross-modal perception for interactive-enhanced audio-visual video pars- ing

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:11.863548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:02.010610Z digest=sha256:ad88cfa29a9164c7f99b14d37cd22e6503a3dc0376fbe0d221fae62be9cc0723

Observation cf45796f-a654-4725-807c-735e07bcc76c · outbound

This paper cites Joint-modal label denois- ing for weakly-supervised audio-visual video parsing.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Joint-modal label denois- ing for weakly-supervised audio-visual video parsing

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:11.570835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:02.075272Z digest=sha256:0ed04a432e4e251464f16b3e0e193df5b9d6d1b88634164aaa155058454f2033

Observation 560c4621-b056-414b-a676-215b7a3ed103 · outbound

This paper cites Autoaugment: Learning augmentation strategies from data.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Autoaugment: Learning augmentation strategies from data

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:11.340473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:02.118701Z digest=sha256:017f0a8bbf6e9ba41dabe605ef184484b1a0b8cd325d9c62d04f5fbdb8c1cdc8

Observation 2c8e5958-95e2-4ba1-bbc2-991294eacc14 · outbound

This paper cites Randaugment: Practical automated data augmen- tation with a reduced search space.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Randaugment: Practical automated data augmen- tation with a reduced search space

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:02.188493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:02.188493Z digest=sha256:04ab4db7a7f218736b16dea36659a0e93ad999ff5738a5c13c6af01e139e8206

Observation 452e0cba-d5f6-4472-bb85-350eb5a0da76 · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:02.255554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:02.255554Z digest=sha256:2dd158bfe4dcf3aa30259e35ecc1822cbda911c6930b0ba5f2c27ccdb5a75c64

Observation ce5ceb2c-df9c-4c00-9e7b-463252d0f3dd · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:02.359791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:02.359791Z digest=sha256:2b2329ebed93d93fd9a0c3a4c369364dd9de2a694efe2f94b9d2c1841a1e8bf9

Observation 5aeec829-a26a-4c11-8b4a-a7640cf62ed5 · outbound

This paper cites Cross-modal prompts: Adapting large pre- trained models for audio-visual downstream tasks.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Cross-modal prompts: Adapting large pre- trained models for audio-visual downstream tasks

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:11.081426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:02.464731Z digest=sha256:8721f92b17c30b94ef238ecc6b5680f142c6893cb046fe1033c6968e9bbe1bf0

Observation 226553c4-125e-47df-b314-b453fcafac38 · outbound

This paper cites Revisit weakly-supervised audio-visual video parsing from the lan- guage perspective.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Revisit weakly-supervised audio-visual video parsing from the lan- guage perspective

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:10.846932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:02.567430Z digest=sha256:7872210b41e25d55247d018580abad451bdb64ebce0ae4d0daffc658757b976f

Observation 73015c3d-b739-4224-aefc-aaec94e1dfbb · outbound

This paper cites Col- lecting cross-modal presence-absence evidence for weakly- supervised audio-visual event perception.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Col- lecting cross-modal presence-absence evidence for weakly- supervised audio-visual event perception

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:10.607106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:02.639069Z digest=sha256:eefc168730512b081366ee4777e46242334cf67998a7b7be31f83d3f84b06313

Observation ced513ac-8747-43e7-a380-f0e0a33944b9 · outbound

This paper cites AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:58:08.144148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:02.742986Z digest=sha256:756991c0bd111244ed2181e55dfbbcdb1b7221fe774b166da602cb71d5654ec8

Observation 1a739b27-4492-43ed-a0f6-645d7774e92b · outbound

This paper cites Explaining and Harnessing Adversarial Examples.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Explaining and Harnessing Adversarial Examples

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:02.882381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:02.882381Z digest=sha256:a393acf2ab1231d9cbf47fbef46073e15690f9f239a7247ead4cf9692e18bc91

Observation 151aea69-dc35-4d0a-a214-421f773f98c9 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:02.986186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:02.986186Z digest=sha256:5aef2f892bcc04036c3a7f1f490757ce11471e6308f15a2c6b2ea79a8228eb87

Observation bec3576d-f24e-49b7-bed3-54c3068c74da · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Efficiently Modeling Long Sequences with Structured State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:03.089617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:03.089617Z digest=sha256:6642d84d064d9dadf139f329585d2b1cfdb6edce3d691ec223bfb388c17840d4

Observation f9d2b1b8-0d15-4049-857e-bb6fea2968ad · outbound

This paper cites Combining recurrent, convolutional, and continuous-time models with linear state space layers.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Combining recurrent, convolutional, and continuous-time models with linear state space layers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:10.324169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:03.194398Z digest=sha256:9203fbb7f27470e2c0279564ae02063bceb297ccc7eedb015c2c02ce6d5bcc99

Observation 0b33f58f-cfac-4ef9-93c0-2674aa8c290d · outbound

This paper cites Liquid Structural State-Space Models.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Liquid Structural State-Space Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:03.300791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:03.300791Z digest=sha256:a312f80ba305d938d91c12af1acf05e0d7459800c6a4c37fd1356e8e44da2b56

Observation d2e70f55-3791-4d6a-afba-899058713594 · outbound

This paper cites MambaVision: A Hybrid Mamba-Transformer Vision Backbone.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing MambaVision: A Hybrid Mamba-Transformer Vision Backbone

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:03.377920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:03.377920Z digest=sha256:9e770925ddd4ddf6e943c367b33b45f88d1353000f4a60877ac0715a5370b058

Observation e2dc4534-61f6-4fe7-bdc7-f9fd3e26e966 · outbound

This paper cites Deep residual learning for image recognition.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Deep residual learning for image recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:03.489721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:03.489721Z digest=sha256:62279d2e322450f20818df7a06d597ce25134cf9430f6710eef9265b66714cd8

Observation 27fe4404-ecc1-4c00-ae4b-1174922e344c · outbound

This paper cites AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing AugMix: A Simple Data Processing Method to Improve Robustness and Uncertainty

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:03.569426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:03.569426Z digest=sha256:6c8330f7579cec76b9284630fd89b8511af269fc81fa5e134007069ef3b6976b

Observation 89cb7c1d-36e9-4b82-abe0-ebb6b891b01d · outbound

This paper cites Cnn archi- tectures for large-scale audio classification.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Cnn archi- tectures for large-scale audio classification

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:10.188755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:03.698324Z digest=sha256:18905cafd6674742fa8702301a5b10ab359c2520a635293a7873f9ab4907cac6

Observation 8e53fcd3-3a9c-4f42-96a6-e519130c7bcc · outbound

This paper cites LocalMamba: Visual State Space Model with Windowed Selective Scan.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing LocalMamba: Visual State Space Model with Windowed Selective Scan

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:03.812866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:03.812866Z digest=sha256:db3653fd17142e1ec4a4cdb262196dbc96776b82b7a185c1e63dd1a6f716d65d

Observation 9301dcee-7cf8-4901-b4b6-acc4378b16f5 · outbound

This paper cites Learning tem- porally invariant and localizable features via data augmenta- tion for video recognition.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Learning tem- porally invariant and localizable features via data augmenta- tion for video recognition

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:10.027441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:03.907929Z digest=sha256:1534838a354d8a1a5b355667bea75567bd24c4852b4ddc74c193413f6ec93ef2

Observation 48ceaf10-62c6-4f91-9ae6-f2f7892359f4 · outbound

This paper cites Modality-independent teachers meet weakly-supervised audio-visual event parser.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Modality-independent teachers meet weakly-supervised audio-visual event parser

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:09.899372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:04.022837Z digest=sha256:85618306c4d60a64aed3a3c62ee32a29ff9eb97ef74b87bc2d599be5d22b6862

Observation 0cc45d7e-b710-426a-8604-a17bb499f518 · outbound

This paper cites Coupled Mamba: Enhanced Multi-modal Fusion with Coupled State Space Model.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Coupled Mamba: Enhanced Multi-modal Fusion with Coupled State Space Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:04.083548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:04.083548Z digest=sha256:5256df80aeaed418b28ded8a13d247e4125f5a2b7a55df298b28809c4322071c

Observation 5e794d33-65de-4e8c-96bf-bd59dadc2849 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Jamba: A Hybrid Transformer-Mamba Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:04.186913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:04.186913Z digest=sha256:3abc51054dad90289b0e82fcb6301d28a0f8d53638631992bc0a498680d402f2

Observation ef5608eb-864b-42fd-8792-7251e2f8cf24 · outbound

This paper cites Vmamba: Visual state space model.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Vmamba: Visual state space model

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:09.775664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:04.291318Z digest=sha256:7dba723d1b4764fe4c1abaa5587c7a90099dd104f8ea216c78ed415f06762b72

Observation 65adcd42-ccc8-4f29-a9ab-0883e487c2f5 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Swin transformer: Hierarchical vision transformer using shifted windows

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:04.394262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:04.394262Z digest=sha256:b87e75dcd8e05a847e0c3fa4c42e93c86665cefcb5cf8b7e7b3a7cfa5785daa5

Observation e6c89b0a-ea2d-470f-b872-c34fa8f808a7 · outbound

This paper cites RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing RGBT Tracking via All-layer Multimodal Interactions with Progressive Fusion Mamba

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:58:07.739579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:04.499289Z digest=sha256:b3264599a68629f5632418b70eb4a81221739fce967bd48fce50958264ff54eb

Observation 38faa884-73fb-4956-a11e-51faa5c0e065 · outbound

This paper cites Multi-modal grouping network for weakly-supervised audio-visual video parsing.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Multi-modal grouping network for weakly-supervised audio-visual video parsing

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:09.666464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:04.636366Z digest=sha256:9cbb06c59e20bd6ba70acf94f6acc4f21eabb7c3e36ab14f1bceee1bda228683

Observation 9c56607c-b636-4ffe-b403-bee1d4f7e3a3 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Learning transferable visual models from natural language supervi- sion

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:04.777248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:04.777248Z digest=sha256:ab3d6ff32cdcd75d817d27128f42a03d312c90f34cb6a134ff8e9e9b2466f5aa

Observation 655afca1-a5dc-4906-a1e6-e8d47a33df33 · outbound

This paper cites Coleaf: A contrastive-collaborative learning framework for weakly supervised audio-visual video pars- ing.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Coleaf: A contrastive-collaborative learning framework for weakly supervised audio-visual video pars- ing

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:09.519233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:04.890114Z digest=sha256:4c6b7113daf06603ba5fcfca3468381fc5c0f7482b925180bf801170367b5e1d

Observation 655ab876-e5e7-409e-ad01-f5f0d5185dc8 · outbound

This paper cites Simplified State Space Layers for Sequence Modeling.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Simplified State Space Layers for Sequence Modeling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:04.971510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:04.971510Z digest=sha256:483612abc6b8c3e263a174f08c147a7a19359315026c46ed489fc786f30266d5

Observation 94fc1c35-2f04-4555-a62d-7c0d41389f16 · outbound

This paper cites Audio-visual event localization in unconstrained videos.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Audio-visual event localization in unconstrained videos

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:05.068195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:05.068195Z digest=sha256:e2a5770eb4dc96c063704dfb2ba9e163c06d5ecebc827f6a2720724d997965b2

Observation a4ee575d-c746-4e32-8eb6-9f0ad5b9ba5c · outbound

This paper cites Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:09.382318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:05.149955Z digest=sha256:bedf27d6b3838ad9c826c4103d195ac8df81453c28b4ecb3f0c1ae9f2689e8cf

Observation 562c4318-668e-41cf-bfb4-410212aa1dc0 · outbound

This paper cites Link: Adaptive modality interaction for audio-visual video parsing.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Link: Adaptive modality interaction for audio-visual video parsing

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:09.246132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:05.255570Z digest=sha256:8be2155a6a7d28ab828011f7ee7442c0f4109cc3a202092a6fada46a7b14f7cd

Observation 0605fbe7-85df-4b97-94d1-f91383469d89 · outbound

This paper cites Cbam: Convolutional block attention module.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Cbam: Convolutional block attention module

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:09.136167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:05.306415Z digest=sha256:92f376aa932bbd02d12b0e4a29a90f7d6a2ba24e7d447a099fa786e599b59ebf

Observation 73efefe9-3b44-43bc-b36b-68bdd364fe6b · outbound

This paper cites Exploring heterogeneous clues for weakly-supervised audio-visual video parsing.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Exploring heterogeneous clues for weakly-supervised audio-visual video parsing

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:05.398357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:05.398357Z digest=sha256:5906470f1c69fce565c75dc9283845410391111ad6a1c6b7bf63c9357869c36d

Observation 9882aed9-e2eb-40e2-b052-e07aff7e68c0 · outbound

This paper cites Dual attention matching for audio-visual event localization.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Dual attention matching for audio-visual event localization

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:09.020659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:05.478358Z digest=sha256:e589a4519f780f1e8a3168b9d32af67fc1ca72c3f559ac2d0eacc410d1c13e44

Observation 1d827604-f9d1-4ab5-a83f-76e170833c64 · outbound

This paper cites Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Large-scale con- trastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:05.570892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:05.570892Z digest=sha256:9935f4b75f2c33fb7e90bf3e379236f1118f6d88cb455eaf7bdf2a850f774486

Observation 5ed0db64-e4dd-4aca-90dd-7496901d27f4 · outbound

This paper cites FusionMamba: Dynamic Feature Enhancement for Multimodal Image Fusion with Mamba.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing FusionMamba: Dynamic Feature Enhancement for Multimodal Image Fusion with Mamba

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:05.643282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:05.643282Z digest=sha256:8d8cadbff4db8a78fa34daa7b7c4cbb85ed7cdd73a736a4f8250bde23be8f34a

Observation d365f679-fdf1-4cea-b324-81e40585f0ab · outbound

This paper cites Language- driven all-in-one adverse weather removal.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Language- driven all-in-one adverse weather removal

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:08.921212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:05.718015Z digest=sha256:b1b7e848ce1e655b66cb35d4d3b602659295505001662c0c149bd59413270149

Observation 60fd1cc5-0115-4de8-a785-3cadecebb228 · outbound

This paper cites DepMamba: Progressive Fusion Mamba for Multimodal Depression Detection.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing DepMamba: Progressive Fusion Mamba for Multimodal Depression Detection

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:58:07.426883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:05.783727Z digest=sha256:2051f1c247c725118c952d5c899f9b311f2bb97e89ed66c9488bb039209376c7

Observation d5fd2380-7cf8-4973-af4e-a6605d2aa624 · outbound

This paper cites Text-if: Leveraging semantic text guidance for degradation-aware and interactive image fusion.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Text-if: Leveraging semantic text guidance for degradation-aware and interactive image fusion

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:08.797355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:05.882496Z digest=sha256:e72bc50a1280694a7d4e78624a8d229cd34cb1cc9cd4a9c65efac786b41f9377

Observation 3d7ce16a-c48d-4525-aaaf-f072feda7598 · outbound

This paper cites Mm-pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Mm-pyramid: Multimodal pyramid attentional network for audio-visual event localization and video pars- ing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:08.687711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:06.030283Z digest=sha256:ad378fae2adf395932a46e959b664a1f2ab2727dc427c2e891036da3a65062d2

Observation 8b6998f6-0317-4d24-b6c9-93447fb321da · outbound

This paper cites MambaOut: Do We Really Need Mamba for Vision?.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing MambaOut: Do We Really Need Mamba for Vision?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:06.156919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:06.156919Z digest=sha256:1e8201cfc5f939fac540e38cc42f3fc6303e394d82408d73f6c2a170532cdcd3

Observation b0a227d4-5167-4fd2-b780-5e17d58baadf · outbound

This paper cites Cutmix: Regu- larization strategy to train strong classifiers with localizable features.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Cutmix: Regu- larization strategy to train strong classifiers with localizable features

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T20:58:08.453348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:06.259664Z digest=sha256:37b3d63ecd1ede336edd74bd37e67c1331845c635ea94356d3689b76978b92e1

Observation 9786d213-3270-4def-ba54-3f9774615a9f · outbound

This paper cites VideoMix: Rethinking Data Augmentation for Video Classification.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing VideoMix: Rethinking Data Augmentation for Video Classification

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:06.406992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:06.406992Z digest=sha256:756125545657fe4590b45b7db378b41a89d3e1716a51e81bdfd71d9c8875f509

Observation ede876a5-9423-4eb5-b7a3-9d0cc4aba866 · outbound

This paper cites Don't Judge by the Look: Towards Motion Coherent Video Representation.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Don't Judge by the Look: Towards Motion Coherent Video Representation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:06.514806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:06.514806Z digest=sha256:fde34219e2c37aa5a0c92ed177b0d5c8cefbaecda90d377ed35cc937e272daba

Observation e0ed3644-c223-4aed-be6c-570b4f284f05 · outbound

This paper cites Label-anticipated Event Disentanglement for Audio-Visual Video Parsing.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Label-anticipated Event Disentanglement for Audio-Visual Video Parsing

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:58:07.091761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T20:58:06.622079Z digest=sha256:6a9db97142b9b18eabd79f12dbe2fad1a272326d17af042562ce9fb4f358fb9c

Observation 6e9efd3a-ee14-4b9d-96a0-8c80449c38de · outbound

This paper cites Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model.

MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:06.771420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:06.771420Z digest=sha256:f5d0ea6237c63cff83e9603feee958b122b949bd07497ba751ceb99e45534748

Pith citing papers

Observation fbc685ee-a8fd-4b6e-bbc9-03e16cefc264 · inbound

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing cites this paper.

EAR: Enhancing Uni-Modal Representations for Weakly Supervised Audio-Visual Video Parsing MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:06:19.437712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T03:03:13.496634Z digest=sha256:7d5830debb6684cfeac40de7d41926a9721546f028a4f44cb2986a7e80fd4848

Observation 23fc593e-51da-478f-b991-062c5c2beefc · inbound

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization cites this paper.

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-11T18:34:29.549966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T18:34:29.549966Z digest=sha256:29a27f654dc2468c9e80f5a5a18f606f879d64ee4d6ad0db9ff66e5312eece26

Observation 20d44590-0dd1-45c2-81d0-6b1c111972b9 · inbound

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization cites this paper.

UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization MUG: Pseudo Labeling Augmented Audio-Visual Mamba Network for Audio-Visual Video Parsing

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:53.734483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:40:53.734483Z digest=sha256:a183944ee33c81445fa53b33e81bf3ec21f45b741ccb8dee73e78f1008a2ceda