Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:51:59.930418Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2506.23283.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T21:51:59.930418Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
89 of 89 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d0566584-8b1a-4e89-bc30-b1d8a7a49a6e · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Vivit: A video vision transformer
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a8fd3e-b682-462a-913c-6f3ea3a4878d · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition BEiT: BERT Pre-Training of Image Transformers
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97822a62-148e-4a60-ad27-4c0290eedf20 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Is space-time attention all you need for video under- standing? In Proceedings of the International Confer- ence on Machine Learning (ICML), July 2021
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dd5bef8-715c-41db-a180-87c6bea1eb99 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Coyo-700m: Image-text pair dataset, 2022
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 066d303e-31f0-495a-a9d5-e53f30ab3fcf · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Quo vadis, ac- tion recognition? a new model and the kinetics dataset
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b1b048b-180e-4829-9f5b-7141258a6a96 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 745bc955-cd20-424a-bb12-2e2c8673636e · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition A simple framework for con- trastive learning of visual representations, 2020
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35f85598-e66d-4960-b93e-a9b3d5b956d3 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Feature-wise transformations
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae2ebe95-7d61-4550-8588-d8c48b4c95d4 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Multiscale vision transformers
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56401364-3da8-488c-ad07-e74ededb08f6 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition X3d: Expanding architec- tures for efficient video recognition
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cfc5cce-ae53-43a3-9dfa-93799361f043 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Masked Autoencoders As Spatiotemporal Learners
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c259ede9-b86a-4879-9f91-ae3ec10321ff · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Slowfast networks for video recog- nition
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5da590e6-82da-4603-8c2e-6b59e9ac8884 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition The” something something” video database for learning and evaluating visual common sense
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68690598-06ab-4175-870e-f52779ee66e8 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39b35f9b-124f-4e9d-90f8-3ea5f91027d2 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Efficiently Modeling Long Sequences with Structured State Spaces
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70f67fe2-2691-497b-a62c-b808bbe5a4b0 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Open-vocabulary object detection via vision and language knowledge distillation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3fe1164-933b-4afb-abab-edf756748487 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Trustworthy machine learning: From data to models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 245cfa19-8e54-4bed-a6b5-431633cdeba3 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Turbo training with token dropout
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 75dbef3c-98ee-4f7e-9e9e-9b8c55438318 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning spatio-temporal features with 3d residual networks for action recognition
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fd048668-ac8d-4cbe-87be-bcb5b743aee7 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition MambaVision: A Hybrid Mamba-Transformer Vision Backbone
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4995629-e6d5-4ab9-a680-a5d3da6ba861 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Clipscore: A reference- free evaluation metric for image captioning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation afba095f-b5d7-4a4d-a56f-5b12f2cf32f8 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Arbitrary style trans- fer in real-time with adaptive instance normalization
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 250059a2-427d-4720-9901-4c9cb5b8384e · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition VideoGraph: Recognizing Minutes-Long Human Activities in Videos
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75807e7a-5368-46d2-982f-815c6503e810 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Smeulders
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c017a9e-706f-40cf-af86-3ef5a43b02f3 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Long movie clip classification with state-space video mod- els
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 034efcb6-bfc0-4177-9cbf-e9149da6fe12 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Scaling up visual and vision- language representation learning with noisy text su- pervision
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06acb9ba-58e8-401a-b5ea-eddbfbff387e · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Laine, and Timo Aila
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c231f89a-87ef-43fe-b2bd-7cd483c811f4 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition The Kinetics Human Action Video Dataset
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e4c4956-8e4e-498e-9b53-449f2cedceab · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Kuehne, H
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8821f97d-59ee-4773-9bee-8fc6a466d9c6 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition The lan- guage of actions: Recovering the syntax and semantics of goal-directed human activities
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11c6d7db-7c34-4bcf-985a-a185bbf26bb1 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 965818a5-d127-4beb-b212-05a651f0ad89 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Mipmap-GS: Let Gaussians Deform with Scale-specific Mipmap for Anti-aliasing Rendering
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6c983faf-b075-426a-beb7-3bd8197b733e · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Videomamba: State space model for efficient video understanding, 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 02fcb1cf-32a1-4da0-b55e-7c84cb0018dc · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Unmasked teacher: Towards training-efficient video foundation models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27f5fb1b-f812-4cce-97ce-0a366490a1c7 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Uniformer: Uni- fied transformer for efficient spatial-temporal repre- sentation learning
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 26abb7e7-33e8-4293-b6cf-5b079c84c844 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Mvitv2: Improved multiscale vision transformers for classification and detection
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1141e522-dbbe-4950-b689-f4c2bf95f6ed · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Fo- caldreamer: Text-driven 3d editing via focal-fusion assembly
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83f7f27f-ffb3-4433-80db-6750c1eb9ce2 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Pointmamba: A simple state space model for point cloud analysis
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 096c446c-94c6-459f-8729-50371225bc84 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Jamba: A Hybrid Transformer-Mamba Language Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e29d25ed-702c-46a4-ac21-893c111fc511 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning to recognize procedural activities with dis- tant supervision
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 82ff3b0f-f212-4e6f-88d6-b2b3490609bf · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Frozen CLIP Models are Efficient Video Learners
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9898fcdc-ced0-4e8e-9241-e170e65e6353 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Annotation-free Audio-Visual Segmentation
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b070c1c1-bf0e-4edd-b688-d0623e3d3f36 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition VMamba: Visual State Space Model
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3b47763b-25ec-4a8e-bbc0-c867acec771c · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Video swin transformer
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6170fa3d-63f0-4201-b1f0-b538fecb1aa3 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Freesegdiff: Annotation-free saliency segmentation with diffusion models
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03621015-443a-4b08-92ce-7b5cb1f70656 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition DiffusionSeg: Adapting Diffusion Towards Unsupervised Object Discovery
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e52b8b84-b3f9-4342-bf72-92861215bdd6 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Open-vocabulary semantic segmenta- tion with frozen vision-language models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 69c49ab9-c49e-4727-87ca-5df1356d7e02 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Attrseg: open- vocabulary semantic segmentation via attribute decomposition-aggregation
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2573ff8e-3f54-4079-b207-f54ad18e4c47 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Scaling open-vocabulary object detection
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 46177889-2581-457d-bd97-0e86f0a016ed · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ClipCap: CLIP Prefix for Image Captioning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bad67b6-8264-424e-98ae-4d4ea9723705 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Expanding Language-Image Pretrained Models for General Video Recognition
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 00a9a893-7320-4bb5-884a-67d9c6c1167c · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Text-Only Training for Image Captioning using Noise-Injected CLIP
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fceb07c-de6f-4b34-9b3b-72af3549bc11 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 600a5af7-0e65-45fa-9ef4-dc11436ed1f6 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 47ad84a8-cfe1-489d-8460-5ab7d35ae11f · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition MambaSCI: Efficient Mamba-UNet for Quad-Bayer Patterned Video Snapshot Compressive Imaging
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 01124fa0-a327-49ef-ac19-31cb7f5beac9 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Dual-path adaptation from image to video transform- ers
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56679b2e-1c87-4056-a665-fe6bc03ccdb6 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Peebles and Saining Xie
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e3362da8-a59f-4667-bb1e-fe3b44df7dbd · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Unresolved cited work
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 09d3aab8-921b-417a-ba0d-7e92b0719a3c · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Dis- entangling spatial and temporal learning for efficient image-to-video transfer learning
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c805022-7f59-4adb-be59-353adf757540 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning transferable visual models from natural lan- guage supervision
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ead28b01-3e89-46be-8e2f-ef446ded4382 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Token- learner: Adaptive space-time tokenization for videos
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7c17898b-796b-46d0-afdd-8836a651f2b1 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition LAION-5B: An open large-scale dataset for training next generation image-text models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b354a31-49bd-4357-9a09-10a6cd7e5e34 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Only time can tell: Discovering temporal data for temporal modeling
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c0e1250-f422-4fe9-9601-1a08ccdc984d · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Darf: Depth- aware generalizable neural radiance field
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 28917599-655e-4971-8114-058987c3fdcb · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 779064c5-4c2f-41b4-a898-307d77d7c2ca · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Coin: A large-scale dataset for comprehensive instruc- tional video analysis
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fbdc6a50-7eec-48f7-8c45-b34de633b22f · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Dim: Diffusion mamba for efficient high-resolution image synthesis, 2024
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af449d3e-434b-4f99-ada6-63a5585c1387 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c40be8d-6286-4410-9506-79a9f25f2837 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition An Empirical Study of Mamba-based Language Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3092b77-db40-4209-bc01-513dabb0a114 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fefb174e-7ad2-4264-838b-62483d589634 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ActionCLIP: A New Paradigm for Video Action Recognition
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 897aba94-bd40-4570-a160-c81813c23320 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cdd6186-d78a-4fba-8888-cf116c13b917 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition PoinTramba: A Hybrid Transformer-Mamba Framework for Point Cloud Analysis
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94f0fbf7-8e97-43b8-a1e3-e50cb9b47a84 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition DualNeRF: Text-Driven 3D Scene Editing via Dual-Field Representation
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d526405-4e77-4079-a62e-68b17786e9ee · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90e734f1-29e5-4dc2-8f66-3a7369f14b76 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Multiview transformers for video recognition
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 68feddfa-c382-4fae-bb7d-49084f20ac09 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition MambaMIL: Enhancing Long Sequence Modeling with Sequence Reordering in Computational Pathology
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2ed0136-122a-4a41-902c-67b430d24944 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition AIM: Adapting image mod- els for efficient video action recognition
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8c6e1585-036c-4a36-987c-68a3ef9dc655 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Multi- modal prototypes for open-world semantic segmen- tation
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f07b56a4-579a-47ed-ad1b-744a7fafac25 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Remamber: Referring image segmentation with mamba twister
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c229ec09-62d6-4d04-b8e3-a67848f1fee9 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning with multi- class auc: Theory and algorithms
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9a409808-e05f-4f5c-a661-b69f11de62a6 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Optimizing two-way partial auc with an end-to-end framework
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7cbfadaf-943e-4673-8f37-e4d2aad83fb1 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Sigmoid Loss for Language Image Pre-Training
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c448c7b1-0717-4d7e-baff-928868910d1e · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Point Cloud Mamba: Point Cloud Learning via State Space Model
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 613a4a46-2840-47e1-a6a4-eb0402e8d530 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c101d2c5-237a-4331-8e65-f43b9565ab7e · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Vidtr: Video transformer without convo- lutions
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8014a4aa-3a2c-4def-8347-74f7b1beac90 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Regionclip: Region-based language-image pretraining
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6deeac2c-3a55-4d5d-8cb7-97eb4bc75e57 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Graph-based high-order relation modeling for long-term action recognition
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ce255339-f313-4b06-8e78-b21533f4e672 · outbound
MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.