Pith. sign in

Paper Citation Record · LEDGER

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition

As of 9 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 0 inbound Pith citation observations for arXiv:2506.23283.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23283 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:51:59.930418Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

89 of 89 outbound references displayed

  • verified exact4
  • verified fuzzy43
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0566584-8b1a-4e89-bc30-b1d8a7a49a6e · outbound

This paper cites Vivit: A video vision transformer.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Vivit: A video vision transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.381794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.381794Z digest=sha256:b26e42bd0f1dc9c0aa10900142dd50980acd885a5a17beb2ef0e062e87185e1d

Observation 18a8fd3e-b682-462a-913c-6f3ea3a4878d · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition BEiT: BERT Pre-Training of Image Transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.473450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.473450Z digest=sha256:877684fd48bcf0382948898dfc89e9a6f62275c0d422faa3d86e92a7919c5fea

Observation 97822a62-148e-4a60-ad27-4c0290eedf20 · outbound

This paper cites Is space-time attention all you need for video under- standing? In Proceedings of the International Confer- ence on Machine Learning (ICML), July 2021.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Is space-time attention all you need for video under- standing? In Proceedings of the International Confer- ence on Machine Learning (ICML), July 2021

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.566560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.566560Z digest=sha256:78763d06b93d87956228ae11b08087342d1215a422e265bcc9c11490ac92a109

Observation 0dd5bef8-715c-41db-a180-87c6bea1eb99 · outbound

This paper cites Coyo-700m: Image-text pair dataset, 2022.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Coyo-700m: Image-text pair dataset, 2022

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.684318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.684318Z digest=sha256:c69f013154db7f6d21036128b91dc5ed68af0a28331cbad7135926281bb29978

Observation 066d303e-31f0-495a-a9d5-e53f30ab3fcf · outbound

This paper cites Quo vadis, ac- tion recognition? a new model and the kinetics dataset.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Quo vadis, ac- tion recognition? a new model and the kinetics dataset

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.823473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.823473Z digest=sha256:bc6192ff61fd3d56cdb383c96f2f7b8a41e5c9a64fbb1fa304d8215b14c796e5

Observation 6b1b048b-180e-4829-9f5b-7141258a6a96 · outbound

This paper cites Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:50.963310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:50.963310Z digest=sha256:1077a23ba8ac03b5dc0db158cc3ee20168a0ffa1eb6d07425f07f84023aef40d

Observation 745bc955-cd20-424a-bb12-2e2c8673636e · outbound

This paper cites A simple framework for con- trastive learning of visual representations, 2020.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition A simple framework for con- trastive learning of visual representations, 2020

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.088191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.088191Z digest=sha256:9d6321c954ced0aba8a7c52bf5c7757332bef0cbfdd2e84cb6c7aee87d017c47

Observation 35f85598-e66d-4960-b93e-a9b3d5b956d3 · outbound

This paper cites Feature-wise transformations.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Feature-wise transformations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.215965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.215965Z digest=sha256:09d92daf74b30d9f3bf01dce3c01d072e9b2dbed9573502fd5fa58b86ac7a0a4

Observation ae2ebe95-7d61-4550-8588-d8c48b4c95d4 · outbound

This paper cites Multiscale vision transformers.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Multiscale vision transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.324371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.324371Z digest=sha256:dde13b9164745448762bc3a8f62fb8551d5ec7cfbb11fb45c1b67754436a8c73

Observation 56401364-3da8-488c-ad07-e74ededb08f6 · outbound

This paper cites X3d: Expanding architec- tures for efficient video recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition X3d: Expanding architec- tures for efficient video recognition

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.451686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.451686Z digest=sha256:7c71643e82250e80f2f5d2ac7bde9a0bb584381a850f5e917b1a5da21ecde48f

Observation 2cfc5cce-ae53-43a3-9dfa-93799361f043 · outbound

This paper cites Masked Autoencoders As Spatiotemporal Learners.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Masked Autoencoders As Spatiotemporal Learners

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.583049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.583049Z digest=sha256:2dc96e1a04f32b0246b35f342418dadd7f89367f2b217b9c8b3139fd7d892a97

Observation c259ede9-b86a-4879-9f91-ae3ec10321ff · outbound

This paper cites Slowfast networks for video recog- nition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Slowfast networks for video recog- nition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.697762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.697762Z digest=sha256:823a7b8f1acd7ac4c22261e2651edddfe5a89cff44a9f96830aad823c4bb8cd2

Observation 5da590e6-82da-4603-8c2e-6b59e9ac8884 · outbound

This paper cites The” something something” video database for learning and evaluating visual common sense.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition The” something something” video database for learning and evaluating visual common sense

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:12.327642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:51.830505Z digest=sha256:ffdf8de8602da0fc174358d7295a685ead8b18d47685915a0c2caa9e5d9911e3

Observation 68690598-06ab-4175-870e-f52779ee66e8 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:51.926035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:51.926035Z digest=sha256:78ffaa4a872b63d8d9b8121661bb036279ce1dc0ced2348f42e817c00a8ebab2

Observation 39b35f9b-124f-4e9d-90f8-3ea5f91027d2 · outbound

This paper cites Efficiently Modeling Long Sequences with Structured State Spaces.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Efficiently Modeling Long Sequences with Structured State Spaces

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:52.045734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:52.045734Z digest=sha256:547a67b7ee2701a7fc4412414db6cac515c7549878835f7c80f6479914679ce2

Observation 70f67fe2-2691-497b-a62c-b808bbe5a4b0 · outbound

This paper cites Open-vocabulary object detection via vision and language knowledge distillation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Open-vocabulary object detection via vision and language knowledge distillation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:12.279929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:52.127801Z digest=sha256:2c84f652160f500770b923e8a5f4186770ed4e560b120a17334e914fa4718a74

Observation d3fe1164-933b-4afb-abab-edf756748487 · outbound

This paper cites Trustworthy machine learning: From data to models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Trustworthy machine learning: From data to models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:12.195299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:52.256779Z digest=sha256:188381ea11e678fe9ae970ac7e98e71a358bb399ee42710cb637bf732e575181

Observation 245cfa19-8e54-4bed-a6b5-431633cdeba3 · outbound

This paper cites Turbo training with token dropout.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Turbo training with token dropout

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:12.084240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:52.390034Z digest=sha256:99aaf2562df227b98c51a02a67849c8f685d6f26ef216b5b77f09c826593c73a

Observation 75dbef3c-98ee-4f7e-9e9e-9b8c55438318 · outbound

This paper cites Learning spatio-temporal features with 3d residual networks for action recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning spatio-temporal features with 3d residual networks for action recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.978319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:52.549772Z digest=sha256:142b74a0c6eb7926dce84ce433dcd21909bde7889689167a1712760482865cbe

Observation fd048668-ac8d-4cbe-87be-bcb5b743aee7 · outbound

This paper cites MambaVision: A Hybrid Mamba-Transformer Vision Backbone.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition MambaVision: A Hybrid Mamba-Transformer Vision Backbone

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:52.633842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:52.633842Z digest=sha256:0a1e4c9b134a1ef662d18f3d59101e14da7b2fedfa67d1a7c5329871b9409042

Observation a4995629-e6d5-4ab9-a680-a5d3da6ba861 · outbound

This paper cites Clipscore: A reference- free evaluation metric for image captioning.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Clipscore: A reference- free evaluation metric for image captioning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.875378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:52.745648Z digest=sha256:4a4690e5ff8b69aa8cfef9cfa1495965670d400110a24ba8084e2322cc32e9c7

Observation afba095f-b5d7-4a4d-a56f-5b12f2cf32f8 · outbound

This paper cites Arbitrary style trans- fer in real-time with adaptive instance normalization.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Arbitrary style trans- fer in real-time with adaptive instance normalization

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.750886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:52.831996Z digest=sha256:05f170efeefdbdd26a52e4c52ff9f1ab624865d7c440d15426fda38c20caa50e

Observation 250059a2-427d-4720-9901-4c9cb5b8384e · outbound

This paper cites VideoGraph: Recognizing Minutes-Long Human Activities in Videos.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition VideoGraph: Recognizing Minutes-Long Human Activities in Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:52.943841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:52.943841Z digest=sha256:254ce8fe4bac6bc4a2b01d8dec1cdcc104d7633fa9ba8d11aa1ac37f9dc78c1d

Observation 75807e7a-5368-46d2-982f-815c6503e810 · outbound

This paper cites Smeulders.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Smeulders

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.626186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:53.022920Z digest=sha256:3a86b96ff3a033bd3f62c8c83528859f61ba6ecda1a522e55bf737028c5ff52f

Observation 5c017a9e-706f-40cf-af86-3ef5a43b02f3 · outbound

This paper cites Long movie clip classification with state-space video mod- els.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Long movie clip classification with state-space video mod- els

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.504852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:53.133192Z digest=sha256:c31a72f8ff91db6880f62e67bc88735ae010fd140bd03b11ac3ec3c1ea855ab5

Observation 034efcb6-bfc0-4177-9cbf-e9149da6fe12 · outbound

This paper cites Scaling up visual and vision- language representation learning with noisy text su- pervision.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Scaling up visual and vision- language representation learning with noisy text su- pervision

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.396778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:53.286576Z digest=sha256:075c2e1a042a1bbf521e117d33804f1cb0de20b7e8162da50fffa3cff0856623

Observation 06acb9ba-58e8-401a-b5ea-eddbfbff387e · outbound

This paper cites Laine, and Timo Aila.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Laine, and Timo Aila

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.285784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:53.420848Z digest=sha256:950d8132353d1c6897a8d3fe545f38d53a713eb6780e573d6bdf1b69090aa9fa

Observation c231f89a-87ef-43fe-b2bd-7cd483c811f4 · outbound

This paper cites The Kinetics Human Action Video Dataset.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition The Kinetics Human Action Video Dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:53.493512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:53.493512Z digest=sha256:a9959363ebff7bdf1b5e68b8632b65719994040c5cdf54e7b399305e43c7b26b

Observation 4e4c4956-8e4e-498e-9b53-449f2cedceab · outbound

This paper cites Kuehne, H.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Kuehne, H

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.170200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:53.561748Z digest=sha256:0009288721a9af5e525b7053fce715485639757f6689faf50ccdcfac190232c9

Observation 8821f97d-59ee-4773-9bee-8fc6a466d9c6 · outbound

This paper cites The lan- guage of actions: Recovering the syntax and semantics of goal-directed human activities.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition The lan- guage of actions: Recovering the syntax and semantics of goal-directed human activities

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:11.036537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:53.693704Z digest=sha256:f7c21e3a15a6a468e3377672926e2309bd679e3837feaa812b1e3d4bbd985708

Observation 11c6d7db-7c34-4bcf-985a-a185bbf26bb1 · outbound

This paper cites F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition F-VLM: Open-Vocabulary Object Detection upon Frozen Vision and Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:53.789728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:53.789728Z digest=sha256:a4f4ad40eeefd60afdfa6312279339cd06c6df4ba13ecde5e6252047f29e177b

Observation 965818a5-d127-4beb-b212-05a651f0ad89 · outbound

This paper cites Mipmap-GS: Let Gaussians Deform with Scale-specific Mipmap for Anti-aliasing Rendering.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Mipmap-GS: Let Gaussians Deform with Scale-specific Mipmap for Anti-aliasing Rendering

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:52:01.162745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:53.897777Z digest=sha256:42f9b98562009ce1114ae830082a3788793765f8a683733c6f0d554ff5d7d183

Observation 6c983faf-b075-426a-beb7-3bd8197b733e · outbound

This paper cites Videomamba: State space model for efficient video understanding, 2024.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Videomamba: State space model for efficient video understanding, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:10.830280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:54.013868Z digest=sha256:6103e9ec406febbae48b6d739659b3876b65d9400db86ea238f51b5803c2cb43

Observation 02fcb1cf-32a1-4da0-b55e-7c84cb0018dc · outbound

This paper cites Unmasked teacher: Towards training-efficient video foundation models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Unmasked teacher: Towards training-efficient video foundation models

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:10.653523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:54.127628Z digest=sha256:423c24e909dff885e5f567dfed818e7b5fd32f0c2b804203ff93cd83ad7485aa

Observation 27f5fb1b-f812-4cce-97ce-0a366490a1c7 · outbound

This paper cites Uniformer: Uni- fied transformer for efficient spatial-temporal repre- sentation learning.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Uniformer: Uni- fied transformer for efficient spatial-temporal repre- sentation learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:10.465539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:54.238037Z digest=sha256:827f5f88eddef5140acba134ee9432005440a24b37cb51064eb00b76571aa552

Observation 26abb7e7-33e8-4293-b6cf-5b079c84c844 · outbound

This paper cites Mvitv2: Improved multiscale vision transformers for classification and detection.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Mvitv2: Improved multiscale vision transformers for classification and detection

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:10.330276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:54.367511Z digest=sha256:5c1983deb3be80f65f0b2a060ad62bfef7bef116cbb32c379a7f4e6ac5b28708

Observation 1141e522-dbbe-4950-b689-f4c2bf95f6ed · outbound

This paper cites Fo- caldreamer: Text-driven 3d editing via focal-fusion assembly.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Fo- caldreamer: Text-driven 3d editing via focal-fusion assembly

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:10.141385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:54.431653Z digest=sha256:0fe182e6d53f6c1322fe5701aa7636f3a241ce2710a723528b1e1e6e299ce191

Observation 83f7f27f-ffb3-4433-80db-6750c1eb9ce2 · outbound

This paper cites Pointmamba: A simple state space model for point cloud analysis.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Pointmamba: A simple state space model for point cloud analysis

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.960882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:54.524479Z digest=sha256:3974b588710a7d8837ed2dff0aa11f777c67815a7b97cfbb420235ea145dc249

Observation 096c446c-94c6-459f-8729-50371225bc84 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Jamba: A Hybrid Transformer-Mamba Language Model

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:54.612897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:54.612897Z digest=sha256:e807b3e158ae7b8670cab7e7f7eb8a9acd60cae4d1afc2b29d3bdfb9312b3ba9

Observation e29d25ed-702c-46a4-ac21-893c111fc511 · outbound

This paper cites Learning to recognize procedural activities with dis- tant supervision.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning to recognize procedural activities with dis- tant supervision

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.822263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:54.696571Z digest=sha256:2c4dc3d142e37c27097997e99b27cef73110a2778469887cc9a8d099afe70e33

Observation 82ff3b0f-f212-4e6f-88d6-b2b3490609bf · outbound

This paper cites Frozen CLIP Models are Efficient Video Learners.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Frozen CLIP Models are Efficient Video Learners

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:54.788646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:54.788646Z digest=sha256:d3d5dd9f9e731436b9bb2fd69912ccc4359e946c7ce691340070bd28bed73801

Observation 9898fcdc-ced0-4e8e-9241-e170e65e6353 · outbound

This paper cites Annotation-free Audio-Visual Segmentation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Annotation-free Audio-Visual Segmentation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:54.889200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:54.889200Z digest=sha256:2cdf59166815b1528e66051d8c430f0536d57cdb61b094adfd62409936e280ca

Observation b070c1c1-bf0e-4edd-b688-d0623e3d3f36 · outbound

This paper cites VMamba: Visual State Space Model.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition VMamba: Visual State Space Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:54.973660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:54.973660Z digest=sha256:40659b0643c26d64e6dd0b63907399d284fde30940711b3502167d8cb2971e8b

Observation 3b47763b-25ec-4a8e-bbc0-c867acec771c · outbound

This paper cites Video swin transformer.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Video swin transformer

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.670936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:55.081355Z digest=sha256:9a014063909321685ba350732780c1182f5ea77af1d210576da878db0997cc26

Observation 6170fa3d-63f0-4201-b1f0-b538fecb1aa3 · outbound

This paper cites Freesegdiff: Annotation-free saliency segmentation with diffusion models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Freesegdiff: Annotation-free saliency segmentation with diffusion models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.524175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:55.165046Z digest=sha256:495f7f0498e6f1cbf5a5cdb500d08ddda63605e302542426f52fd21deca7f5e7

Observation 03621015-443a-4b08-92ce-7b5cb1f70656 · outbound

This paper cites DiffusionSeg: Adapting Diffusion Towards Unsupervised Object Discovery.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition DiffusionSeg: Adapting Diffusion Towards Unsupervised Object Discovery

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:55.235358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:55.235358Z digest=sha256:eff615919cfc60229d72ab2dcd3c76879d38557dab7650eed45168d053cee818

Observation e52b8b84-b3f9-4342-bf72-92861215bdd6 · outbound

This paper cites Open-vocabulary semantic segmenta- tion with frozen vision-language models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Open-vocabulary semantic segmenta- tion with frozen vision-language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.336866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:55.327089Z digest=sha256:28cd55649140f5ff801a9a0f665c0e077598d556fc687f32e863a554381de0d0

Observation 69c49ab9-c49e-4727-87ca-5df1356d7e02 · outbound

This paper cites Attrseg: open- vocabulary semantic segmentation via attribute decomposition-aggregation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Attrseg: open- vocabulary semantic segmentation via attribute decomposition-aggregation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:09.120762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:55.420685Z digest=sha256:bbbc475cc52d81e45caeb25d798bd8bc74a26dd30513c79808d90c3beefac8f6

Observation 2573ff8e-3f54-4079-b207-f54ad18e4c47 · outbound

This paper cites Scaling open-vocabulary object detection.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Scaling open-vocabulary object detection

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:07.890235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:55.482406Z digest=sha256:98675f4903169a5287c34ef7bfe9b0827f2383810e0ecd47b2ed4d9d44141bbc

Observation 46177889-2581-457d-bd97-0e86f0a016ed · outbound

This paper cites ClipCap: CLIP Prefix for Image Captioning.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ClipCap: CLIP Prefix for Image Captioning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:55.557921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:55.557921Z digest=sha256:1554996e44604561f190ae4db274bd9a14a2303fca4ae53955e3ff3cde2ef5d5

Observation 8bad67b6-8264-424e-98ae-4d4ea9723705 · outbound

This paper cites Expanding Language-Image Pretrained Models for General Video Recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Expanding Language-Image Pretrained Models for General Video Recognition

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:52:00.859981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:55.663136Z digest=sha256:c4c32f02c3b5109c9a9bd42908f1de970aa2c24169b2a390d9a842a72f1cd39b

Observation 00a9a893-7320-4bb5-884a-67d9c6c1167c · outbound

This paper cites Text-Only Training for Image Captioning using Noise-Injected CLIP.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Text-Only Training for Image Captioning using Noise-Injected CLIP

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:55.746646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:55.746646Z digest=sha256:3882a47e90b07e4891aaab63a9fdae23d5af697b28a465fb180d183c4e8de113

Observation 1fceb07c-de6f-4b34-9b3b-72af3549bc11 · outbound

This paper cites an unresolved cited work.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:07.162702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:55.855563Z digest=sha256:a01d8471d431e9ff9c1a5bc024e23143f83d3dc2bcc95685375d24e729c54a95

Observation 600a5af7-0e65-45fa-9ef4-dc11436ed1f6 · outbound

This paper cites ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ST-Adapter: Parameter-Efficient Image-to-Video Transfer Learning

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:52:00.638728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:55.932723Z digest=sha256:84d3479f16efecb19c2fd58d5d9000a89c4e1acdb58dea987d2975f58e0e8da0

Observation 47ad84a8-cfe1-489d-8460-5ab7d35ae11f · outbound

This paper cites MambaSCI: Efficient Mamba-UNet for Quad-Bayer Patterned Video Snapshot Compressive Imaging.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition MambaSCI: Efficient Mamba-UNet for Quad-Bayer Patterned Video Snapshot Compressive Imaging

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:52:00.425970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:56.065967Z digest=sha256:e8eae3145b469ea711007359c4e87abef57ab5d429e6a0a4d0b9a90e9b813331

Observation 01124fa0-a327-49ef-ac19-31cb7f5beac9 · outbound

This paper cites Dual-path adaptation from image to video transform- ers.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Dual-path adaptation from image to video transform- ers

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:06.981489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:56.170231Z digest=sha256:ef9a2dd9eb1a4b4c08244d0ef732c3896267559dd629ce36e3e381abc2ee8683

Observation 56679b2e-1c87-4056-a665-fe6bc03ccdb6 · outbound

This paper cites Peebles and Saining Xie.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Peebles and Saining Xie

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:06.766762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:56.291434Z digest=sha256:da8e3ab07bc08e39f2189909d9a337dcc151bb74ffef8d172ad11c50842c280b

Observation e3362da8-a59f-4667-bb1e-fe3b44df7dbd · outbound

This paper cites an unresolved cited work.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:52:06.501535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:56.418109Z digest=sha256:d8281cc540fd01e404d01f0fed7ad8da307cdf911a47543064e3baf671b7cb2e

Observation 09d3aab8-921b-417a-ba0d-7e92b0719a3c · outbound

This paper cites Dis- entangling spatial and temporal learning for efficient image-to-video transfer learning.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Dis- entangling spatial and temporal learning for efficient image-to-video transfer learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:06.295700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:56.546195Z digest=sha256:b5046e11deba4d30167236cbfc26fdd23ca5293c31b5074e5d164e881b76dd30

Observation 4c805022-7f59-4adb-be59-353adf757540 · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning transferable visual models from natural lan- guage supervision

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:06.119792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:56.659418Z digest=sha256:0c25ce68af23f7d69b117b2fec2dbcb79b364749b1a465e754f2305454d0821e

Observation ead28b01-3e89-46be-8e2f-ef446ded4382 · outbound

This paper cites Token- learner: Adaptive space-time tokenization for videos.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Token- learner: Adaptive space-time tokenization for videos

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:05.975865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:56.759866Z digest=sha256:94bc367bf5b0521f601c257b2590935bdb3d5b0a357cc7c061c807eae5848edd

Observation 7c17898b-796b-46d0-afdd-8836a651f2b1 · outbound

This paper cites LAION-5B: An open large-scale dataset for training next generation image-text models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition LAION-5B: An open large-scale dataset for training next generation image-text models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:56.863163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:56.863163Z digest=sha256:32c5933b9bdc38ef4a92e36cc77f32cd0529b993c746ad50f0af033b8dc17fb4

Observation 8b354a31-49bd-4357-9a09-10a6cd7e5e34 · outbound

This paper cites Only time can tell: Discovering temporal data for temporal modeling.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Only time can tell: Discovering temporal data for temporal modeling

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:05.790063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:56.962660Z digest=sha256:54ae4d5c036ba73a8c1570fdbc5a947ff09400e68fe6c72925035ad2acb2802a

Observation 8c0e1250-f422-4fe9-9601-1a08ccdc984d · outbound

This paper cites Darf: Depth- aware generalizable neural radiance field.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Darf: Depth- aware generalizable neural radiance field

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:05.620347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:57.082523Z digest=sha256:cc207ef1cc7a18b24a5a767e31d0a301301cd8ba6463763caf4b914f32b43b05

Observation 28917599-655e-4971-8114-058987c3fdcb · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.234563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.234563Z digest=sha256:d6f6fb6774ee1e2e8f38c3f70a65a2f31041dab51de50106dd1b8a4b889987c1

Observation 779064c5-4c2f-41b4-a898-307d77d7c2ca · outbound

This paper cites Coin: A large-scale dataset for comprehensive instruc- tional video analysis.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Coin: A large-scale dataset for comprehensive instruc- tional video analysis

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:05.432865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:57.360492Z digest=sha256:88806449e2c882ac53b0e16e29e10809dacfeaa515e6e40e22d317183d76ee8e

Observation fbdc6a50-7eec-48f7-8c45-b34de633b22f · outbound

This paper cites Dim: Diffusion mamba for efficient high-resolution image synthesis, 2024.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Dim: Diffusion mamba for efficient high-resolution image synthesis, 2024

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:03.938654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:57.464332Z digest=sha256:c434bec0cb8b18235b3b56d16900c3789c1aa170610b3259aa10181adda69659

Observation af449d3e-434b-4f99-ada6-63a5585c1387 · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.583666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.583666Z digest=sha256:0494c45426d224734008294bf8ead0e4c5d60d475aebed7716c45b3d15d075dc

Observation 0c40be8d-6286-4410-9506-79a9f25f2837 · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition An Empirical Study of Mamba-based Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.712328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.712328Z digest=sha256:9a5a1cc2b63b3d24b5ededc9ba3c9fe8d36e978f49604491b383142aa95cbd96

Observation a3092b77-db40-4209-bc01-513dabb0a114 · outbound

This paper cites Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Sigma: Siamese Mamba Network for Multi-Modal Semantic Segmentation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.852368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.852368Z digest=sha256:f7c69269b949c7591911946386d404f0be9500c25af3d2b33451f56f78273974

Observation fefb174e-7ad2-4264-838b-62483d589634 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition ActionCLIP: A New Paradigm for Video Action Recognition

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:57.921969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:57.921969Z digest=sha256:634ceb5e8931bafce9ed27df2119c9c92ede8ca0caa521a3419b93e809cca4c1

Observation 897aba94-bd40-4570-a160-c81813c23320 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.007202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.007202Z digest=sha256:f22bf3ea993f54cf0de22a1616fa1e8f718f2956f7b30b1c93eaa87b304f697a

Observation 6cdd6186-d78a-4fba-8888-cf116c13b917 · outbound

This paper cites PoinTramba: A Hybrid Transformer-Mamba Framework for Point Cloud Analysis.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition PoinTramba: A Hybrid Transformer-Mamba Framework for Point Cloud Analysis

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.110217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.110217Z digest=sha256:76478c2f5cb8d77e72835cee870436c9926d1f0e525ab19cdd680011a4e3c137

Observation 94f0fbf7-8e97-43b8-a1e3-e50cb9b47a84 · outbound

This paper cites DualNeRF: Text-Driven 3D Scene Editing via Dual-Field Representation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition DualNeRF: Text-Driven 3D Scene Editing via Dual-Field Representation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.193805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.193805Z digest=sha256:669dfb73156ffaa8fc839137b8d706acbdfb6ede9e469c558b410e6e45d1a954

Observation 7d526405-4e77-4079-a62e-68b17786e9ee · outbound

This paper cites CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition CLIP-ViP: Adapting Pre-trained Image-Text Model to Video-Language Representation Alignment

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.251948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.251948Z digest=sha256:4e455c3c56886ab0efa6f30007925178d35bcb64a7f589ec43f026a58fcd4ee6

Observation 90e734f1-29e5-4dc2-8f66-3a7369f14b76 · outbound

This paper cites Multiview transformers for video recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Multiview transformers for video recognition

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:03.304085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:58.324442Z digest=sha256:a77f1a4185f6ff3c5e5232c4fb267958b47999b1861af50a1ce468a48b87d661

Observation 68feddfa-c382-4fae-bb7d-49084f20ac09 · outbound

This paper cites MambaMIL: Enhancing Long Sequence Modeling with Sequence Reordering in Computational Pathology.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition MambaMIL: Enhancing Long Sequence Modeling with Sequence Reordering in Computational Pathology

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:58.389943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:58.389943Z digest=sha256:8570e617cf9b31d24d175d5713ec79c0ddf9db97cdb20b5b3ec48fa1a75e6eca

Observation f2ed0136-122a-4a41-902c-67b430d24944 · outbound

This paper cites AIM: Adapting image mod- els for efficient video action recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition AIM: Adapting image mod- els for efficient video action recognition

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:03.146644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:58.460734Z digest=sha256:0d6604b13826c7ecd927f73c76d69a32db05b9166d79995593960585deb7e76e

Observation 8c6e1585-036c-4a36-987c-68a3ef9dc655 · outbound

This paper cites Multi- modal prototypes for open-world semantic segmen- tation.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Multi- modal prototypes for open-world semantic segmen- tation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:02.959704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:58.591097Z digest=sha256:1b90ee1f259db8e814a391c32634a0d6c0d9409247f4d9374782cfda179aab93

Observation f07b56a4-579a-47ed-ad1b-744a7fafac25 · outbound

This paper cites Remamber: Referring image segmentation with mamba twister.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Remamber: Referring image segmentation with mamba twister

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:02.772877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:58.763767Z digest=sha256:4d0c3037bd485916b70361e9f5411a3bf4ad603161b4c4b4dd9e7de73008ef47

Observation c229ec09-62d6-4d04-b8e3-a67848f1fee9 · outbound

This paper cites Learning with multi- class auc: Theory and algorithms.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Learning with multi- class auc: Theory and algorithms

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:02.585717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:58.912661Z digest=sha256:f22775c308834a7707c1df6e92905af62103660a394ef3e14a25044a08854f22

Observation 9a409808-e05f-4f5c-a661-b69f11de62a6 · outbound

This paper cites Optimizing two-way partial auc with an end-to-end framework.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Optimizing two-way partial auc with an end-to-end framework

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:02.417014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:59.152808Z digest=sha256:6923e1a170a5d94c1c69682bc556253c76f5449460e6078fd4b0883b7bf16085

Observation 7cbfadaf-943e-4673-8f37-e4d2aad83fb1 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Sigmoid Loss for Language Image Pre-Training

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:59.342925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:59.342925Z digest=sha256:d1c0b0a7c2a7435d1c9168e8750e9691d41dc1e16e9de105ece002c20d379373

Observation c448c7b1-0717-4d7e-baff-928868910d1e · outbound

This paper cites Point Cloud Mamba: Point Cloud Learning via State Space Model.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Point Cloud Mamba: Point Cloud Learning via State Space Model

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:59.480981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:59.480981Z digest=sha256:4c501f65517f7e891c86ec2e7ff59d269302a6cf017a7e5698b3e99dcf304700

Observation 613a4a46-2840-47e1-a6a4-eb0402e8d530 · outbound

This paper cites G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition G4Seg: Generation for Inexact Segmentation Refinement with Diffusion Models

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:59.564079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:59.564079Z digest=sha256:1e37ebd36f6f3cf671762e0ffd23402e2fe81da42cb9357fddef660db3d2e3b8

Observation c101d2c5-237a-4331-8e65-f43b9565ab7e · outbound

This paper cites Vidtr: Video transformer without convo- lutions.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Vidtr: Video transformer without convo- lutions

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:02.118114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:59.662077Z digest=sha256:9c34c5baa314fe186e7692ad078347b1e7a1eccab624cd109043b8d94ffa89d3

Observation 8014a4aa-3a2c-4def-8347-74f7b1beac90 · outbound

This paper cites Regionclip: Region-based language-image pretraining.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Regionclip: Region-based language-image pretraining

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:01.837632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:59.733742Z digest=sha256:29fa8549681be11267e123903ab9e1f144a57e1a8a68eeaf9da08d3a23b458fd

Observation 6deeac2c-3a55-4d5d-8cb7-97eb4bc75e57 · outbound

This paper cites Graph-based high-order relation modeling for long-term action recognition.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Graph-based high-order relation modeling for long-term action recognition

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:52:01.571641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:51:59.857627Z digest=sha256:9c761ebdc38f01a2b1d3057d32854d2a168928395ef2a73026809f264901513c

Observation ce255339-f313-4b06-8e78-b21533f4e672 · outbound

This paper cites Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model.

MoMa: Modulating Mamba for Adapting Image Foundation Models to Video Recognition Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:59.930418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:59.930418Z digest=sha256:0c6695057c51373ca12c7d82463285da5f1f2981ff40096f9d8ccfb55470e23e

Pith citing papers

No inbound Pith citation observations are available.