Pith. sign in

Paper Citation Record · LEDGER

VideoMamba: State Space Model for Efficient Video Understanding

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2403.06977.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.06977 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:00:30.625789Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:30:07.123917Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation fc3e6eb1-35bb-48ac-b4bf-c0883b8af3c4 · inbound

A Survey of Mamba cites this paper.

A Survey of Mamba VideoMamba: State Space Model for Efficient Video Understanding

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:13:30.939914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-23T22:09:19.917854Z digest=sha256:79acae083510f62b4c9787fda831e2a8435ef9461f925c79680a55d51ae2868d

Observation 25ac5823-7017-48d0-a5aa-8f0bdc729f79 · inbound

MambaLCT: Boosting Tracking via Long-term Context State Space Model cites this paper.

MambaLCT: Boosting Tracking via Long-term Context State Space Model VideoMamba: State Space Model for Efficient Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T13:00:30.625789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:00:30.625789Z digest=sha256:26d9bb066c11c8a512617e970a5b69ed81da22ab4e6f9791ba81dc5b508adc15

Observation 7d18cdbb-f0af-447a-9a99-39726de1a73f · inbound

Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking cites this paper.

Exploiting Multimodal Spatial-temporal Patterns for Video Object Tracking VideoMamba: State Space Model for Efficient Video Understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:12:52.333311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:12:52.333311Z digest=sha256:1ea3353f163e757137cfbdca0783c9f026a6275f3469a4e737e09fdf60e299b5

Observation f8ce410b-9b52-42f1-bfad-48c2aae2d9c7 · inbound

V"Mean"ba: Visual State Space Models only need 1 hidden dimension cites this paper.

V"Mean"ba: Visual State Space Models only need 1 hidden dimension VideoMamba: State Space Model for Efficient Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T10:29:37.003688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T10:29:37.003688Z digest=sha256:0896019adcca7e027da5171ae69de96c919601d3484c811c99a0b94bb951211f

Observation c574a8b3-4361-48b2-8816-b02ceddbf2c2 · inbound

MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training Smoothing cites this paper.

MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training Smoothing VideoMamba: State Space Model for Efficient Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T23:40:07.883499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:40:07.883499Z digest=sha256:68a68e4f2f90eaec3aad9f9ad6fad6d869f041c61ffa870d61a3c949eedc935f

Observation 9521212c-fc34-4fb2-979e-440ae27c7e8f · inbound

H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving cites this paper.

H-MBA: Hierarchical MamBa Adaptation for Multi-Modal Video Understanding in Autonomous Driving VideoMamba: State Space Model for Efficient Video Understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:11.569111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:42:11.569111Z digest=sha256:0226d46914b4dc80f84365c04021d10a619187da740e5bfbdf34e439bc3cc290

Observation 2c0f48e6-09af-4641-b416-3321b1f3bd60 · inbound

AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation cites this paper.

AVS-Mamba: Exploring Temporal and Multi-modal Mamba for Audio-Visual Segmentation VideoMamba: State Space Model for Efficient Video Understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-10T20:38:56.483125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:38:56.483125Z digest=sha256:f61027d6e64aea8b6870b87d3a2ad035fee323954ec617909d026b7579cdb8b7

Observation de4aa581-649f-4391-b74f-f759561fe1ec · inbound

MV-GMN: State Space Model for Multi-View Action Recognition cites this paper.

MV-GMN: State Space Model for Multi-View Action Recognition VideoMamba: State Space Model for Efficient Video Understanding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T15:37:30.555080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:37:30.555080Z digest=sha256:2318f9ba53d99a1e4bd0f1eaee01f2ba2b6458105711a1862c54bb23cf51e20e

Observation 3faf318b-c026-4671-8a37-75a93ea98b8e · inbound

Sparsified State-Space Models are Efficient Highway Networks cites this paper.

Sparsified State-Space Models are Efficient Highway Networks VideoMamba: State Space Model for Efficient Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:54:45.758010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:54:45.758010Z digest=sha256:2457be8391c40123be5e8104a8b506ffed67c8abdd7bbd95a0e5e54013534079

Observation d2465311-0db7-4e7e-8740-81f8cbadd9eb · inbound

Mamba Drafters for Speculative Decoding cites this paper.

Mamba Drafters for Speculative Decoding VideoMamba: State Space Model for Efficient Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.919222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:12.919222Z digest=sha256:bef3961027021a663631f68965c42fde3a3120c379dcfed7c00b36949cced3c2

Observation 03a5c6c0-ad48-4a67-be9d-51d07e94ec7f · inbound

DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos cites this paper.

DySS: Dynamic Queries and State-Space Learning for Efficient 3D Object Detection from Multi-Camera Videos VideoMamba: State Space Model for Efficient Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:23.682312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:23.682312Z digest=sha256:d6b003e05aab206d3170a2eda673b023c947e90ba83e77e125e330c80b75bf12

Observation 60fd24f9-7f8e-4d8d-bade-a4d8a55cc9a8 · inbound

Comparing Learning Paradigms for Egocentric Video Summarization cites this paper.

Comparing Learning Paradigms for Egocentric Video Summarization VideoMamba: State Space Model for Efficient Video Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:41.396361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:41.396361Z digest=sha256:7a45cc08d11bb852f6af9df9f0f86d2b70a5d41cc3995598c45baf2c7c6b6960

Observation 421ad7d2-17c9-475e-b20f-1894381d7cff · inbound

QuarterMap: Efficient Post-Training Token Pruning for Visual State Space Models cites this paper.

QuarterMap: Efficient Post-Training Token Pruning for Visual State Space Models VideoMamba: State Space Model for Efficient Video Understanding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:59:28.795259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:59:28.795259Z digest=sha256:b576587755e2348c630e055dccf1cab5ca1f8efe90663fbc05065526ea64201d

Observation ad9977a5-0e2f-4aa0-909f-dee0562fc13e · inbound

Few-Shot Object Detection via Spatial-Channel State Space Model cites this paper.

Few-Shot Object Detection via Spatial-Channel State Space Model VideoMamba: State Space Model for Efficient Video Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T15:39:16.242741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:39:16.242741Z digest=sha256:cd743bce8c26c8c372063d9afb6a11a5455a52c61052ffa7caf34a67b68af699

Observation cf99f8eb-e72c-4ddc-9c7e-33afe00406fe · inbound

HydraMamba: Multi-Head State Space Model for Global Point Cloud Learning cites this paper.

HydraMamba: Multi-Head State Space Model for Global Point Cloud Learning VideoMamba: State Space Model for Efficient Video Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T14:04:41.966729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:04:41.966729Z digest=sha256:e0a9051891bbf1dec2101134b729c818a0fb3787c0ca0a49cb91d82407e31392

Observation 8897b41a-c3b7-4c25-9cc7-e293b2f6b83a · inbound

Straightforward Bayesian A/B testing with Dirichlet posteriors cites this paper.

Straightforward Bayesian A/B testing with Dirichlet posteriors VideoMamba: State Space Model for Efficient Video Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T21:43:17.884122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:43:17.884122Z digest=sha256:6c7498168958e4a3fad253459ad5b4e6d6b74ae3a3825dcb1ca1990c800eeac4

Observation c77fc7dd-61f4-410d-aa78-68d13ba08b1e · inbound

Animate-X++: Universal Character Image Animation with Dynamic Backgrounds cites this paper.

Animate-X++: Universal Character Image Animation with Dynamic Backgrounds VideoMamba: State Space Model for Efficient Video Understanding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T21:09:10.359648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:09:10.359648Z digest=sha256:4023260b40aaf3ea91e5a89ae616da76608c78a4c85194e63842bc9699c99f4d

Observation 1a643685-9f82-42b7-a214-1ecb92e22736 · inbound

Hierarchical Spatio-temporal Segmentation Network for Ejection Fraction Estimation in Echocardiography Videos cites this paper.

Hierarchical Spatio-temporal Segmentation Network for Ejection Fraction Estimation in Echocardiography Videos VideoMamba: State Space Model for Efficient Video Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:21:44.196052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:21:44.196052Z digest=sha256:0454454aa75a89f503045bbaf08d18372f004843ae0f9c8adabe1eea01d6ca8d

Observation cbb424ea-83e3-4e0a-95eb-569bb96cbb12 · inbound

Boosting Micro-Expression Analysis via Prior-Guided Video-Level Regression cites this paper.

Boosting Micro-Expression Analysis via Prior-Guided Video-Level Regression VideoMamba: State Space Model for Efficient Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T16:13:33.351148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:13:33.351148Z digest=sha256:967ae7bdffedfd088d8a7927a66fc66c30d8ce0803b1c9a3d59a836727c1e19c

Observation 8b02b6aa-54c2-4cd9-9c75-b4f3705b80aa · inbound

Time-Scaling State-Space Models for Dense Video Captioning cites this paper.

Time-Scaling State-Space Models for Dense Video Captioning VideoMamba: State Space Model for Efficient Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.179915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.179915Z digest=sha256:5c04d51b16345100202fbe274d9ed27ec906b0072d554d43b2aba4e8e2bbfe93

Observation 36de079a-1ff4-4fe6-a456-9424b8091984 · inbound

MambaADv2: Evolving Duality-enhanced State Space Model for Unsupervised Anomaly Detection cites this paper.

MambaADv2: Evolving Duality-enhanced State Space Model for Unsupervised Anomaly Detection VideoMamba: State Space Model for Efficient Video Understanding

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:49:44.636582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T09:26:51.456652Z digest=sha256:099174fe098861c750936c6560accbb3352fca9d373b898ba8aa8340cfab5aab

Observation f26f67de-cf65-4323-9ad8-1fd3fdad76a3 · inbound

Efficient Remote Sensing Instance Segmentation with Linear-Time State Space Distilled Visual Foundation Models cites this paper.

Efficient Remote Sensing Instance Segmentation with Linear-Time State Space Distilled Visual Foundation Models VideoMamba: State Space Model for Efficient Video Understanding

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:30:07.126013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-25T21:20:21.295192Z digest=sha256:cfcfd83c408c056b5647803a0a9105d95fe288743896bd131dbec3adc02f9be5

Observation eec9768f-e5f3-4c09-90a0-481b673ef511 · inbound

Consistent and Editable: A Balanced Framework for Text-Guided Video Editing cites this paper.

Consistent and Editable: A Balanced Framework for Text-Guided Video Editing VideoMamba: State Space Model for Efficient Video Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-11T09:28:19.511968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T09:28:19.511968Z digest=sha256:1393d4bad23abd32f8bedf88eff264e98c89607e78118967b6cca58be0a95632

Observation 4869c109-36eb-443f-a9e6-7d320f6d3e54 · inbound

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding cites this paper.

VideoSEMA: a scalable and efficient Mamba-like attention for video understanding VideoMamba: State Space Model for Efficient Video Understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T01:25:52.300928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:25:52.300928Z digest=sha256:7fecc977d7e63833f2e0998414765965b5a42e6aba18a7b0114e3d228a6fa042