Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:59:05.327148Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2509.03426.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T10:59:05.327148Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3ef1adcf-14c5-4b03-a712-e911b72d70ff · outbound
Time-Scaling State-Space Models for Dense Video Captioning Flamingo: a visual language model for few-shot learning,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e15d9f35-d9c6-4729-b85c-6ad279a9fa2b · outbound
Time-Scaling State-Space Models for Dense Video Captioning Vivit: A video vision transformer
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 260eb709-a8c2-4dc1-bebb-556f5fdc0e06 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 79b155be-b617-4707-b39a-7e7b2f728240 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Is space-time attention all you need for video understanding? 2021
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d67e6e22-0bf9-4486-9277-49357766d9a5 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Hierarchical State Space Models for Continuous Sequence-to-Sequence Modeling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0e15e12-3a49-4a32-889a-51c7eb382f47 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Quo vadis, action recognition? a new model and the kinetics dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 546ba720-7323-449a-8a91-93e2b9098c1d · outbound
Time-Scaling State-Space Models for Dense Video Captioning Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a4c9b6d-09f8-4b7d-8e4c-6f4da2e4900b · outbound
Time-Scaling State-Space Models for Dense Video Captioning PaLI-X: On Scaling up a Multilingual Vision and Language Model
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ee1896d-7be7-494f-9607-160ef8634304 · outbound
Time-Scaling State-Space Models for Dense Video Captioning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22a43343-5c80-46be-b353-e8b2756ea021 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Multiscale Vision Transformers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 763d8db4-5f01-442a-b3cb-73fac83ac806 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Soda: Story oriented dense video captioning evaluation framework
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 52a83bc8-c1c2-4693-87f4-c5e46e46ec73 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4b37c31-0b7d-450d-b761-0c4bacff8110 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Combining recurrent, convolutional, and continuous-time models with the structured learnable linear state space layer
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1d47f72a-7a7a-471c-95d9-65e23fb11b7a · outbound
Time-Scaling State-Space Models for Dense Video Captioning Efficiently modeling long sequences with structured state spaces
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 02eb2a0b-fa0e-4a97-89b8-d43a88dba24e · outbound
Time-Scaling State-Space Models for Dense Video Captioning On the parameterization and initialization of diagonal state space models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a96d9ce7-a49e-49c9-a47c-9a59b83bc134 · outbound
Time-Scaling State-Space Models for Dense Video Captioning VideoLLM: Modeling Video Sequence with Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d1310a3-123d-4a09-9dc4-db1c334cfdc1 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Towards Evaluating the Robustness of Visual State Space Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f39f5100-1a88-40ec-81ab-444754944496 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Ac- tivitynet: A large-scale video benchmark for human activity understanding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b9de3557-d17c-4b82-b232-4570f09191b1 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Multimodal pretraining for dense video captioning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c800fdae-84fe-43ce-9c00-b1591e5e755a · outbound
Time-Scaling State-Space Models for Dense Video Captioning A better use of audio-visual cues: Dense video cap- tioning with bi-modal transformer
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3d29b200-f18b-47e7-aaf1-1ebca35ebf60 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Long movie clip classification with state- space video model
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b839122e-ee85-4824-b877-f48905c8d382 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Long movie clip classification with state-space video model
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 21afb9be-2d44-4e7e-a02d-34f394414243 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Simplified State Space Layers for Sequence Modeling
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db29ba23-07d7-49f1-bf0e-4250ac02d87f · outbound
Time-Scaling State-Space Models for Dense Video Captioning MaMMUT: A simple architecture for joint learning for multimodal tasks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 8b02b6aa-54c2-4cd9-9c75-b4f3705b80aa · outbound
Time-Scaling State-Space Models for Dense Video Captioning VideoMamba: State Space Model for Efficient Video Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab8b26f-5d87-41fa-b6bf-254942890bf9 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Video swin transformer
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b9ce09-b7b6-4c85-9f2a-646781b75f3b · outbound
Time-Scaling State-Space Models for Dense Video Captioning A convnet for the 2020s
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 32d52451-a406-4fe0-b256-332c2401d76f · outbound
Time-Scaling State-Space Models for Dense Video Captioning Downs, Preey Shah, Tri Dao, Stephen A
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 27d19b19-9f62-493a-8a90-c85f26493ea3 · outbound
Time-Scaling State-Space Models for Dense Video Captioning SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa2611a9-9b26-4ddb-8bc8-7cab31ca8ec6 · outbound
Time-Scaling State-Space Models for Dense Video Captioning VideoMamba: Spatio-Temporal Selective State Space Model
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d1096f00-3e6c-4016-b9ab-16b2e3d017ec · outbound
Time-Scaling State-Space Models for Dense Video Captioning Rethinking video vits: Sparse video tubes for joint image and video learning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 6666fc41-e35c-43e4-a9e6-225d4f339952 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Dynamic pretraining of vision-language models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c7bfcf73-dd47-4dc2-a241-e88a69dfca1b · outbound
Time-Scaling State-Space Models for Dense Video Captioning Mirasol3B: A multimodal autoregressive model for time-aligned and con- textual modalities
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 05b6e8c7-1d46-4af4-81c0-2c282d6d95d4 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Timechat: A time-sensitive multimodal large language model for long video understanding
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b33834a9-be83-4272-b023-ffe06f022bb8 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fdcfeb93-88ae-4938-86a0-319a6302e0b0 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Cider: Consensus- based image description evaluation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation dfae2cd1-8124-40fb-8e84-203c0db2fe08 · outbound
Time-Scaling State-Space Models for Dense Video Captioning GIT: A Generative Image-to-text Transformer for Vision and Language
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89ad7e88-ac55-45b2-b712-7ea2327f3ac8 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Bidirectional attentive fusion with context gating for dense video captioning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b004eba0-e72e-4671-a7d2-ea3858afc324 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Selective structured state-spaces for long-form video understanding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 79fc654f-ca8e-4874-b01e-f665e61a48b2 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Omnivid: A generative framework for universal video understanding
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0c3f3c76-adad-4ea5-945d-c3afc4e5e9d4 · outbound
Time-Scaling State-Space Models for Dense Video Captioning End- to-end dense video captioning with parallel decoding
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0742e15a-10e6-476a-8cf5-6b6c38583dd5 · outbound
Time-Scaling State-Space Models for Dense Video Captioning The Illusion of State in State-Space Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab7d1216-38f4-44f8-bfac-296252236029 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 66b8ffce-91e3-4b18-9aec-5375f4ffe0c6 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Vid2seq: Large-scale pretraining of a visual language model for dense video captioning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 095ad276-acb1-4764-8beb-0644dc2bde8a · outbound
Time-Scaling State-Space Models for Dense Video Captioning Hierarchical video-moment retrieval and step-captioning
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e3b22238-bb17-499a-a7e7-883e5526a1ad · outbound
Time-Scaling State-Space Models for Dense Video Captioning Merlot: Multimodal neural script knowledge models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation cce30a80-326c-41e7-8b40-7bd4addec768 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Unifying event detection and captioning as sequence generation via pre-training
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 010fbf44-0d72-4c54-81d8-e80060f91912 · outbound
Time-Scaling State-Space Models for Dense Video Captioning Towards automatic learning of pro- cedures from web instructional videos
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 34ea5893-63ca-41e8-be67-a2fdbafbdcd3 · outbound
Time-Scaling State-Space Models for Dense Video Captioning End- to-end dense video captioning with masked transformer
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fcecb641-1158-4264-bf34-bea5569e83bb · outbound
Time-Scaling State-Space Models for Dense Video Captioning Streaming dense video captioning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 7dee1653-5de3-427f-884a-53f77c052c4e · outbound
Time-Scaling State-Space Models for Dense Video Captioning Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebbb1fd9-7e14-4c77-a4ae-e0ad3e36a2bd · outbound
Time-Scaling State-Space Models for Dense Video Captioning Thapliyal, William Yang Wang, and Radu Soricut
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f97a4890-fbba-44b6-b68c-42ff01c071f5 · outbound
Time-Scaling State-Space Models for Dense Video Captioning State space models for event cameras
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 97c46b42-392a-4835-a400-0d62648a5eae · outbound
Time-Scaling State-Space Models for Dense Video Captioning Flamingo: a Visual Language Model for Few-Shot Learning
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.