Pith. sign in

Paper Citation Record · LEDGER

Time-Scaling State-Space Models for Dense Video Captioning

As of 21 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2509.03426.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.03426 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T10:59:05.327148Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3ef1adcf-14c5-4b03-a712-e911b72d70ff · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Time-Scaling State-Space Models for Dense Video Captioning Flamingo: a visual language model for few-shot learning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:06.238261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.046435Z digest=sha256:54068a23d7b4c869aa71ce039e1876a4fac3d047630bc20d72903db6bfa7ab31

Observation e15d9f35-d9c6-4729-b85c-6ad279a9fa2b · outbound

This paper cites Vivit: A video vision transformer.

Time-Scaling State-Space Models for Dense Video Captioning Vivit: A video vision transformer

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.058702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.058702Z digest=sha256:cfa166ceed52d62d0d822a6859d93b69e28e2e7759976513f48d5483f20cd356

Observation 260eb709-a8c2-4dc1-bebb-556f5fdc0e06 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments.

Time-Scaling State-Space Models for Dense Video Captioning Meteor: An automatic metric for mt evaluation with improved correlation with human judgments

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:06.205518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.064648Z digest=sha256:f62c5a5a666d892c371eaae0f1e130b29b0159a141060b934a0afc479e113bdf

Observation 79b155be-b617-4707-b39a-7e7b2f728240 · outbound

This paper cites Is space-time attention all you need for video understanding? 2021.

Time-Scaling State-Space Models for Dense Video Captioning Is space-time attention all you need for video understanding? 2021

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:06.184422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.070716Z digest=sha256:d9e3d6f317cc5088a98e29d3141a4dc38d0c12e14ffef51ba53485a5ce960b34

Observation d67e6e22-0bf9-4486-9277-49357766d9a5 · outbound

This paper cites Hierarchical State Space Models for Continuous Sequence-to-Sequence Modeling.

Time-Scaling State-Space Models for Dense Video Captioning Hierarchical State Space Models for Continuous Sequence-to-Sequence Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.075845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.075845Z digest=sha256:26ba069cfc7f6c53713685c11c42831f4a1a6292dc7cba458ee46573dd4b40ec

Observation b0e15e12-3a49-4a32-889a-51c7eb382f47 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Time-Scaling State-Space Models for Dense Video Captioning Quo vadis, action recognition? a new model and the kinetics dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.081046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.081046Z digest=sha256:a521df07194dd1e32640e8e01fe3f95a02d53d156d67979767ffe43b56d7e8cd

Observation 546ba720-7323-449a-8a91-93e2b9098c1d · outbound

This paper cites Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding.

Time-Scaling State-Space Models for Dense Video Captioning Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.085854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.085854Z digest=sha256:05efe565d9c945af7c53db51c631bddbdc1e48e3e27e6c67265dca22d5bced12

Observation 2a4c9b6d-09f8-4b7d-8e4c-6f4da2e4900b · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

Time-Scaling State-Space Models for Dense Video Captioning PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.091693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.091693Z digest=sha256:f928b21c311d09f92ea90cd052495d301a13708d8ddbf03f10d241e8a5f106b4

Observation 8ee1896d-7be7-494f-9607-160ef8634304 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Time-Scaling State-Space Models for Dense Video Captioning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.096843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.096843Z digest=sha256:ef2d6fd9eb8017453d738387e19997607c03fe0f99d39ec551d7d24ca13fbdc3

Observation 22a43343-5c80-46be-b353-e8b2756ea021 · outbound

This paper cites Multiscale Vision Transformers.

Time-Scaling State-Space Models for Dense Video Captioning Multiscale Vision Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.102006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.102006Z digest=sha256:f5442f5a93643617b52f65d24cbb228f57ef6a2cc7499a29dfb583121bb8d5cf

Observation 763d8db4-5f01-442a-b3cb-73fac83ac806 · outbound

This paper cites Soda: Story oriented dense video captioning evaluation framework.

Time-Scaling State-Space Models for Dense Video Captioning Soda: Story oriented dense video captioning evaluation framework

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:06.158015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.107388Z digest=sha256:c23718103e7a9027c8d4cae62c5f11dbd229b7deab14e2c9bd7fde32a0013f86

Observation 52a83bc8-c1c2-4693-87f4-c5e46e46ec73 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

Time-Scaling State-Space Models for Dense Video Captioning Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.112247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.112247Z digest=sha256:81b96851589180449ac337219427f5e71a1431d75c5e237919a07733cf85f7ae

Observation a4b37c31-0b7d-450d-b761-0c4bacff8110 · outbound

This paper cites Combining recurrent, convolutional, and continuous-time models with the structured learnable linear state space layer.

Time-Scaling State-Space Models for Dense Video Captioning Combining recurrent, convolutional, and continuous-time models with the structured learnable linear state space layer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:06.142564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.117319Z digest=sha256:40050aaa78244fda9b9144eb4f2ca7ae726f465247b41c5e484e625d9d0f427e

Observation 1d47f72a-7a7a-471c-95d9-65e23fb11b7a · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

Time-Scaling State-Space Models for Dense Video Captioning Efficiently modeling long sequences with structured state spaces

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:06.125885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.122127Z digest=sha256:9291b67b862c7b560c4b8225c283f542d0e40a1a576398499ae763e45ecf2726

Observation 02eb2a0b-fa0e-4a97-89b8-d43a88dba24e · outbound

This paper cites On the parameterization and initialization of diagonal state space models.

Time-Scaling State-Space Models for Dense Video Captioning On the parameterization and initialization of diagonal state space models

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:06.109945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.127507Z digest=sha256:f60bf3a52257cb1a336fdc3d0e53be5a3b3ead63e3827c045ba91f656250ce00

Observation a96d9ce7-a49e-49c9-a47c-9a59b83bc134 · outbound

This paper cites VideoLLM: Modeling Video Sequence with Large Language Models.

Time-Scaling State-Space Models for Dense Video Captioning VideoLLM: Modeling Video Sequence with Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.132385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.132385Z digest=sha256:f1068f796d9a8c94f346c2607f2d0dcd09a21b709dc33aaa1bfa62e4a7be01ef

Observation 6d1310a3-123d-4a09-9dc4-db1c334cfdc1 · outbound

This paper cites Towards Evaluating the Robustness of Visual State Space Models.

Time-Scaling State-Space Models for Dense Video Captioning Towards Evaluating the Robustness of Visual State Space Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.137814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.137814Z digest=sha256:1cc3abc485d39c3270706d85c2c3c3fd2f40b95e094a7cd4a9d9124bde419788

Observation f39f5100-1a88-40ec-81ab-444754944496 · outbound

This paper cites Ac- tivitynet: A large-scale video benchmark for human activity understanding.

Time-Scaling State-Space Models for Dense Video Captioning Ac- tivitynet: A large-scale video benchmark for human activity understanding

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:06.091962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.143278Z digest=sha256:e76c0dfc8cd86f2f6b7670e4b61dd579d3a92a4133b39fc51534bbef79afc6b4

Observation b9de3557-d17c-4b82-b232-4570f09191b1 · outbound

This paper cites Multimodal pretraining for dense video captioning.

Time-Scaling State-Space Models for Dense Video Captioning Multimodal pretraining for dense video captioning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:06.071800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.148295Z digest=sha256:ebc90aa291e5fe474aa14876368339c0fb4ed7e80d2f6f5a0cc1fb5e855fe302

Observation c800fdae-84fe-43ce-9c00-b1591e5e755a · outbound

This paper cites A better use of audio-visual cues: Dense video cap- tioning with bi-modal transformer.

Time-Scaling State-Space Models for Dense Video Captioning A better use of audio-visual cues: Dense video cap- tioning with bi-modal transformer

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:06.055720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.153609Z digest=sha256:70bbb5908656f48c0e1a58d373928b31dc9e43aafaf0ca7497baf44c1059aaf3

Observation 3d29b200-f18b-47e7-aaf1-1ebca35ebf60 · outbound

This paper cites Long movie clip classification with state- space video model.

Time-Scaling State-Space Models for Dense Video Captioning Long movie clip classification with state- space video model

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:06.039558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.158954Z digest=sha256:23dc07f17061ce71ab83e5dfbbec301f20522134193c2930001cc785a05d3a42

Observation b839122e-ee85-4824-b877-f48905c8d382 · outbound

This paper cites Long movie clip classification with state-space video model.

Time-Scaling State-Space Models for Dense Video Captioning Long movie clip classification with state-space video model

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:06.023782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.164858Z digest=sha256:f74e768e6b2f3ca408319e20b37cda897f18705c4aab8d73d54a33caf887226a

Observation 21afb9be-2d44-4e7e-a02d-34f394414243 · outbound

This paper cites Simplified State Space Layers for Sequence Modeling.

Time-Scaling State-Space Models for Dense Video Captioning Simplified State Space Layers for Sequence Modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.169637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.169637Z digest=sha256:c68ce6441c33faefd57ffc33a9ff0592f3a99426abfcbfe1675367dd110c6b97

Observation db29ba23-07d7-49f1-bf0e-4250ac02d87f · outbound

This paper cites MaMMUT: A simple architecture for joint learning for multimodal tasks.

Time-Scaling State-Space Models for Dense Video Captioning MaMMUT: A simple architecture for joint learning for multimodal tasks

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:06.008388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.174622Z digest=sha256:66c8ab03e13326db7a1d0f3a537863160aadf7f50c17741f24053e1974294ff0

Observation 8b02b6aa-54c2-4cd9-9c75-b4f3705b80aa · outbound

This paper cites VideoMamba: State Space Model for Efficient Video Understanding.

Time-Scaling State-Space Models for Dense Video Captioning VideoMamba: State Space Model for Efficient Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.179915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.179915Z digest=sha256:734ea04d2733466a10d1877a9d42abab72aed516748360363e71c2bb0db8b4bc

Observation 8ab8b26f-5d87-41fa-b6bf-254942890bf9 · outbound

This paper cites Video swin transformer.

Time-Scaling State-Space Models for Dense Video Captioning Video swin transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.185580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.185580Z digest=sha256:738f47d40bf4bb05cd2d0577f7a030784cc4cf2979d2e1dd666dee72b810cbdd

Observation 25b9ce09-b7b6-4c85-9f2a-646781b75f3b · outbound

This paper cites A convnet for the 2020s.

Time-Scaling State-Space Models for Dense Video Captioning A convnet for the 2020s

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.981270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.190421Z digest=sha256:ec2b8f98eb0cb1a12088072313c7d71de32ebe2ef83e2de97b7b71de0b9a6cf1

Observation 32d52451-a406-4fe0-b256-332c2401d76f · outbound

This paper cites Downs, Preey Shah, Tri Dao, Stephen A.

Time-Scaling State-Space Models for Dense Video Captioning Downs, Preey Shah, Tri Dao, Stephen A

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.966685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.195376Z digest=sha256:78e85e11c98c50855a565efa901bd21933c2f76e7c98f1c0174a1d196d65df28

Observation 27d19b19-9f62-493a-8a90-c85f26493ea3 · outbound

This paper cites SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces.

Time-Scaling State-Space Models for Dense Video Captioning SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.201010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.201010Z digest=sha256:a7f22fca65662133fea0e8c56a2fa214826e4b865627f24ff915d53a0377fc0e

Observation aa2611a9-9b26-4ddb-8bc8-7cab31ca8ec6 · outbound

This paper cites VideoMamba: Spatio-Temporal Selective State Space Model.

Time-Scaling State-Space Models for Dense Video Captioning VideoMamba: Spatio-Temporal Selective State Space Model

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-05T10:59:05.426653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.206923Z digest=sha256:46d67a9e2cf09dbf4cdb3a6b15386837e8742bf9e5f60d2389c81c475ae6c368

Observation d1096f00-3e6c-4016-b9ab-16b2e3d017ec · outbound

This paper cites Rethinking video vits: Sparse video tubes for joint image and video learning.

Time-Scaling State-Space Models for Dense Video Captioning Rethinking video vits: Sparse video tubes for joint image and video learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.951879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.213214Z digest=sha256:fd2820f55194cbf6b56265c54cdddaf0dfb058bf622533fffab3c798420b2beb

Observation 6666fc41-e35c-43e4-a9e6-225d4f339952 · outbound

This paper cites Dynamic pretraining of vision-language models.

Time-Scaling State-Space Models for Dense Video Captioning Dynamic pretraining of vision-language models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.936144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.218342Z digest=sha256:2315a8063fcdb2adea4961253fd08a2ec4c78cdbc05732ea11d3225441d908b3

Observation c7bfcf73-dd47-4dc2-a241-e88a69dfca1b · outbound

This paper cites Mirasol3B: A multimodal autoregressive model for time-aligned and con- textual modalities.

Time-Scaling State-Space Models for Dense Video Captioning Mirasol3B: A multimodal autoregressive model for time-aligned and con- textual modalities

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.920264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.224715Z digest=sha256:044a7e39cf9def3476f556f41eed4523eb17346aa8f4b4da1b0387d12e17f506

Observation 05b6e8c7-1d46-4af4-81c0-2c282d6d95d4 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding.

Time-Scaling State-Space Models for Dense Video Captioning Timechat: A time-sensitive multimodal large language model for long video understanding

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.903696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.230302Z digest=sha256:cf5b23f6e9f0b5e7e87e91c9324fce935849ae5d0aaf404e40a323965727c521

Observation b33834a9-be83-4272-b023-ffe06f022bb8 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Time-Scaling State-Space Models for Dense Video Captioning Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.888158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.235876Z digest=sha256:a7527af548d9db72abe92017dfea590696a11b66bb894714ff388ba5d8ce75b6

Observation fdcfeb93-88ae-4938-86a0-319a6302e0b0 · outbound

This paper cites Cider: Consensus- based image description evaluation.

Time-Scaling State-Space Models for Dense Video Captioning Cider: Consensus- based image description evaluation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.873713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.241325Z digest=sha256:873ea1acd7eb3306ffa676ad13ca17fb4fa507cbb3206fe65930ed5573a4fe61

Observation dfae2cd1-8124-40fb-8e84-203c0db2fe08 · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

Time-Scaling State-Space Models for Dense Video Captioning GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.246519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.246519Z digest=sha256:616e1c6048f92a80c461a9e06ef48ccc0742204e00e229be99f38440cb0ba17b

Observation 89ad7e88-ac55-45b2-b712-7ea2327f3ac8 · outbound

This paper cites Bidirectional attentive fusion with context gating for dense video captioning.

Time-Scaling State-Space Models for Dense Video Captioning Bidirectional attentive fusion with context gating for dense video captioning

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.857024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.251579Z digest=sha256:b1555e37d672eef314e0c7904f0928b892eefcbc1f06f3edac7ce832f55d3790

Observation b004eba0-e72e-4671-a7d2-ea3858afc324 · outbound

This paper cites Selective structured state-spaces for long-form video understanding.

Time-Scaling State-Space Models for Dense Video Captioning Selective structured state-spaces for long-form video understanding

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.840726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.256934Z digest=sha256:4464b43e1abc3f19477050a9fab9df506b791f1553c2d4355b352542e0d26aad

Observation 79fc654f-ca8e-4874-b01e-f665e61a48b2 · outbound

This paper cites Omnivid: A generative framework for universal video understanding.

Time-Scaling State-Space Models for Dense Video Captioning Omnivid: A generative framework for universal video understanding

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.823966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.261979Z digest=sha256:dc652c03125b9f0311c06ae44cae1106dcadf0aa0435b0fef5550523a96fd094

Observation 0c3f3c76-adad-4ea5-945d-c3afc4e5e9d4 · outbound

This paper cites End- to-end dense video captioning with parallel decoding.

Time-Scaling State-Space Models for Dense Video Captioning End- to-end dense video captioning with parallel decoding

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.807389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.266645Z digest=sha256:d57e4ea4d038fc3d3c36ba42945938dd390fd07efd264eb5bed098c5092cd5b3

Observation 0742e15a-10e6-476a-8cf5-6b6c38583dd5 · outbound

This paper cites The Illusion of State in State-Space Models.

Time-Scaling State-Space Models for Dense Video Captioning The Illusion of State in State-Space Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.271932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.271932Z digest=sha256:8c275738149cdd1133053b7bbc1c97ccaf81e108062c74ca78aa228cfffef870

Observation ab7d1216-38f4-44f8-bfac-296252236029 · outbound

This paper cites Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement.

Time-Scaling State-Space Models for Dense Video Captioning Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.791726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.276722Z digest=sha256:3481cb074a8176361b0c74fb2a33b65be01cfb74e942bab12752e23d51010849

Observation 66b8ffce-91e3-4b18-9aec-5375f4ffe0c6 · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning.

Time-Scaling State-Space Models for Dense Video Captioning Vid2seq: Large-scale pretraining of a visual language model for dense video captioning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.775654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.281749Z digest=sha256:84f73f4676aa35d4ae5d7be264d66b084341fc1ecf88ff5befd5ac3e63453552

Observation 095ad276-acb1-4764-8beb-0644dc2bde8a · outbound

This paper cites Hierarchical video-moment retrieval and step-captioning.

Time-Scaling State-Space Models for Dense Video Captioning Hierarchical video-moment retrieval and step-captioning

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.759187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.286212Z digest=sha256:cca0362ffbdd993ed017deea38fcd0c01b5cbb93726317d557159a9f34665990

Observation e3b22238-bb17-499a-a7e7-883e5526a1ad · outbound

This paper cites Merlot: Multimodal neural script knowledge models.

Time-Scaling State-Space Models for Dense Video Captioning Merlot: Multimodal neural script knowledge models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.743057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.291033Z digest=sha256:456078cdf7a62ef5f9864da54a608b6f44f493e465a9033f892d9fa077447a85

Observation cce30a80-326c-41e7-8b40-7bd4addec768 · outbound

This paper cites Unifying event detection and captioning as sequence generation via pre-training.

Time-Scaling State-Space Models for Dense Video Captioning Unifying event detection and captioning as sequence generation via pre-training

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.724265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.295700Z digest=sha256:7cef84eb138afd1aa950770661d3d1b2f544daa3cacc4793adc85193ad866c28

Observation 010fbf44-0d72-4c54-81d8-e80060f91912 · outbound

This paper cites Towards automatic learning of pro- cedures from web instructional videos.

Time-Scaling State-Space Models for Dense Video Captioning Towards automatic learning of pro- cedures from web instructional videos

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.705813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.300370Z digest=sha256:d25aad5b212900bce7b883749f3374ea609e3200bdbb1a2ee9db9a72d079f04d

Observation 34ea5893-63ca-41e8-be67-a2fdbafbdcd3 · outbound

This paper cites End- to-end dense video captioning with masked transformer.

Time-Scaling State-Space Models for Dense Video Captioning End- to-end dense video captioning with masked transformer

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.690821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.305534Z digest=sha256:7c524185369b899de66c34af313adb81ddefa9f48ba23f9585b384dda64c0607

Observation fcecb641-1158-4264-bf34-bea5569e83bb · outbound

This paper cites Streaming dense video captioning.

Time-Scaling State-Space Models for Dense Video Captioning Streaming dense video captioning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.673618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.310438Z digest=sha256:cf3a0e86fbc060680eb2aa9608aec57f8bd506cd8304459c130a8292301a6eb2

Observation 7dee1653-5de3-427f-884a-53f77c052c4e · outbound

This paper cites Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model.

Time-Scaling State-Space Models for Dense Video Captioning Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.316796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.316796Z digest=sha256:8d9539db1c6af9ec25516d179884409ace1fe6e249b9bb2dcc2f433bd1fe006a

Observation ebbb1fd9-7e14-4c77-a4ae-e0ad3e36a2bd · outbound

This paper cites Thapliyal, William Yang Wang, and Radu Soricut.

Time-Scaling State-Space Models for Dense Video Captioning Thapliyal, William Yang Wang, and Radu Soricut

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.655316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.322511Z digest=sha256:3bc9605f2f7dc23a4d3782186260f376ace7ba56d6d7743a720522017dad07f6

Observation f97a4890-fbba-44b6-b68c-42ff01c071f5 · outbound

This paper cites State space models for event cameras.

Time-Scaling State-Space Models for Dense Video Captioning State space models for event cameras

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T10:59:05.637715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-05T10:59:05.327148Z digest=sha256:934e869d07d17d36d2dfa3d84590acf884ed6de04354390b16617cbbf7e7012a

Observation 97c46b42-392a-4835-a400-0d62648a5eae · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Time-Scaling State-Space Models for Dense Video Captioning Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T10:59:05.052003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:59:05.052003Z digest=sha256:ddd97610c9ce7b3d55a92646923b739e2cf61e7dd0a0eef88c4b852d716db524

Pith citing papers

No inbound Pith citation observations are available.