Pith. sign in

Paper Citation Record · LEDGER

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding

As of 21 August 2026, this Paper Citation Record lists 100 of 120 outbound references and 1 inbound Pith citation observation for arXiv:2507.02591.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02591 v3

Coverage vector

measured 100 of 120 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:29:56.537267Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:57:29.171219Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 120 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved97
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fd7260af-fe08-4fda-8399-745342c69b9c · outbound

This paper cites GPT-4 Technical Report.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.170619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.170619Z digest=sha256:b918e30b599a692e02e397360f0ed63e9d7396046522fc7a078d878ada40b53a

Observation 26ac5815-ddda-4ff1-8f70-55b952b16ead · outbound

This paper cites HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.222083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.222083Z digest=sha256:4f7eed3ae4cb04b8d0fa507e6471d8ed02950d87a01193ca8f599be5ed187668

Observation 6e71df41-cddb-4ed0-9cba-c88cf129020e · outbound

This paper cites In- finibench: A comprehensive benchmark for large multi- modal models in very long video understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding In- finibench: A comprehensive benchmark for large multi- modal models in very long video understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.286948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.286948Z digest=sha256:a81c00109d95d52c0d417b4e68499208e547a1fd15c4cd38cb7fcb4e0a71ddd4

Observation 050ede92-a899-4723-9dfa-8c1005e9b05d · outbound

This paper cites Qwen2.5-VL Technical Report.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.428876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.428876Z digest=sha256:7eb66f332d21369d5c0e15bb88501e99ec8952c6a62024c9600273dc95c615d2

Observation 82ac5bc2-d64c-4066-bf61-72049155ccc7 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding PaliGemma: A versatile 3B VLM for transfer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.584546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.584546Z digest=sha256:aaf33c413b0bbf2a781748ee036bb21ccc0e4d4858ac242874a8805d796ade6d

Observation 86aa583b-e38f-4500-9435-3eea1cb6fa26 · outbound

This paper cites VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.711985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.711985Z digest=sha256:16fb6f7c471c244a3c5eef347c3dc137e2fa81e3922ded926d5c6345a3068a00

Observation 400a4ef2-710e-4470-b21a-1d604f15c3cb · outbound

This paper cites Token merging: Your ViT but faster.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Token merging: Your ViT but faster

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.814996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.814996Z digest=sha256:e4b61bcc01fd93d9f6fc0b7c7c49baa744699b24775ac2deca59270e609edcb3

Observation 7fb399f3-d917-44ac-88d5-35187acd98c2 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Activitynet: A large-scale video benchmark for human activity understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.918119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.918119Z digest=sha256:3eb6281eed3ef07095840e927715889b0c9c14339d629f89fb45c8dd486a9b15

Observation 0f3e4acc-c4f0-4034-9e3f-43b4e2628177 · outbound

This paper cites Matryoshka Multimodal Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Matryoshka Multimodal Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.998290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.998290Z digest=sha256:ac1848e1b2a46918a5f4e699682f77719104057208ee5eba58c1e4fed3c4ad23

Observation 6eb9324b-c854-49d8-9a63-c23479ecc6c8 · outbound

This paper cites View transformer layers from online optimization perspective, 2025.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding View transformer layers from online optimization perspective, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.089471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.089471Z digest=sha256:05a27c3ec3a6c288bb9a0e1bee5f6d8dd4fd3cd1c475a48672e717f6e2102e6e

Observation 7b21e852-8f3b-4f2d-9372-9a4dce0f8077 · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.227081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.227081Z digest=sha256:7693998c0650468feb7f02fbeaf2f8981529565de857c4ded4d400f84e75dbc5

Observation 05cfcee0-61aa-4dad-95f3-7633ca240c40 · outbound

This paper cites Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.369805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.369805Z digest=sha256:053008667eba85f833b8b1c8daa518ef3bf69a1f4953a8dce7f1d584e4356654

Observation 67515cb2-5b29-4a06-be65-585b37f24353 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.489965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.489965Z digest=sha256:5caa16e765564b31b1019a842212edb059f733c05aac0e715ec275b78a5bd2fb

Observation eef22555-b744-4038-8e00-8a46b63f3888 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.568606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.568606Z digest=sha256:faff378d0ddae1ddeaa97b44fb8c6454ceaf82c4021bb8182df2d6524fdb4bf1

Observation 54b94232-4936-4c5b-83ae-3360eb5bea48 · outbound

This paper cites MotionLLM: Understanding Human Behaviors from Human Motions and Videos.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.674994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.674994Z digest=sha256:5a9ec0a76478bb74f70fe0dd2f32bf3e4a6445d06d1faf974a8f6ff55504cfa1

Observation 518e2809-35a0-4685-8c4c-de6170ccf924 · outbound

This paper cites Stuffed mamba: State col- lapse and state capacity of rnn-based long-context model- ing.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Stuffed mamba: State col- lapse and state capacity of rnn-based long-context model- ing

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.973755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.973755Z digest=sha256:ecdd7307f22703a00f9df4a7b7bc6f24d0b85a05c39269a7d527b95ee2f80414

Observation 658d614a-3946-4d30-b409-c44a47429924 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.077308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.077308Z digest=sha256:ac863ad7cbec4a0adc57c3e214e310e0b88df67a991cc81cd77bd884e8b69131

Observation 2159393d-359b-4bf4-b3c1-2c44b1b3c8ca · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.151067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.151067Z digest=sha256:cd4198c18340e42ac7e277193d3012d47e7ebb7d7efff97cdafeda0ac829e0ab

Observation 5125826a-e05b-44f4-8e0e-d056a65494fe · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.254227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.254227Z digest=sha256:016c70cc7f300e906017999a898665efa8e89d5a2d46c404054044bc96d9de74

Observation e5c1393e-2194-4f0a-9adc-f040f4e1f40c · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.334244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.334244Z digest=sha256:348cad227b92dfbe5df8849d2c00c8a0944fa717004786fe85eb0e1e8812c5a9

Observation f7e63bfc-07a6-4c67-a0fe-1cbf2929b32e · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.455631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.455631Z digest=sha256:52396dc575df4dea6f4f8149e5f8fdf216312d7d6ab725134f7d7cf88e1e845d

Observation 889cdfca-77ea-444e-be11-5f1d974dba96 · outbound

This paper cites Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.624319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.624319Z digest=sha256:85429e69bc6d4ece451b1718648982c58f96b53969972ed5a915c1d4b2e96dba

Observation d6032bc0-b3ab-48ea-937f-15b7b414e895 · outbound

This paper cites Gate-variants of gated re- current unit (gru) neural networks.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Gate-variants of gated re- current unit (gru) neural networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.743701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.743701Z digest=sha256:64e39c83c6182f642b62df976e565c49f56d7f1d1f2eae4bf713145bab92b563

Observation 864436db-4c6d-4a50-bea4-c18eb3217e64 · outbound

This paper cites Towards Event-oriented Long Video Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Towards Event-oriented Long Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.820033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.820033Z digest=sha256:38761b594618fbdd748e76788e5843abef8d6814d58ffc60b9ebe0e26d203371

Observation 67b408c5-29ca-49df-a29c-9846938969cd · outbound

This paper cites Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.891728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.891728Z digest=sha256:981593ed11d6986cebe3e8d49fbc100c8866947459803cc9a6e166ff3fb95fa4

Observation 7adf30ca-8359-47d1-833d-9416aa5f6e3c · outbound

This paper cites Videoagent: A memory-augmented mul- timodal agent for video understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Videoagent: A memory-augmented mul- timodal agent for video understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.028784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.028784Z digest=sha256:fa98cdcaac6cd217f6f8fb5ad3f8bc246870ad9b33f5e61ed8d724f9c2015fff

Observation cad1da19-64d5-40e3-9867-bd0534405420 · outbound

This paper cites Were RNNs All We Needed?.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Were RNNs All We Needed?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.131873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.131873Z digest=sha256:cacd9294364423a86244f20fc934d6ad75899fcc3449c7a4844f1f99e150b801

Observation 7568dce7-d255-4613-9cc3-9f90767b4b4c · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.229493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.229493Z digest=sha256:14543b63b961e91ff05c7be7b09a480f0787892d0ae218b0250031bd203d4349

Observation bc7552e6-52e4-4ee8-971f-3618e1c7ba2d · outbound

This paper cites Long short-term memory.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Long short-term memory

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.302154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.302154Z digest=sha256:99fa9f9a66149b779836e5cec67aea72ac0023d698d57cc7fb30f00e4f781d08

Observation 77d5682a-b30c-42db-b433-4261c4fdda77 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.386856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.386856Z digest=sha256:3a5fd293dfa6ef47fab5ed99f24d537002bb47e8b7fb97cc062bb36b79c80761

Observation 6f53dc11-111a-4c9e-9ca0-ef9b3f4c1988 · outbound

This paper cites Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.482506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.482506Z digest=sha256:74f3c05c497dc72f9cb35ff7dcba8e033c2b216300bbfe80682951fd01e6015c

Observation fa2f5207-8970-413d-bf4e-4d804d98ed10 · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.583701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.583701Z digest=sha256:6e56cec6f1e8a9279219efcc4a1a76fe2c4eb32e79329f1df10c6872e6739eb8

Observation 5638984e-b43e-4833-a759-52ab312519ab · outbound

This paper cites Masked autoencoders are scal- able vision learners.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Masked autoencoders are scal- able vision learners

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.658548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.658548Z digest=sha256:1336bd8785ead68015ec077ffc23f9890d96be704e39ff29e6a5aadcc6941126

Observation 6fb40779-dca4-40bb-abb8-87b94a84d98a · outbound

This paper cites VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:29:57.988732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:29:50.789507Z digest=sha256:0dea231d5fd461ee5db40edd8800cd57e1fe23feae5ba5ac8f6d020df739486c

Observation 6e4b89be-19af-4698-b465-3a75f94544c0 · outbound

This paper cites Token compensator: Altering in- ference cost of vision transformer without re-tuning.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Token compensator: Altering in- ference cost of vision transformer without re-tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.888495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.888495Z digest=sha256:e58c92ae57580a220ba88991fb75c9468c6e09d5849e6894902aa7bf726fabb5

Observation 65cf98aa-0a57-421f-8ddf-f8a99652eb0a · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.012187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.012187Z digest=sha256:167d64f7d9952985ff48835d23aed4fe833bea168f6ddfe740d6bcee10e15e59

Observation cd13014e-7ead-46af-9005-82a217effe0b · outbound

This paper cites Exploring Enhanced Contextual Information for Video-Level Object Tracking.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Exploring Enhanced Contextual Information for Video-Level Object Tracking

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.094833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.094833Z digest=sha256:b67c0b2e18d974843fb77bce9dab6fbc0071115f10a880a977496e87f272cd63

Observation e660e977-e55a-46bf-b95b-d8a5b4637ed1 · outbound

This paper cites Transformers are rnns: Fast autore- gressive transformers with linear attention.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Transformers are rnns: Fast autore- gressive transformers with linear attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.173839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.173839Z digest=sha256:7a2f84861364c9c90e78205ed8645a9333bc3366ef75baa1f2954edb2e770f54

Observation 7b51d115-5159-4ae8-ad73-38f1c729ed51 · outbound

This paper cites Rethinking Positional Encoding in Language Pre-training.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Rethinking Positional Encoding in Language Pre-training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.253360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.253360Z digest=sha256:e022156706b032e3d5585dee55fce31a443fc239f8432a19a3a010ec7cf9850e

Observation 1859956c-fa26-4f03-acbe-d34f0e8363e9 · outbound

This paper cites Video Token Merging for Long-form Video Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video Token Merging for Long-form Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.350310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.350310Z digest=sha256:fa39884162560d1bc616f941ee86d86a8d46dd4636b5cf9a870dbfa0f169057e

Observation b5b48a0b-b677-4c05-b7c6-cca924d5aa4e · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.447003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.447003Z digest=sha256:f37cce1df952f1d12e33067af0be0ada30ec6e3cc27aed7e5aecabc3e0ef664c

Observation 41e8f07a-f02d-405e-83f9-e0ba620521e1 · outbound

This paper cites Lmms-eval: Accelerating the development of large multimoal models, 2024.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Lmms-eval: Accelerating the development of large multimoal models, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.537802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.537802Z digest=sha256:2039989153233121e4247c75d9b4d1b949f73e5290aef9b14e29635a8f980fef

Observation ce37d5cf-a163-49c2-bf22-184a2702f662 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.630112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.630112Z digest=sha256:7cf1bdb28c23620a80ffa9207bb203bc8fa98207ed0deeba23b4b83d6c98d423

Observation 67cf1236-7fa1-4118-8cbe-6711b0775924 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.704165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.704165Z digest=sha256:27ccfc5a99d138fa50f16a4686250df3f6c473da19b389d4d8e64c69e7b2e506

Observation 96efacf1-46f1-4ad7-8ead-ec7d9589b0bb · outbound

This paper cites Mvbench: A comprehensive multi-modal video under- standing benchmark.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Mvbench: A comprehensive multi-modal video under- standing benchmark

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.787241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.787241Z digest=sha256:e7a5044a149210eeba4d73b608944535cbec23d580c4aaa5a2b73bac5bb7b495

Observation cc3df688-87ae-43be-9cbf-90e6559d12ff · outbound

This paper cites Videomamba: State space model for efficient video understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Videomamba: State space model for efficient video understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.870161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.870161Z digest=sha256:0f934527dde495f5666051aae2eab4ffe0edb72886d0782755db20058275b1d6

Observation eb4f2bdf-9eaf-4eca-9789-49d5d5680fa3 · outbound

This paper cites Independently recurrent neural network (indrnn): Building a longer and deeper rnn.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Independently recurrent neural network (indrnn): Building a longer and deeper rnn

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.962556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.962556Z digest=sha256:b6146ee49dca1ee76486fbfa8df8a2a5e36f227eb7ce14ace10b67b7e9e6dbd5

Observation 4d725fa6-89c5-47d4-8899-7fd448e4b2c5 · outbound

This paper cites Mamba- nd: Selective state space modeling for multi-dimensional data.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Mamba- nd: Selective state space modeling for multi-dimensional data

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.058317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.058317Z digest=sha256:c50136de6eabfffa83b97e0026ceeb94e00b3f72872000c2d23256867cd01526

Observation 2efa0987-9240-45d2-aa06-4a09a803cc06 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.144931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.144931Z digest=sha256:0aaca902324c402b8413294a9e1ddc9cbe82a2929105e72e259b7ab1ec74379d

Observation 427d13d9-fd81-4bd6-88c9-36743cc2950d · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Llama-vid: An image is worth 2 tokens in large language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.217239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.217239Z digest=sha256:c2bdc454769f55e3a573da7b9f191b6de69f907e1117dfd4d4c4a16626bce805

Observation 29839910-d5e9-4898-a35d-64c2d808633e · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.299845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.299845Z digest=sha256:c3261d4826a3b104af24288764741d887b270f12069e56a175bd43330cd5bf2d

Observation 91d85d33-4de6-46c1-9662-ac35ff987eda · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VILA: On Pre-training for Visual Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.368568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.368568Z digest=sha256:a355288ff0ab92b5d7fbbbb925babfe977ab48d5ec5cb846d54951a6e00d9a41

Observation 1268ee51-1cde-4edf-9513-2bd023e36f54 · outbound

This paper cites Visual instruction tuning.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Visual instruction tuning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.438157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.438157Z digest=sha256:7801765f603494c4ae3f7a91827ff3c0c73109785da77035787e18ea0566891c

Observation 097793e8-a78c-4584-a32b-f5bb696a9267 · outbound

This paper cites Improved baselines with visual instruction tuning.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Improved baselines with visual instruction tuning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.544631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.544631Z digest=sha256:1fab09b5748db08433faedbab6c13a92dba6e3708fd4b78fb213854676527527

Observation 67cf2745-2a8b-4f86-9da3-44dafe224d18 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.642874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.642874Z digest=sha256:1a9becac152709a98beab398af4f47eae6dc4110e5cfef7b7a09875fdaf2d640

Observation ee3968c0-a0cf-49e6-afc9-d3ab647abbf9 · outbound

This paper cites Visual instruction tuning.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Visual instruction tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.709002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.709002Z digest=sha256:e8b6c81f16276803cb0009a5792bf20b33bbda34068cb4b7661c0db6fb745c3d

Observation ec5d62f4-428d-4497-bbd9-08a8a5b08701 · outbound

This paper cites E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.796096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.796096Z digest=sha256:55693852e53b6491db3708741b3beaa586e0b69a0c599d6e2aff543e5ba96849

Observation 131cfbbc-8a0e-42de-a5d0-d4c2040a6ada · outbound

This paper cites Snakes and Ladders: Two Steps Up for VideoMamba.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Snakes and Ladders: Two Steps Up for VideoMamba

Reference 59

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T20:29:57.830138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:29:52.891899Z digest=sha256:786d94cf80d976b3a7c1a98150ae16624adb5c9470a7add43981fe2790e7db6c

Observation fef4cecf-0169-4367-8e2e-009733dc94b6 · outbound

This paper cites Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.976080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.976080Z digest=sha256:50721bf38fd45d99d7aed8a948a1ea05d09ac3c511345231fc02651fba78ab01

Observation 2d58ae8a-7056-4704-8753-f10239e6cc0e · outbound

This paper cites Video Token Sparsification for Efficient Multimodal LLMs in Autonomous Driving.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video Token Sparsification for Efficient Multimodal LLMs in Autonomous Driving

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:29:57.794271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T20:29:53.072085Z digest=sha256:503489cfc31744182d5c563316471874cc767132f0c6cf0e591934fc5b78805e

Observation 08f1d850-86c2-4028-af62-fef6abd66539 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.161164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.161164Z digest=sha256:5fa710986ec352fdd678d7d550884a6c3a541ad06a56802d7c46c01d75775c6f

Observation 550f1580-0770-401c-8fdc-f1bf961f5b2d · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.257731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.257731Z digest=sha256:7b1cd8c688f8bb4ecfa1e524320b56d9748644bad087f9a288a123274bab05c5

Observation 174d41dd-2a7a-4ac1-91c0-ecce84463bed · outbound

This paper cites Videomamba: Spatio-temporal selective state space model.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Videomamba: Spatio-temporal selective state space model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.353271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.353271Z digest=sha256:c40e00697b56296976a0fe0b398d116869de34a5228a1d8b8fc849fe4a6c08fc

Observation 3c35457c-e2e5-4a8f-8bc9-e1ee253910da · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding RWKV: Reinventing RNNs for the Transformer Era

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.451063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.451063Z digest=sha256:14aaa4cd94395d2b06a6b5342516a1c77eb7e0706ce17859b58388ca5009e1e6

Observation 712d3c3f-5c91-418d-b940-6534d53f7b75 · outbound

This paper cites Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.546916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.546916Z digest=sha256:9386bdf30fbffb135f5b6c3b06a4b681e558a57d94ae290adc8e6b6b9899695a

Observation 99aa023c-f10a-4c49-99c2-00644696aa80 · outbound

This paper cites RWKV-7 "Goose" with Expressive Dynamic State Evolution.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding RWKV-7 "Goose" with Expressive Dynamic State Evolution

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.623195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.623195Z digest=sha256:81d62923746ddbb55cf40c1a00adb96450794f325156f523b3fadd74dd4b9773

Observation a9d8c3be-0243-4298-95b9-3ee921ca8923 · outbound

This paper cites VL-Mamba: Exploring State Space Models for Multimodal Learning.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VL-Mamba: Exploring State Space Models for Multimodal Learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.692582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.692582Z digest=sha256:f7314e4d82368ebc951eca950a1841f182d65d4fcf76f93faa582d68569a81e9

Observation 210f87e3-5250-4913-835f-ce77de68d1b6 · outbound

This paper cites Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.778431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.778431Z digest=sha256:820015d0d892ed5edf31a7a4b6e78fa9e7f4ae74388f197671fb5dd8aaa7449f

Observation 91df7b2a-5fb5-4c48-a2a8-6008dc862506 · outbound

This paper cites Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.879224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.879224Z digest=sha256:95dc62c27053d6b80ef3c64d15c83d0e53706592f3d275b410e862657a0a41e3

Observation b496daaa-7fa5-4104-8fc1-3028415a32f5 · outbound

This paper cites Automated as- sistance for creative writing with an rnn language model.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Automated as- sistance for creative writing with an rnn language model

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.933269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.933269Z digest=sha256:e46d4b17d792770f7b19ec45d3959b8a88db654f40bb9696c24a8ba979093afc

Observation a387132b-e8ff-4ff4-8820-df3656c28326 · outbound

This paper cites Bidirectional recur- rent neural networks.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Bidirectional recur- rent neural networks

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.985796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.985796Z digest=sha256:bdc138808e71eaa90bc2cce9d79f8f9db7dd5c0c1356ba728b2243af5ecaca5d

Observation 16a14370-b6e2-41eb-8cd6-6a4b73f12e0b · outbound

This paper cites Llava-prumerge: Adaptive token reduc- tion for efficient large multimodal models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Llava-prumerge: Adaptive token reduc- tion for efficient large multimodal models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.036800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.036800Z digest=sha256:94a81d2f32c77a47701257f6ece33038b0b720b3bf4bd11f9595fb27bacbfd17

Observation 743e4ac4-13eb-4ee1-8f93-8f2fc22b05e8 · outbound

This paper cites Disan: Directional self-attention network for rnn/cnn-free language understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Disan: Directional self-attention network for rnn/cnn-free language understanding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.092514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.092514Z digest=sha256:54c71d4367816b8c5861106ef3cdf20ab6a5c8743077ac2533e47514692dca8b

Observation 8a5837f7-493f-489d-bcdb-c2b0723c4d99 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.158609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.158609Z digest=sha256:9d2945e91035ad164f540e60327b9ad8226f7d8d841866f5359a801c3b1e9249

Observation 11789708-fe6c-41c0-b5a3-1c112e649775 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.223612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.223612Z digest=sha256:f709d75ed0917d79a76d095a0becf45d4b2263f7d3692623fd7816a23057d6f3

Observation 00d8e688-0c33-48cb-ab53-274cf486fbfb · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.291657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.291657Z digest=sha256:dbf22dd93e885c2212859aee9e446a737b8f120458fe34f2555f1a1e52b39ca6

Observation c4ebeb15-4443-4c06-9dcd-0a56d5531c28 · outbound

This paper cites MovieChat+: Question-aware Sparse Memory for Long Video Question Answering.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MovieChat+: Question-aware Sparse Memory for Long Video Question Answering

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.353632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.353632Z digest=sha256:dd13ddfcf5a10535eb7c438630c8386b9a946ccec2b08e4fcec9468ccac192f4

Observation 605c7da0-e77d-421b-a97a-1546f94eb854 · outbound

This paper cites Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.414573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.414573Z digest=sha256:912b04fb8a9bd002019fa5c3a705f2d82e393f8678b8de8fe314974bfa07b2c8

Observation c64bc55b-2e82-4402-8aae-ec3bec8db202 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding, 2021.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Roformer: Enhanced transformer with rotary position embedding, 2021

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.473573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.473573Z digest=sha256:9a6674bd5b89038fcab68d19bde8f090d9555c8c947e20007f01b1500222896e

Observation 1f64888b-413f-4e14-ac5b-6f1b0bd7b83a · outbound

This paper cites Koala: Key frame-conditioned long video-llm.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Koala: Key frame-conditioned long video-llm

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.531733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.531733Z digest=sha256:a26e262c93ef8aa413c043d04e12038960b0e2e1ed301e816df1c75122c7a912

Observation 9fdda23a-e379-4a88-8fda-78e5a5fba147 · outbound

This paper cites DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.595717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.595717Z digest=sha256:9275eb78cf55b5c77c50338348d859233e0477d33788d4db499df0da022195f2

Observation 4f764f96-db39-421c-a41d-58e3607f5b97 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.651903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.651903Z digest=sha256:a95915ce31dc1955de208aecb4c621e147585933ff037d0e0f6ea7e6e3ec551f

Observation 56dee312-c2e3-4278-9be9-1b72b046e87f · outbound

This paper cites Attention is all you need.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Attention is all you need

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.714517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.714517Z digest=sha256:eb8fa0a7e417f254650519f69a0c21306c443ba53e46fdf71946831bcf01c285

Observation 5c8a4b1f-7233-4bcb-9af4-7f93cd16f279 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.835249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.835249Z digest=sha256:dfe33d683c980327b53f5af63d616d29166e2246a2ac1d4c34983a3c15e98c30

Observation d0d78680-692b-4287-ba9b-5a160a7bad8c · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding CogVLM: Visual Expert for Pretrained Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.904223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.904223Z digest=sha256:eb9cef38307ac38828d06644a31ec887fc4ed04dae0b0124fa645d38104e209b

Observation 537224d8-9dba-45f2-902d-a7db50547967 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.965989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.965989Z digest=sha256:5a0beed1a8144e717073cd9702addcc3d2d91e1bf3cb3f5fe9159fa9c9be5592

Observation 335c0eca-0909-4bfa-ae52-8fd8012f6ea6 · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video-and-language research.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Vatex: A large-scale, high-quality multilingual dataset for video-and-language research

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.031669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.031669Z digest=sha256:8b53bb4592c155a2baae34b3e9fb7ae2e0ce7fbfe63e262096318928686861dd

Observation b0c89a3f-5ec9-4803-ab29-2c4119d62c42 · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.094037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.094037Z digest=sha256:0287c3dae518742f6b939b13a79bbf932f0f23fd26989bf6269aa3092f6f9de3

Observation a46fe6e4-f825-4248-a277-b170d844826d · outbound

This paper cites Longvlm: Efficient long video under- standing via large language models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Longvlm: Efficient long video under- standing via large language models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.171694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.171694Z digest=sha256:128067c05edb98aeba12f74df24e192c187684ed5abd9960210fb2bd0b72a091

Observation 1e423a24-2f0a-46a3-8cf4-c44392ca6812 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.224420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.224420Z digest=sha256:2083367ea3783203628a8cdb3844cb280464520b6adc0fb53e963b0b63319687

Observation e1f5b7ff-f3f4-47cf-84c7-597dfc837803 · outbound

This paper cites MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.275065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.275065Z digest=sha256:f17addc90d5cc458e0341a347746df6444a0cf921c5d2c52e77f65d7b81459ca

Observation 4bd87471-658a-43fb-b255-b133fc4b9a93 · outbound

This paper cites Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.355032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.355032Z digest=sha256:ede9bad22647ba9ade0307bb0f7f4b89c302f43a2b965330f569693008933fbb

Observation 7d5a20d2-bc5b-4ea3-b24e-2cd0facb5bf6 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.468866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.468866Z digest=sha256:b79701442cbaf5941c85c8d607c194cd97f1a4083b3b1ea24e4da250e904370d

Observation e7dcce52-2094-48fc-add4-b66e602eeb3a · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding xgen-mm (blip-3): A family of open large multimodal models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.594072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.594072Z digest=sha256:fa15b7517b626806641123ce4940db021a1e0dbfdee719e35f3e42402c07dc2f

Observation a52bc8ad-4b85-42e8-bc8f-a01a40173fc7 · outbound

This paper cites Qwen2 Technical Report.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Qwen2 Technical Report

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.704606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.704606Z digest=sha256:1564c306f4bca99fce0b4db389189cb5525daa2da59d3d371147986378e74caa

Observation 0038c185-3efb-4b91-8d47-5a0d438fc751 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.864749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.864749Z digest=sha256:8f6dea9822a07a052da4e285f472adeb662f57c24026214e20b465ff47eefb6f

Observation 78ba265f-537d-4ff3-9b1d-371127275889 · outbound

This paper cites Gated Delta Networks: Improving Mamba2 with Delta Rule.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Gated Delta Networks: Improving Mamba2 with Delta Rule

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:56.031745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:56.031745Z digest=sha256:c42ac998332be4d1b291180673d61fd575f65a7787a20ab8124015baa38bded8

Observation 2b0430c1-0fd4-4b5a-82b4-bf8817005e6e · outbound

This paper cites Parallelizing linear transformers with the delta rule over sequence length.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Parallelizing linear transformers with the delta rule over sequence length

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:56.156114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:56.156114Z digest=sha256:ad10693e2cf33e48e44b6936f9bfce4804c8e0e0b89fcf432dc5f602d620079c

Observation f8077658-29a0-4f35-a5a3-9d480e1c43d4 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:56.364545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:56.364545Z digest=sha256:442eb168d0d7280573cf490fd663118e3fcd5e61ad27fe5eb1fd69c6de93120e

Observation 43514bb3-a1bf-4712-adca-c41d2225b911 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:56.537267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:56.537267Z digest=sha256:5c0d522d121273bbb12565a11d6bb19758156a264fa2eabc727e68035d6165f8

Pith citing papers

Observation 908cfaf1-2bcd-40b4-8f85-5f8f63e346e6 · inbound

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models cites this paper.

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T14:57:29.171219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:57:29.171219Z digest=sha256:e70cd03ea98edde74b57db027c7bada4665b1804a6e75fa1f414f70b35f66c35