Pith. sign in

Paper Citation Record · LEDGER

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding

As of 10 August 2026, this Paper Citation Record lists 100 of 120 outbound references and 1 inbound Pith citation observation for arXiv:2507.02591.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.02591 v3

Coverage vector

measured 100 of 120 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:29:56.537267Z

measured 101 of 101 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:57:29.171219Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 120 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved97
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fd7260af-fe08-4fda-8399-745342c69b9c · outbound

This paper cites GPT-4 Technical Report.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.170619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.170619Z digest=sha256:06b5037ad47bc615fdc4826a405aaae35833f51b88e8b885bfb7a583356347bd

Observation 26ac5815-ddda-4ff1-8f70-55b952b16ead · outbound

This paper cites HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.222083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.222083Z digest=sha256:6c5bf83db0b1739f3b9b6e551e4998338cf2b781eac618ee045fff601e536471

Observation 6e71df41-cddb-4ed0-9cba-c88cf129020e · outbound

This paper cites In- finibench: A comprehensive benchmark for large multi- modal models in very long video understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding In- finibench: A comprehensive benchmark for large multi- modal models in very long video understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.286948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.286948Z digest=sha256:c5939311523499b4a43dd41941b875e1e1a19196513c556ec8fe76eaf6d18d31

Observation 050ede92-a899-4723-9dfa-8c1005e9b05d · outbound

This paper cites Qwen2.5-VL Technical Report.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.428876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.428876Z digest=sha256:11ca2e98810c94b831fb17a563adad6eaee6859286129e71b212c99219305a06

Observation 82ac5bc2-d64c-4066-bf61-72049155ccc7 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding PaliGemma: A versatile 3B VLM for transfer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.584546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.584546Z digest=sha256:9579143828e1c050f4af86670b77039a1ca0d6cf632c9d108f3d689afd0236dc

Observation 86aa583b-e38f-4500-9435-3eea1cb6fa26 · outbound

This paper cites VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.711985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.711985Z digest=sha256:62ed6ed19bfbc90699cee2d8fff6423fe80b6689ad8b9e77c81a188074ce1949

Observation 400a4ef2-710e-4470-b21a-1d604f15c3cb · outbound

This paper cites Token merging: Your ViT but faster.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Token merging: Your ViT but faster

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.814996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.814996Z digest=sha256:df1d5b1aa1ac3aeb5da43ff07b566febd101fdfac2b4318d00c6bd66e5c4e681

Observation 7fb399f3-d917-44ac-88d5-35187acd98c2 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Activitynet: A large-scale video benchmark for human activity understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.918119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.918119Z digest=sha256:b386209665279de87fb33eef782148b9bbb139657331dfaa60bf4cde6d575070

Observation 0f3e4acc-c4f0-4034-9e3f-43b4e2628177 · outbound

This paper cites Matryoshka Multimodal Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Matryoshka Multimodal Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:47.998290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:47.998290Z digest=sha256:c56bdc4c63cad4923a40bdc1045236cfa10bf1fde66acc38fe3de82d9689c474

Observation 6eb9324b-c854-49d8-9a63-c23479ecc6c8 · outbound

This paper cites View transformer layers from online optimization perspective, 2025.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding View transformer layers from online optimization perspective, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.089471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.089471Z digest=sha256:a324cd7c184a678d605238a82946653cdd0435281c3fa9138c23b121c0fb0662

Observation 7b21e852-8f3b-4f2d-9372-9a4dce0f8077 · outbound

This paper cites AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.227081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.227081Z digest=sha256:a8f17dc44ead10b67ab93de8ea98effd0f6d017c985c4bfbcc02f5839d145cd4

Observation 05cfcee0-61aa-4dad-95f3-7633ca240c40 · outbound

This paper cites Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.369805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.369805Z digest=sha256:3f5e4fe47fd38175573fd3cd71f73ea4c69f6b6c91e183fe61c38bf440892b8c

Observation 67515cb2-5b29-4a06-be65-585b37f24353 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.489965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.489965Z digest=sha256:af8bad8ec70a2bd35990b0ef053ef32d88c470d199869b323fd24030715af607

Observation eef22555-b744-4038-8e00-8a46b63f3888 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.568606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.568606Z digest=sha256:94163d2837e22cd48b4043fe688aa5978c0f1e0d0985f6e53813690be26657e2

Observation 54b94232-4936-4c5b-83ae-3360eb5bea48 · outbound

This paper cites MotionLLM: Understanding Human Behaviors from Human Motions and Videos.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.674994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.674994Z digest=sha256:ded29c244b54d7ea327d0ed68eba20695b308580e5ad001c496a441e89b50881

Observation 518e2809-35a0-4685-8c4c-de6170ccf924 · outbound

This paper cites Stuffed mamba: State col- lapse and state capacity of rnn-based long-context model- ing.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Stuffed mamba: State col- lapse and state capacity of rnn-based long-context model- ing

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:48.973755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:48.973755Z digest=sha256:a88d74ef0a78d3c17b0a41c4ec63dde5c925c96a067faf440de5a554e2063d0e

Observation 658d614a-3946-4d30-b409-c44a47429924 · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.077308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.077308Z digest=sha256:c99712ed137961d05db6007fbcbb28e0c3ccc926c7da9e4ce5b3b1f892d550d5

Observation 2159393d-359b-4bf4-b3c1-2c44b1b3c8ca · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.151067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.151067Z digest=sha256:705ae11e3b10b19ddf01934bef4df3298839460bc1d708f6999bf570c851779f

Observation 5125826a-e05b-44f4-8e0e-d056a65494fe · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.254227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.254227Z digest=sha256:d9d7586690c1083ccbb2b0d8072ce9d00f02ec306b45917663b10e2e77aa1cc4

Observation e5c1393e-2194-4f0a-9adc-f040f4e1f40c · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.334244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.334244Z digest=sha256:a74615db696d2ad5d3cb0444ea41699a821d7019374ec1e71eb5d60ac1dce245

Observation f7e63bfc-07a6-4c67-a0fe-1cbf2929b32e · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.455631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.455631Z digest=sha256:52b70cdcc57a18d69512301e52dd81ba850e865dd498eac06e06590367e6576f

Observation 889cdfca-77ea-444e-be11-5f1d974dba96 · outbound

This paper cites Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.624319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.624319Z digest=sha256:4bed11a5f099489a763c38eb887842d1e3e2357f6520108b60e553486b018d51

Observation d6032bc0-b3ab-48ea-937f-15b7b414e895 · outbound

This paper cites Gate-variants of gated re- current unit (gru) neural networks.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Gate-variants of gated re- current unit (gru) neural networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.743701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.743701Z digest=sha256:aa0d0355465a625bbf8f42bf07704fdfec6d9295ba3fd605c5edf56ae6c30ceb

Observation 864436db-4c6d-4a50-bea4-c18eb3217e64 · outbound

This paper cites Towards Event-oriented Long Video Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Towards Event-oriented Long Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.820033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.820033Z digest=sha256:45a3b7cad02bde5c408eda823f30c8595cb2330b4a7fa61eb574b8988cfb82b8

Observation 67b408c5-29ca-49df-a29c-9846938969cd · outbound

This paper cites Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:49.891728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:49.891728Z digest=sha256:96a7488cb16e1c05e902440e7104ca1643128194f4d3163a79655f6498293bc0

Observation 7adf30ca-8359-47d1-833d-9416aa5f6e3c · outbound

This paper cites Videoagent: A memory-augmented mul- timodal agent for video understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Videoagent: A memory-augmented mul- timodal agent for video understanding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.028784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.028784Z digest=sha256:abd3709d66e8ee3cc63af9e86eabeb2e3dbabce4f8b01dd5423a9d3d5093b278

Observation cad1da19-64d5-40e3-9867-bd0534405420 · outbound

This paper cites Were RNNs All We Needed?.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Were RNNs All We Needed?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.131873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.131873Z digest=sha256:9125eec3f785b1bab8e69237696d985791e7012dbea4b6207fd278b9da9fc094

Observation 7568dce7-d255-4613-9cc3-9f90767b4b4c · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.229493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.229493Z digest=sha256:91322a15c71a29b81f8bcb9b1bc013582a36d032dbdb8b528972d192760fa687

Observation bc7552e6-52e4-4ee8-971f-3618e1c7ba2d · outbound

This paper cites Long short-term memory.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Long short-term memory

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.302154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.302154Z digest=sha256:5e9342b43aa6e98cb3e520464e74ef3c501ed18353e049f4ada9cd308313f613

Observation 77d5682a-b30c-42db-b433-4261c4fdda77 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.386856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.386856Z digest=sha256:0d4ea99ba3c9c72fd5f65ad71af25f5333f2a35358592935d5ca0fa02ec091dd

Observation 6f53dc11-111a-4c9e-9ca0-ef9b3f4c1988 · outbound

This paper cites Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.482506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.482506Z digest=sha256:f8267e24c2bfc4c58a0bd908bc7d928bfc7cc471d1ad5ee8b9f39d2e2bb3719d

Observation fa2f5207-8970-413d-bf4e-4d804d98ed10 · outbound

This paper cites Ma-lmm: Memory-augmented large multimodal model for long-term video understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Ma-lmm: Memory-augmented large multimodal model for long-term video understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.583701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.583701Z digest=sha256:d0ad238b44120182e5f9a9db09abd3eb5453d4d35ef2dd57ca577fb7e683d4f3

Observation 5638984e-b43e-4833-a759-52ab312519ab · outbound

This paper cites Masked autoencoders are scal- able vision learners.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Masked autoencoders are scal- able vision learners

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.658548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.658548Z digest=sha256:159eff3ce30bfcc2043ddd0c3183c4768f47d98d8bfd06b040878c21843fd0ab

Observation 6fb40779-dca4-40bb-abb8-87b94a84d98a · outbound

This paper cites VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:29:57.988732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:29:50.789507Z digest=sha256:442d685d19c82f3189adf85d84bb2069b2ce03c7c7fd60ff002985509362e6e9

Observation 6e4b89be-19af-4698-b465-3a75f94544c0 · outbound

This paper cites Token compensator: Altering in- ference cost of vision transformer without re-tuning.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Token compensator: Altering in- ference cost of vision transformer without re-tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.888495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.888495Z digest=sha256:03d5cd01a1ed1113aa580dd610e9802405938336a19548411db82c1539529a7c

Observation 65cf98aa-0a57-421f-8ddf-f8a99652eb0a · outbound

This paper cites Chat-univi: Unified visual representation em- powers large language models with image and video under- standing.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Chat-univi: Unified visual representation em- powers large language models with image and video under- standing

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.012187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.012187Z digest=sha256:b8683f67d6f9e69b0012e8d4bf44ae1dc2bdc92fe560b93141f7aefa9a012548

Observation cd13014e-7ead-46af-9005-82a217effe0b · outbound

This paper cites Exploring Enhanced Contextual Information for Video-Level Object Tracking.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Exploring Enhanced Contextual Information for Video-Level Object Tracking

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.094833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.094833Z digest=sha256:8ab7f9cdbcc9c73bc54eb4e36b027f892e4fba9c2b42b042900c9b7d688d3b18

Observation e660e977-e55a-46bf-b95b-d8a5b4637ed1 · outbound

This paper cites Transformers are rnns: Fast autore- gressive transformers with linear attention.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Transformers are rnns: Fast autore- gressive transformers with linear attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.173839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.173839Z digest=sha256:dab1668c9a7d7bb8eff63b98b2ac0c5d9ad194219b030bb75f9215920b9f52e8

Observation 7b51d115-5159-4ae8-ad73-38f1c729ed51 · outbound

This paper cites Rethinking Positional Encoding in Language Pre-training.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Rethinking Positional Encoding in Language Pre-training

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.253360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.253360Z digest=sha256:7e5b67d833e6c0fc37548afe21b0ea5a8989021f9bb95e6e9dbdd951f843f9f4

Observation 1859956c-fa26-4f03-acbe-d34f0e8363e9 · outbound

This paper cites Video Token Merging for Long-form Video Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video Token Merging for Long-form Video Understanding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.350310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.350310Z digest=sha256:ed6130f8417fa657cc9e13ec6c8030bd05ebc018df478eecb39036405ec20eee

Observation b5b48a0b-b677-4c05-b7c6-cca924d5aa4e · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.447003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.447003Z digest=sha256:ac32ebad1a2a959c80d96d25aa99c82e0e20045b73690395bc13cf8b0c8d9de8

Observation 41e8f07a-f02d-405e-83f9-e0ba620521e1 · outbound

This paper cites Lmms-eval: Accelerating the development of large multimoal models, 2024.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Lmms-eval: Accelerating the development of large multimoal models, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.537802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.537802Z digest=sha256:57894714a3695c82cab75d289b2080bfe2de71cd71a41907d1e14d71f4db6de5

Observation ce37d5cf-a163-49c2-bf22-184a2702f662 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LLaVA-OneVision: Easy Visual Task Transfer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.630112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.630112Z digest=sha256:5661ae9cd345aa58529f34a4dd96a0f573a4b250993bdb89643973df4648dafd

Observation 67cf1236-7fa1-4118-8cbe-6711b0775924 · outbound

This paper cites Aria: An Open Multimodal Native Mixture-of-Experts Model.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.704165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.704165Z digest=sha256:ab4c809c899b2dd118b718431ed0bdc47e64750032f78a41d8075398a017c428

Observation 96efacf1-46f1-4ad7-8ead-ec7d9589b0bb · outbound

This paper cites Mvbench: A comprehensive multi-modal video under- standing benchmark.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Mvbench: A comprehensive multi-modal video under- standing benchmark

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.787241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.787241Z digest=sha256:a079d94cb6a263de40fb2bdcddd7538cc91aae22064501275f96904422e994b0

Observation cc3df688-87ae-43be-9cbf-90e6559d12ff · outbound

This paper cites Videomamba: State space model for efficient video understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Videomamba: State space model for efficient video understanding

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.870161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.870161Z digest=sha256:eac356382ecd2f4274e5876e42189d7ef2338c34acc6c856c1f4186d232c7d5b

Observation eb4f2bdf-9eaf-4eca-9789-49d5d5680fa3 · outbound

This paper cites Independently recurrent neural network (indrnn): Building a longer and deeper rnn.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Independently recurrent neural network (indrnn): Building a longer and deeper rnn

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:51.962556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:51.962556Z digest=sha256:c93d411d9e5a89b7ec94ec0a974f662f9e49c8116e189612fddf65af12b9f71f

Observation 4d725fa6-89c5-47d4-8899-7fd448e4b2c5 · outbound

This paper cites Mamba- nd: Selective state space modeling for multi-dimensional data.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Mamba- nd: Selective state space modeling for multi-dimensional data

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.058317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.058317Z digest=sha256:4c96f5775d2b7fd873cad155784231e49e67434021f4e11c400d832cfe08ab4b

Observation 2efa0987-9240-45d2-aa06-4a09a803cc06 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.144931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.144931Z digest=sha256:8ad811dfeb6ed8c1a4009bf37a9b17d784cebb520becc6230063c96624bb12d7

Observation 427d13d9-fd81-4bd6-88c9-36743cc2950d · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Llama-vid: An image is worth 2 tokens in large language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.217239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.217239Z digest=sha256:3e842b41ef83ca6542a66dba31956032f21feb10d8ccaefe3ba0561475912ea6

Observation 29839910-d5e9-4898-a35d-64c2d808633e · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.299845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.299845Z digest=sha256:17ad104a2cae524bdd8a49ded2acd1820221753786e5305a57edc5c4d819bb04

Observation 91d85d33-4de6-46c1-9662-ac35ff987eda · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VILA: On Pre-training for Visual Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.368568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.368568Z digest=sha256:08d644ff268665a9a5495a4643bb57c9516ca2dbba144d79c902a30d51bf0f00

Observation 1268ee51-1cde-4edf-9513-2bd023e36f54 · outbound

This paper cites Visual instruction tuning.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Visual instruction tuning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.438157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.438157Z digest=sha256:aefe21637e20e627d6e123d6d5cfe4375eea538b40c5a3e3358bcf6226a3a260

Observation 097793e8-a78c-4584-a32b-f5bb696a9267 · outbound

This paper cites Improved baselines with visual instruction tuning.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Improved baselines with visual instruction tuning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.544631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.544631Z digest=sha256:266c5141941f46d357991e7bbcd41a19da6d42be6546f14ed0f4891dc49dd400

Observation 67cf2745-2a8b-4f86-9da3-44dafe224d18 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.642874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.642874Z digest=sha256:ef175ef8d13420f186751b9f315032c520af422c0a31b776bdcc3cd34c1c00e0

Observation ee3968c0-a0cf-49e6-afc9-d3ab647abbf9 · outbound

This paper cites Visual instruction tuning.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Visual instruction tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.709002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.709002Z digest=sha256:cb60bf00bc6836e0afb2217eae9d58fb54e50001677a3a372f417386e488bc0d

Observation ec5d62f4-428d-4497-bbd9-08a8a5b08701 · outbound

This paper cites E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.796096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.796096Z digest=sha256:a150f4ca57df1185f5a26c53af8669c33ee168350d0acb6038b7be2a6e6517a2

Observation 131cfbbc-8a0e-42de-a5d0-d4c2040a6ada · outbound

This paper cites Snakes and Ladders: Two Steps Up for VideoMamba.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Snakes and Ladders: Two Steps Up for VideoMamba

Reference 59

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T20:29:57.830138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:29:52.891899Z digest=sha256:36c0ee229c045f881b613ed12f338ddf5a45285d7d43ffb5f034d4683bd21558

Observation fef4cecf-0169-4367-8e2e-009733dc94b6 · outbound

This paper cites Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:52.976080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:52.976080Z digest=sha256:e05e30982e8ad251dc50d32949bba8aab7993cadba6a2ab880ff3a1c672d48ed

Observation 2d58ae8a-7056-4704-8753-f10239e6cc0e · outbound

This paper cites Video Token Sparsification for Efficient Multimodal LLMs in Autonomous Driving.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video Token Sparsification for Efficient Multimodal LLMs in Autonomous Driving

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-06T20:29:57.794271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T20:29:53.072085Z digest=sha256:128f790cd520ba80350fd97939350a0a8767100a7d02b52997c7e48ec2f6cae5

Observation 08f1d850-86c2-4028-af62-fef6abd66539 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.161164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.161164Z digest=sha256:c1ff60b83f224d86dc4a43c983564b16fcf3f2cc67ecccfe8caaacb2ee2dbb98

Observation 550f1580-0770-401c-8fdc-f1bf961f5b2d · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.257731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.257731Z digest=sha256:9512a65b9f6b96dba9e9fbf0dd4d522f6eb3470591fdf97c6994bc80024be114

Observation 174d41dd-2a7a-4ac1-91c0-ecce84463bed · outbound

This paper cites Videomamba: Spatio-temporal selective state space model.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Videomamba: Spatio-temporal selective state space model

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.353271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.353271Z digest=sha256:7942dde6cd84f88471e4aea3af41c4f7239dbf75b923b7ffafdfa4fcc06c351d

Observation 3c35457c-e2e5-4a8f-8bc9-e1ee253910da · outbound

This paper cites RWKV: Reinventing RNNs for the Transformer Era.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding RWKV: Reinventing RNNs for the Transformer Era

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.451063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.451063Z digest=sha256:af5f75e40871cc7ffc5ba1da4e473b82e0ee8ceb9a69357ca06e1c64d335036b

Observation 712d3c3f-5c91-418d-b940-6534d53f7b75 · outbound

This paper cites Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.546916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.546916Z digest=sha256:ccb404588cb8c97de714e0ac9d84aa517b719171fb269133075691a5146dd72c

Observation 99aa023c-f10a-4c49-99c2-00644696aa80 · outbound

This paper cites RWKV-7 "Goose" with Expressive Dynamic State Evolution.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding RWKV-7 "Goose" with Expressive Dynamic State Evolution

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.623195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.623195Z digest=sha256:818146afda97e7995876ece3107d1729123bba3b0c285dfddbcd0a29571e21a1

Observation a9d8c3be-0243-4298-95b9-3ee921ca8923 · outbound

This paper cites VL-Mamba: Exploring State Space Models for Multimodal Learning.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VL-Mamba: Exploring State Space Models for Multimodal Learning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.692582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.692582Z digest=sha256:c17fc3eb97c341080e04a1434eff2e867dbe7c9142e22d4f280107604d72be90

Observation 210f87e3-5250-4913-835f-ce77de68d1b6 · outbound

This paper cites Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.778431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.778431Z digest=sha256:20ad0c35c8ee7af34962b33e97f26d6411628854c4fe846d5d01913b5093eb2a

Observation 91df7b2a-5fb5-4c48-a2a8-6008dc862506 · outbound

This paper cites Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.879224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.879224Z digest=sha256:21dde9602662b9ffdc4e11d5aa4085981b5647c2d04637cc7038f789d848d76d

Observation b496daaa-7fa5-4104-8fc1-3028415a32f5 · outbound

This paper cites Automated as- sistance for creative writing with an rnn language model.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Automated as- sistance for creative writing with an rnn language model

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.933269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.933269Z digest=sha256:f5ebec4127a8a375b2aa21622a7abd68448e8af853ef60bc59d1cb064ac245f3

Observation a387132b-e8ff-4ff4-8820-df3656c28326 · outbound

This paper cites Bidirectional recur- rent neural networks.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Bidirectional recur- rent neural networks

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:53.985796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:53.985796Z digest=sha256:dd93d04b2606db6d9bb24c15cfbbf1851bcac06cb8c042428d0c5c47ceece399

Observation 16a14370-b6e2-41eb-8cd6-6a4b73f12e0b · outbound

This paper cites Llava-prumerge: Adaptive token reduc- tion for efficient large multimodal models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Llava-prumerge: Adaptive token reduc- tion for efficient large multimodal models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.036800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.036800Z digest=sha256:0b4f5fa13f6605b0d8a93b1b2df493bb84e30ac8e52681697d4f6e1a3a61f65b

Observation 743e4ac4-13eb-4ee1-8f93-8f2fc22b05e8 · outbound

This paper cites Disan: Directional self-attention network for rnn/cnn-free language understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Disan: Directional self-attention network for rnn/cnn-free language understanding

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.092514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.092514Z digest=sha256:ca81fff1bf551320458af727c9c1882f4229664619a656936f0605f313bee703

Observation 8a5837f7-493f-489d-bcdb-c2b0723c4d99 · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.158609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.158609Z digest=sha256:f8f15ec784fecf8ec6ef637fdbac66cedbac01fa8413e35d62ca5065bbe5891c

Observation 11789708-fe6c-41c0-b5a3-1c112e649775 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.223612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.223612Z digest=sha256:156af6a84abc374202c26819aef014ec7bc0e9a0b4270362680fac546a83db75

Observation 00d8e688-0c33-48cb-ab53-274cf486fbfb · outbound

This paper cites MovieChat: From Dense Token to Sparse Memory for Long Video Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.291657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.291657Z digest=sha256:0a89b38d4fe4affbb005aba95bb6f5c05762c4836a5767dec428219382228bf2

Observation c4ebeb15-4443-4c06-9dcd-0a56d5531c28 · outbound

This paper cites MovieChat+: Question-aware Sparse Memory for Long Video Question Answering.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MovieChat+: Question-aware Sparse Memory for Long Video Question Answering

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.353632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.353632Z digest=sha256:8508048bff4f497383c12e7426e95671bc4beb17fd092ef9a5fe1ae14449747b

Observation 605c7da0-e77d-421b-a97a-1546f94eb854 · outbound

This paper cites Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.414573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.414573Z digest=sha256:335306b3563da8027f8701c481a13705d25f246cac49f1860c22f6e0cf13212c

Observation c64bc55b-2e82-4402-8aae-ec3bec8db202 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding, 2021.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Roformer: Enhanced transformer with rotary position embedding, 2021

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.473573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.473573Z digest=sha256:1d07ae838e503d67487380b4458044ea3e5125f2b580af7bd0f8215ed714961a

Observation 1f64888b-413f-4e14-ac5b-6f1b0bd7b83a · outbound

This paper cites Koala: Key frame-conditioned long video-llm.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Koala: Key frame-conditioned long video-llm

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.531733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.531733Z digest=sha256:4efd73619fc63260c249f23fdefd1e74cc70e632a9232adf3dc8b39d1949d91b

Observation 9fdda23a-e379-4a88-8fda-78e5a5fba147 · outbound

This paper cites DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.595717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.595717Z digest=sha256:6a39acf2c73453c97bb4df90fb9ff883529f50fc680e971c67d4a7788740d85f

Observation 4f764f96-db39-421c-a41d-58e3607f5b97 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LLaMA: Open and Efficient Foundation Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.651903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.651903Z digest=sha256:2f7edf221064680b34982d3a8265088c28aa9ef7201cdc0b9ee6750c7824b6ee

Observation 56dee312-c2e3-4278-9be9-1b72b046e87f · outbound

This paper cites Attention is all you need.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Attention is all you need

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.714517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.714517Z digest=sha256:79bf0b91263f5a4fd053d2abdbdfcea7f577e3e0eb29b094c52be58c6c4d6c8a

Observation 5c8a4b1f-7233-4bcb-9af4-7f93cd16f279 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.835249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.835249Z digest=sha256:49e7c98ddb9baed43b64dc74fa56ff036276d42ba364da7aad39e95d7ce2a283

Observation d0d78680-692b-4287-ba9b-5a160a7bad8c · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding CogVLM: Visual Expert for Pretrained Language Models

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.904223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.904223Z digest=sha256:4036fef2fea471d5c5e994eeb919cf3850a47b3c5546d55a3c6c503ded3d8ade

Observation 537224d8-9dba-45f2-902d-a7db50547967 · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LVBench: An Extreme Long Video Understanding Benchmark

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:54.965989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:54.965989Z digest=sha256:d1af92bb698f4a5aaeb1fbfc7aed2cc1c6350dc2158bf25a419b24c7ea78aeeb

Observation 335c0eca-0909-4bfa-ae52-8fd8012f6ea6 · outbound

This paper cites Vatex: A large-scale, high-quality multilingual dataset for video-and-language research.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Vatex: A large-scale, high-quality multilingual dataset for video-and-language research

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.031669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.031669Z digest=sha256:e104354dab38f73a7993f32bf7c0f0c965c9e08c0f9703ffd35779095b2c841f

Observation b0c89a3f-5ec9-4803-ab29-2c4119d62c42 · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.094037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.094037Z digest=sha256:89f4c94bfdd16c6fb09c4a3a66cc3533e43a9e895e151de6808a3171ace2b621

Observation a46fe6e4-f825-4248-a277-b170d844826d · outbound

This paper cites Longvlm: Efficient long video under- standing via large language models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Longvlm: Efficient long video under- standing via large language models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.171694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.171694Z digest=sha256:9dac9661f4ad3ba7666679bce938188ab3fdd8507f622ed076b0f1fd3ce17f22

Observation 1e423a24-2f0a-46a3-8cf4-c44392ca6812 · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.224420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.224420Z digest=sha256:a1350e500c4b756a9a9591245ac46205e35fa48354a5476b140cf947fad4d555

Observation e1f5b7ff-f3f4-47cf-84c7-597dfc837803 · outbound

This paper cites MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.275065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.275065Z digest=sha256:15b24f1dde35f31ccf287d5b12f5851f306f04b9e5cec6f3d3ce108663636d49

Observation 4bd87471-658a-43fb-b255-b133fc4b9a93 · outbound

This paper cites Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.355032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.355032Z digest=sha256:66c1eceff1e78509c28c7c16ed3cbff9ae5927b1bc25e7fb7df8eb643248a045

Observation 7d5a20d2-bc5b-4ea3-b24e-2cd0facb5bf6 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.468866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.468866Z digest=sha256:eb941eca8cfb16b7549edfe8b32f0128ac7c8eb71f153c0cbbb0b3d0628d500d

Observation e7dcce52-2094-48fc-add4-b66e602eeb3a · outbound

This paper cites xgen-mm (blip-3): A family of open large multimodal models.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding xgen-mm (blip-3): A family of open large multimodal models

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.594072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.594072Z digest=sha256:117cc86295e7346785d0826f931519b387948d0d9404835659e58f7c7e7022fd

Observation a52bc8ad-4b85-42e8-bc8f-a01a40173fc7 · outbound

This paper cites Qwen2 Technical Report.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Qwen2 Technical Report

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.704606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.704606Z digest=sha256:854220e0486c81163bf484abb98ba5c8a826dadfb52bd89d71d2b9546daeb525

Observation 0038c185-3efb-4b91-8d47-5a0d438fc751 · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:55.864749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:55.864749Z digest=sha256:f7d7c8602eb14d3fcc943a36a3198f036ef0b259a4483afb2c9b005b3494c3ee

Observation 78ba265f-537d-4ff3-9b1d-371127275889 · outbound

This paper cites Gated Delta Networks: Improving Mamba2 with Delta Rule.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Gated Delta Networks: Improving Mamba2 with Delta Rule

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:56.031745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:56.031745Z digest=sha256:eee454cbc9cd82ef4a15b22587dfd4a0cfc61d198417a93d6b8c40b59a2f549a

Observation 2b0430c1-0fd4-4b5a-82b4-bf8817005e6e · outbound

This paper cites Parallelizing linear transformers with the delta rule over sequence length.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Parallelizing linear transformers with the delta rule over sequence length

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:56.156114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:56.156114Z digest=sha256:9ead206c590777fe079e373ddbea4feb9e3061417d66209b77fb235d3400110f

Observation f8077658-29a0-4f35-a5a3-9d480e1c43d4 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:56.364545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:56.364545Z digest=sha256:afd8620cafbf5b93d3962c7429d8a4b80066f3abd407e5df685b09f3dc741e09

Observation 43514bb3-a1bf-4712-adca-c41d2225b911 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:56.537267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:56.537267Z digest=sha256:2638fab8181db574d014285510c123a3bf271034df14f071bd1a4cba757f89f3

Pith citing papers

Observation 908cfaf1-2bcd-40b4-8f85-5f8f63e346e6 · inbound

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models cites this paper.

$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T14:57:29.171219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:57:29.171219Z digest=sha256:4fdd25d920476516eb7c5036e33408d8022e63b9caed31fa9fd3507f8fc1efe3