Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:29:56.537267Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 100 of 120 outbound references and 1 inbound Pith citation observation for arXiv:2507.02591.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T20:29:56.537267Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T14:57:29.171219Z
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 120 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fd7260af-fe08-4fda-8399-745342c69b9c · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26ac5815-ddda-4ff1-8f70-55b952b16ead · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding HiRED: Attention-Guided Token Dropping for Efficient Inference of High-Resolution Vision-Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e71df41-cddb-4ed0-9cba-c88cf129020e · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding In- finibench: A comprehensive benchmark for large multi- modal models in very long video understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 050ede92-a899-4723-9dfa-8c1005e9b05d · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Qwen2.5-VL Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82ac5bc2-d64c-4066-bf61-72049155ccc7 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding PaliGemma: A versatile 3B VLM for transfer
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86aa583b-e38f-4500-9435-3eea1cb6fa26 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VANE-Bench: Video Anomaly Evaluation Benchmark for Conversational LMMs
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 400a4ef2-710e-4470-b21a-1d604f15c3cb · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Token merging: Your ViT but faster
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fb399f3-d917-44ac-88d5-35187acd98c2 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Activitynet: A large-scale video benchmark for human activity understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f3e4acc-c4f0-4034-9e3f-43b4e2628177 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Matryoshka Multimodal Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6eb9324b-c854-49d8-9a63-c23479ecc6c8 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding View transformer layers from online optimization perspective, 2025
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b21e852-8f3b-4f2d-9372-9a4dce0f8077 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05cfcee0-61aa-4dad-95f3-7633ca240c40 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video Mamba Suite: State Space Model as a Versatile Alternative for Video Understanding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67515cb2-5b29-4a06-be65-585b37f24353 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eef22555-b744-4038-8e00-8a46b63f3888 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54b94232-4936-4c5b-83ae-3360eb5bea48 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MotionLLM: Understanding Human Behaviors from Human Motions and Videos
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 518e2809-35a0-4685-8c4c-de6170ccf924 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Stuffed mamba: State col- lapse and state capacity of rnn-based long-context model- ing
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 658d614a-3946-4d30-b409-c44a47429924 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2159393d-359b-4bf4-b3c1-2c44b1b3c8ca · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5125826a-e05b-44f4-8e0e-d056a65494fe · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5c1393e-2194-4f0a-9adc-f040f4e1f40c · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7e63bfc-07a6-4c67-a0fe-1cbf2929b32e · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 889cdfca-77ea-444e-be11-5f1d974dba96 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6032bc0-b3ab-48ea-937f-15b7b414e895 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Gate-variants of gated re- current unit (gru) neural networks
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 864436db-4c6d-4a50-bea4-c18eb3217e64 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Towards Event-oriented Long Video Understanding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67b408c5-29ca-49df-a29c-9846938969cd · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Vision-RWKV: Efficient and Scalable Visual Perception with RWKV-Like Architectures
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7adf30ca-8359-47d1-833d-9416aa5f6e3c · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Videoagent: A memory-augmented mul- timodal agent for video understanding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cad1da19-64d5-40e3-9867-bd0534405420 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Were RNNs All We Needed?
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7568dce7-d255-4613-9cc3-9f90767b4b4c · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc7552e6-52e4-4ee8-971f-3618e1c7ba2d · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Long short-term memory
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77d5682a-b30c-42db-b433-4261c4fdda77 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f53dc11-111a-4c9e-9ca0-ef9b3f4c1988 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa2f5207-8970-413d-bf4e-4d804d98ed10 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5638984e-b43e-4833-a759-52ab312519ab · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Masked autoencoders are scal- able vision learners
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fb40779-dca4-40bb-abb8-87b94a84d98a · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6e4b89be-19af-4698-b465-3a75f94544c0 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Token compensator: Altering in- ference cost of vision transformer without re-tuning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65cf98aa-0a57-421f-8ddf-f8a99652eb0a · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Chat-univi: Unified visual representation em- powers large language models with image and video under- standing
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd13014e-7ead-46af-9005-82a217effe0b · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Exploring Enhanced Contextual Information for Video-Level Object Tracking
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e660e977-e55a-46bf-b95b-d8a5b4637ed1 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Transformers are rnns: Fast autore- gressive transformers with linear attention
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b51d115-5159-4ae8-ad73-38f1c729ed51 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Rethinking Positional Encoding in Language Pre-training
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1859956c-fa26-4f03-acbe-d34f0e8363e9 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video Token Merging for Long-form Video Understanding
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5b48a0b-b677-4c05-b7c6-cca924d5aa4e · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MiniMax-01: Scaling Foundation Models with Lightning Attention
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41e8f07a-f02d-405e-83f9-e0ba620521e1 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Lmms-eval: Accelerating the development of large multimoal models, 2024
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce37d5cf-a163-49c2-bf22-184a2702f662 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LLaVA-OneVision: Easy Visual Task Transfer
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67cf1236-7fa1-4118-8cbe-6711b0775924 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Aria: An Open Multimodal Native Mixture-of-Experts Model
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96efacf1-46f1-4ad7-8ead-ec7d9589b0bb · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Mvbench: A comprehensive multi-modal video under- standing benchmark
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc3df688-87ae-43be-9cbf-90e6559d12ff · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Videomamba: State space model for efficient video understanding
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb4f2bdf-9eaf-4eca-9789-49d5d5680fa3 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Independently recurrent neural network (indrnn): Building a longer and deeper rnn
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d725fa6-89c5-47d4-8899-7fd448e4b2c5 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Mamba- nd: Selective state space modeling for multi-dimensional data
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2efa0987-9240-45d2-aa06-4a09a803cc06 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 427d13d9-fd81-4bd6-88c9-36743cc2950d · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Llama-vid: An image is worth 2 tokens in large language models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29839910-d5e9-4898-a35d-64c2d808633e · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91d85d33-4de6-46c1-9662-ac35ff987eda · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VILA: On Pre-training for Visual Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1268ee51-1cde-4edf-9513-2bd023e36f54 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Visual instruction tuning
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 097793e8-a78c-4584-a32b-f5bb696a9267 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Improved baselines with visual instruction tuning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67cf2745-2a8b-4f86-9da3-44dafe224d18 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee3968c0-a0cf-49e6-afc9-d3ab647abbf9 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Visual instruction tuning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec5d62f4-428d-4497-bbd9-08a8a5b08701 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding E.T. Bench: Towards Open-Ended Event-Level Video-Language Understanding
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 131cfbbc-8a0e-42de-a5d0-d4c2040a6ada · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Snakes and Ladders: Two Steps Up for VideoMamba
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fef4cecf-0169-4367-8e2e-009733dc94b6 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Vista-LLaMA: Reducing Hallucination in Video Language Models via Equal Distance to Visual Tokens
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d58ae8a-7056-4704-8753-f10239e6cc0e · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video Token Sparsification for Efficient Multimodal LLMs in Autonomous Driving
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 08f1d850-86c2-4028-af62-fef6abd66539 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 550f1580-0770-401c-8fdc-f1bf961f5b2d · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 174d41dd-2a7a-4ac1-91c0-ecce84463bed · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Videomamba: Spatio-temporal selective state space model
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c35457c-e2e5-4a8f-8bc9-e1ee253910da · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding RWKV: Reinventing RNNs for the Transformer Era
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 712d3c3f-5c91-418d-b940-6534d53f7b75 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99aa023c-f10a-4c49-99c2-00644696aa80 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding RWKV-7 "Goose" with Expressive Dynamic State Evolution
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9d8c3be-0243-4298-95b9-3ee921ca8923 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VL-Mamba: Exploring State Space Models for Multimodal Learning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 210f87e3-5250-4913-835f-ce77de68d1b6 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91df7b2a-5fb5-4c48-a2a8-6008dc862506 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Various Lengths, Constant Speed: Efficient Language Modeling with Lightning Attention
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b496daaa-7fa5-4104-8fc1-3028415a32f5 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Automated as- sistance for creative writing with an rnn language model
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a387132b-e8ff-4ff4-8820-df3656c28326 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Bidirectional recur- rent neural networks
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16a14370-b6e2-41eb-8cd6-6a4b73f12e0b · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Llava-prumerge: Adaptive token reduc- tion for efficient large multimodal models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 743e4ac4-13eb-4ee1-8f93-8f2fc22b05e8 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Disan: Directional self-attention network for rnn/cnn-free language understanding
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a5837f7-493f-489d-bcdb-c2b0723c4d99 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11789708-fe6c-41c0-b5a3-1c112e649775 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00d8e688-0c33-48cb-ab53-274cf486fbfb · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4ebeb15-4443-4c06-9dcd-0a56d5531c28 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MovieChat+: Question-aware Sparse Memory for Long Video Question Answering
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 605c7da0-e77d-421b-a97a-1546f94eb854 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c64bc55b-2e82-4402-8aae-ec3bec8db202 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Roformer: Enhanced transformer with rotary position embedding, 2021
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f64888b-413f-4e14-ac5b-6f1b0bd7b83a · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Koala: Key frame-conditioned long video-llm
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fdda23a-e379-4a88-8fda-78e5a5fba147 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding DyCoke: Dynamic Compression of Tokens for Fast Video Large Language Models
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f764f96-db39-421c-a41d-58e3607f5b97 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LLaMA: Open and Efficient Foundation Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56dee312-c2e3-4278-9be9-1b72b046e87f · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Attention is all you need
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c8a4b1f-7233-4bcb-9af4-7f93cd16f279 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 86
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0d78680-692b-4287-ba9b-5a160a7bad8c · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding CogVLM: Visual Expert for Pretrained Language Models
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 537224d8-9dba-45f2-902d-a7db50547967 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LVBench: An Extreme Long Video Understanding Benchmark
Reference 88
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 335c0eca-0909-4bfa-ae52-8fd8012f6ea6 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Vatex: A large-scale, high-quality multilingual dataset for video-and-language research
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0c89a3f-5ec9-4803-ab29-2c4119d62c42 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding VideoAgent: Long-form Video Understanding with Large Language Model as Agent
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a46fe6e4-f825-4248-a277-b170d844826d · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Longvlm: Efficient long video under- standing via large language models
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e423a24-2f0a-46a3-8cf4-c44392ca6812 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1f5b7ff-f3f4-47cf-84c7-597dfc837803 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding MotionBank: A Large-scale Video Motion Benchmark with Disentangled Rule-based Annotations
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bd87471-658a-43fb-b255-b133fc4b9a93 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Pllava : Parameter-free llava extension from images to videos for video dense captioning, 2024
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d5a20d2-bc5b-4ea3-b24e-2cd0facb5bf6 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7dcce52-2094-48fc-add4-b66e602eeb3a · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding xgen-mm (blip-3): A family of open large multimodal models
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a52bc8ad-4b85-42e8-bc8f-a01a40173fc7 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Qwen2 Technical Report
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0038c185-3efb-4b91-8d47-5a0d438fc751 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Gated Linear Attention Transformers with Hardware-Efficient Training
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78ba265f-537d-4ff3-9b1d-371127275889 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Gated Delta Networks: Improving Mamba2 with Delta Rule
Reference 99
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b0430c1-0fd4-4b5a-82b4-bf8817005e6e · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Parallelizing linear transformers with the delta rule over sequence length
Reference 100
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8077658-29a0-4f35-a5a3-9d480e1c43d4 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 101
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43514bb3-a1bf-4712-adca-c41d2225b911 · outbound
AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 102
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 908cfaf1-2bcd-40b4-8f85-5f8f63e346e6 · inbound
$M^3-Verse$: A "Spot the Difference" Challenge for Large Multimodal Models AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.