Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T04:02:43.261543Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 77 of 77 outbound references and 56 inbound Pith citation observations for arXiv:2501.00574.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-18T04:02:43.261543Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.314074Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T16:49:57.265026Z
77 of 77 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5ccd813f-5326-4801-b799-fd04a1a4215e · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Ht-step: Aligning instructional articles with how-to videos
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e2f55cb3-6f44-4e36-bba9-2e1ad7668ada · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Qwen Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 779244d8-d237-4e0e-8210-a9eda9ce2e6d · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d33aac72-3fad-4d6f-9401-54539ca3cf5d · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Frozen in time: A joint video and image encoder for end-to- end retrieval
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0d0606ad-ee11-41ca-8cc8-d55eb4b5395a · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Fuyu- 8b: A multimodal architecture for ai agents
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cd0461f2-3750-45fa-9588-9ae33a149ad8 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Token Merging: Your ViT But Faster
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cbe8f332-5b99-4c76-8ab7-33c1155b89e5 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling HourVideo: 1-Hour Video-Language Understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7311a77c-640e-4ac4-9681-d6e43ef63cfe · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dd31a8fc-f16e-4dcd-bc52-ab91dbd3186a · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Efficient Large Multi-modal Models via Visual Context Compression
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f7fbad86-8795-4021-b276-98d6a6a2e4a5 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fcd427c2-aacc-45cf-8243-6151d72161a6 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Panda-70m: Captioning 70m videos with multiple cross- modality teachers
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation db4489fd-e6fb-47cb-98ff-2ea6d77c59d0 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 600b90e3-8d40-47d6-b370-789f7f70ca11 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2df3594c-4236-4286-87c5-991fdb42bb55 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 23d85641-bcfe-4bb7-b319-6c101585e709 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0afb140e-3b56-496b-8add-35ee60a28ad6 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 09c80962-1b2b-4976-89df-22ce0a95b191 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c4f74501-f968-4cdd-bcb5-ca61ca3d2d20 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Tall: Temporal activity localization via language query
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5a6407dc-e93e-48ee-93ae-b696bfda9549 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Ego4d: Around the world in 3,000 hours of egocentric video
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation afaf9f30-f354-46ec-a3e9-4f62459c9405 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Online Video Understanding: OVBench and VideoChat-Online
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8f02a67c-db75-4ea4-a44e-9ed3ff269651 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video recap: Recursive captioning of hour-long videos
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f41c7ea8-baff-483e-98b6-afeb880af81c · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling MiraData: A Large-Scale Video Dataset with Long Durations and Structured Captions
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 77257d2f-6ebe-424c-99c5-ee5a0d83ca59 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling The Kinetics Human Action Video Dataset
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b3a99a9e-d760-48bd-9bcf-c4f21aa785d9 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling OtterHD: A High-Resolution Multi-modality Model
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 06e36acb-10d6-4385-93c0-b600a5db02d5 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LLaVA-OneVision: Easy Visual Task Transfer
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 82223398-adaf-490f-b246-cc2e9b3989d8 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7c3bf99c-f3e6-405f-a78d-f220407e7788 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling VideoChat: Chat-Centric Video Understanding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 05f1833f-7345-4679-abda-3b63faa6ddcb · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Unmasked teacher: Towards training-efficient video foundation models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 39366faf-09cb-4205-8c62-99262af4df53 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Mvbench: A comprehensive multi-modal video understanding benchmark
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6a3411f0-f35b-4861-8d54-10f3e21d3b57 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Llama-vid: An image is worth 2 tokens in large language models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e03994eb-65ed-4c9b-acc0-f9a966a39006 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-llava: Learning united visual repre- sentation by alignment before projection
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b19cb18b-dc47-4443-aa35-f12b1b8382c4 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Microsoft coco: Common objects in context
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ca620c43-9dda-49f8-8438-ded2f33821b0 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Visual instruction tuning
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fc2efdfa-89c7-4199-81f7-5ccd0e43f5b6 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Kangaroo: A Powerful Video-Language Model Supporting Long-context Video Input
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fb0f0cfd-d02a-46eb-a989-78fc83bdc244 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling VideoGPT+: Integrating Image and Video Encoders for Enhanced Video Understanding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0dd1d071-ab76-4ad0-b384-85327ae7f995 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b450f8da-bc02-4493-9f13-e6e97c94b1eb · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 106110a7-88ec-446f-8e39-e87aa3da748a · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Spoken moments: Learning joint audio-visual representations from video descriptions
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a653d899-4963-46a4-9e17-1233011a2c27 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling GPT-4 Technical Report
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f899955c-8817-46e0-8d7e-26390ecd5e8c · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9374501d-dfe7-4cda-b36b-7a36fb7f9513 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Perception test: A diagnostic benchmark for multimodal video models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b843d29b-a458-4567-8ccd-fd94e0d8e4d6 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f9b9e478-9dbf-4b84-83e1-faebb67969f9 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling CinePile: A Long Video Question Answering Dataset and Benchmark
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 00d5b85c-c553-4069-8408-12eac93fca39 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e74b5432-69be-4af7-8dd0-32eccc895a00 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Timechat: A time-sensitive multimodal large lan- guage model for long video understanding
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 69a01eff-aa45-4564-9c4f-28685cf5b52c · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Sharegemini: Scaling up video caption data for multi- modal large language models
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 424c255a-4d59-447f-87fa-c36613e5708f · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 731ff6fb-eb7b-48c7-b925-bcf2ff8d0414 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ca6154b8-b0b9-4640-9600-6a93a4c81256 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Moviechat: From dense token to sparse memory for long video understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 450a0c10-27fe-488c-addc-f2015cf948b4 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Koala: Key frame-conditioned long video-llm
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a6425a9b-0607-4e84-91d1-a274210e554b · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5f8b2268-7e4a-457a-8f70-84c0d0615742 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5760c3b7-d10b-43f6-85ce-d9520ab46306 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LVBench: An Extreme Long Video Understanding Benchmark
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d47e1344-6c36-4fc0-b57d-eef20f1ac2e4 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9489c8d4-380e-4e33-aee0-489b66a22283 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling InternVideo: General Video Foundation Models via Generative and Discriminative Learning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 600ba87b-ff0d-4e3e-a1e7-027435d4e235 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Internvideo2: Scaling video foundation models for multimodal video understanding
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ba361a7a-a74e-498e-8ba7-1aab43771bf1 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Visual Context Window Extension: A New Perspective for Long Video Understanding
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ac284730-eee2-40f8-ab43-a6799a557d74 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Longvlm: Efficient long video understand- ing via large language models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4de7f155-5d02-4109-9a9a-0c302c4cdc64 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6feda575-e467-4932-b44d-2da9d80ef809 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a100e30c-9c51-4c45-9148-50e2b78a265e · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b9332801-9f36-4c56-b77f-a8364c0f149c · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Advanc- ing high-resolution video-language representation with large- scale video transcriptions
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5cfa7a2c-3401-4e26-b69e-c5672311255c · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Vript: A Video Is Worth Thousands of Words
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 97ab7361-9b0b-40e7-9f20-7f65a83d98de · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling TimeSuite: Improving MLLMs for Long Video Understanding via Grounded Tuning
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6d06eef6-2d31-4fd3-a6bc-2736bdb5bb57 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Sigmoid loss for language image pre-training
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 47e71157-dafe-4069-b911-66e8c6f76dcd · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 428835ff-04fc-425d-834c-bd2a36c4e492 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LvBench: A Benchmark for Long-form Video Understanding with Versatile Multi-modal Question Answering
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e272bae0-47ed-44e8-9638-54b01db80967 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Long Context Transfer from Language to Vision
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 869fdb9a-b101-4c3a-b5bd-5fdbd436cb60 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 53a85dce-09a4-4b2a-b5c5-1a81b15478ab · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Llava- next: A strong zero-shot video understanding model
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 84022ade-c828-472f-a4f0-b9028aa05dfa · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 93d93d03-bac4-4091-959d-50669fe75564 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMs
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9d9007bb-eb1d-400d-ad90-cecca5f3ca5a · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling MLVU: Benchmarking Multi-task Long Video Understanding
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ed625540-0776-4336-b00a-cf1276dcd26e · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Visual Dropout in LLM Visual token redundancy in LLM inference
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0cdb5f83-0a33-42e7-abbc-9a728e983447 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling Video-Language Connectors As shown in Fig
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fb7abffb-5425-4f12-8bfd-eaa32f4854c0 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling We provide details of the data construc- tion pipeline for each dataset as follows
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d0c7ab25-93cb-4bd3-9c63-697861feefa1 · outbound
VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling 11 and 12) and long video understanding ( Figs
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 59801fb1-50fb-4317-8a32-1acef7e4adef · inbound
Time-R1: Post-Training Large Vision Language Model for Temporal Video Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 304f2a4f-3440-41b9-a78a-5c5e63b912ff · inbound
VideoChat-R1: Enhancing Spatio-Temporal Perception via Reinforcement Fine-Tuning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0e55e230-5270-4ea2-ba9d-10a36c347294 · inbound
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cbb5882-cc3d-430f-b987-5b744e65f3ef · inbound
Vad-R1: Towards Video Anomaly Reasoning via Perception-to-Cognition Chain-of-Thought VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1746ded8-6795-49d0-871f-472260bf428f · inbound
AV-Reasoner: Improving and Benchmarking Clue-Grounded Audio-Visual Counting for MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdef7f07-c1c2-49ed-8e95-b25717da792e · inbound
Flattery in Motion: Benchmarking and Analyzing Sycophancy in Video-LLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 464c5e5e-07f5-46f0-ae02-fcf6e2b9ff09 · inbound
VRBench: A Benchmark for Multi-Step Reasoning in Long Narrative Videos VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdf7c58a-339b-43ed-9661-ce81ff93e999 · inbound
Video-XL-2: Towards Very Long-Video Understanding Through Task-Aware KV Sparsification VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 632099fc-9a4c-47e4-b513-325dddaaf3e7 · inbound
AVC-DPO: Aligned Video Captioning via Direct Preference Optimization VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97ac225c-a44b-48e7-b367-d1e266ecfdf5 · inbound
Infinite Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 215cecf1-67bb-41ca-83fc-6a30a3038aa5 · inbound
VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ceb20892-77c2-4ab6-a4b0-7359bf440543 · inbound
TAR: Temporal Anchor-Constrained Reasoning for Video Temporal Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7022ef98-3fa6-4b6c-b519-e0a1146e13de · inbound
Cambrian-S: Towards Spatial Supersensing in Video VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 001a86da-c99c-44df-80e1-33be32c08153 · inbound
Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8d31a43-6c44-42a5-9ac9-39649bcdf1c2 · inbound
EchoingPixels: Aliasing-Resistant Joint Token Reduction for Audio-Visual LLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b10fd8d-e836-401e-8b70-c49303de9192 · inbound
Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5f773962-080e-4f81-b5a2-6269198da281 · inbound
VideoThinker: Building Agentic VideoLLMs with LLM-Guided Tool Reasoning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation caf712ff-39f2-405b-b02d-3c7e44b8e3ab · inbound
LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 755d7aae-8794-410d-9ab5-18f342f58639 · inbound
Diagnosing Long-Video Quantitative Reasoning in Multimodal LLMs via Enumeration and Counting VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bf6b2fc-d160-4eda-b8b9-fa7e050e45f8 · inbound
Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6f1f7f69-f505-4b72-878e-8402334e5135 · inbound
AdaSpark: Adaptive Sparsity for Efficient Long-Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e2f84104-be0e-42d7-a0af-20076b606867 · inbound
Small Vision-Language Models are Smart Compressors for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 67be4ce6-9f37-4586-b1ca-299ab7ba4e77 · inbound
POINTS-Long: Adaptive Dual-Mode Visual Reasoning in MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation eb49f453-1f16-4af0-8773-198ffa504596 · inbound
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2f95654f-aee8-4488-b844-0c8567bbfff0 · inbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d3a281b8-a4c6-4ec7-9e22-7b5e6f809ea1 · inbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1318ce42-fc44-483c-a7c7-eb48232047fa · inbound
OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ce784b1c-a6ce-413a-ac62-537513023795 · inbound
VisMMOE: Exploiting Visual-Expert Affinity for Efficient Visual-Language MoE Offloading VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 52e12638-e132-44a5-8b65-f8bfc00e0044 · inbound
MedHorizon: Towards Long-context Medical Video Understanding in the Wild VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4adeab72-40fb-4e0b-b7e2-130439272b81 · inbound
TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6ec7e470-e915-41ae-adcc-28dc93d42da5 · inbound
TOC-Bench: A Temporal Object Consistency Benchmark for Video Large Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 66fdcaa3-4c5c-4457-9147-816527f507fb · inbound
EvoGround: Self-Evolving Video Agents for Video Temporal Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fe731a24-2731-4709-a6b4-008686feedc5 · inbound
StreamPro: From Reactive Perception to Proactive Decision-Making in Streaming Video VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ec812b9c-1654-49bd-9cb1-e7a4964b0231 · inbound
PyraVid: Hierarchical Multimodal Memory for Long-Horizon Video Reasoning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f2146df8-67d7-4947-ac5f-ccda5531421b · inbound
MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fb4fd70d-585c-48b0-9d6e-22d13ada3ef0 · inbound
LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fd7d14f7-d3f6-4de2-91e8-9d9a6a2997d2 · inbound
Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 43226218-c24d-4428-810e-2184c620bc5c · inbound
Linear Scaling Video VLMs for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7c215330-0ec7-4050-a3a4-90e738110d57 · inbound
AdaCodec: A Predictive Visual Code for Video MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ce149296-1715-4fee-8bf7-78373f24a336 · inbound
Towards One-to-Many Temporal Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 9be2b6a0-608e-4891-b064-1e0a1a0a2cf6 · inbound
StoryVideoQA: Scaling Deep Video Understanding with a Large-Scale, Multi-Genre and Auto-Generated Dataset VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a32cd9a1-2731-47e8-8f02-3a0d2d0a4957 · inbound
GOPAgen: Motion-Aware and Efficient Agentic Long-Video Understanding with Structural Memory and Hierarchical Reasoning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ecee5536-b925-44dd-9aa9-f5e695be6697 · inbound
MotionEnhancer: Leveraging Video Diffusion for Motion-Enhanced Vision-Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fa29733a-7da2-4aac-bdbf-d5e70755b7cb · inbound
See More, Think Deeper: Query-Expanded Visual Evidence and Answer-Clue Guided Reflection for Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4b1a77c9-1b4d-496e-9916-5f4f058923b4 · inbound
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 168
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a3eb1385-31d6-423c-b654-d54bef28fc13 · inbound
ViCoStream: Streaming VideoLLMs Can Run Beyond 100 FPS with Stage-Wise Coordinated Inference VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2df7dcd3-e675-474f-858a-cee988064085 · inbound
AVOC: Enhancing Hour-Level Audio-Video Understanding in Omni-Modal LLMs via Retrieval-Inspired Token Compression VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8ecf9277-e14b-4f6e-a53e-1cb3ff4a078f · inbound
Learning to Deny: Action Denial in Multimodal Large Language Models VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cde95483-51b5-46fa-9610-b56e3f11fa18 · inbound
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9641bb0-bc0e-4aa9-8d06-47f68c5dbe93 · inbound
DynTrace: Tracking Dynamic Object Evidence for 4D Spatio-Temporal Reasoning in MLLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7e61b8a-07a7-4a88-98f4-a78be078cb89 · inbound
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7669212-e82e-4bbf-b512-faae1e162957 · inbound
TimeLens2: Generalist Video Temporal Grounding with Multimodal LLMs VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3bd2182-e873-4443-8b8e-d01840d6f635 · inbound
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3b7f910-a70e-4475-b4e6-7f994cd1d712 · inbound
Long-Horizon Embodied Decision-Making via Multimodal Memory Compression VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98abc67f-d45e-4fbe-9411-15972855c281 · inbound
CAVE: Competence-Aware Visual Boundary Evidence Alignment for Video Temporal Grounding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8651f988-a6eb-478d-9397-9933409b53f1 · inbound
Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.