Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:44:50.672180Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 2 inbound Pith citation observations for arXiv:2412.12833.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T13:44:50.672180Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:46:40.436903Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
72 of 72 outbound references displayed
External citation measurements
0
pith, observed 2026-08-05T02:28:24.338817Z
Observation 7deb4e54-a544-43cf-9ec6-4ae0d1de3e4e · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be2706f-a41f-4581-80a9-75c517c53179 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Frozen in time: A joint video and image encoder for end-to-end retrieval
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c21441fc-495e-413e-ade8-ce4e89f3dea1 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Memory Consolidation Enables Long-Context Video Understanding
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c36730f4-6ed6-4fee-be05-11bdb36ab249 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Conceptual 12m: Pushing web-scale image-text pre- training to recognize long-tail visual concepts
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e07e8417-78d9-4864-b5c2-a8077e335d3b · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baf08990-0c88-4e02-a7a9-08728f4048a6 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Xmem: Long- term video object segmentation with an atkinson-shiffrin memory model
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebaf58fb-25de-4edf-af88-ca03f5b24131 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3c85fba-9546-4263-8206-8e8ead4454e2 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3bf6c2e-3ffa-4700-881f-4528e9510437 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering PaLM: Scaling Language Modeling with Pathways
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 459de546-0dde-49ba-8586-8f548a6d0328 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Scaling instruction- finetuned language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 42793921-42a3-4e57-ac57-e495bf7b8f8b · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Instructblip: Towards general- purpose vision-language models with instruction tuning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 39776e8a-ce3f-4fb6-b1a0-b2a823373b02 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Egovqa-an egocentric video question answer- ing benchmark dataset
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ecacce4a-9f99-4463-b9bf-3f7744a82d22 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Eva: Exploring the limits of masked visual representa- tion learning at scale
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c3d9022f-9226-42e2-9a86-01c96d77e172 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering An empirical study of end-to-end video-language transformers with masked vi- sual modeling
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5505f809-a4e5-421c-9ce2-9e63472a52d2 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering The faculty of language: what is it, who has it, and how did it evolve? science, 298(5598):1569–1579, 2002
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7cf23d61-6dcd-4c4e-b4bb-ede4d7c5d4af · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Ma-lmm: Memory-augmented large multimodal model for long-term video understanding
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 58ad06cc-9093-4e80-ba80-1d5356e57f6d · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e83701a-aab5-4b91-ae2d-0fce57c80f23 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Chat-univi: Unified visual representation em- powers large language models with image and video un- derstanding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e49bfb55-3144-4a2f-b33b-59b8d134b026 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Language Repository for Long Video Understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 730ded9d-3710-4b50-a4b5-2f055009d5b3 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering The Kinetics Human Action Video Dataset
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbc6bf7b-8517-413f-9621-03f0cf75958d · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Segment any- thing
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9ccc714-7929-4e32-bee2-724ade01dcd7 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Revealing Single Frame Bias for Video-and-Language Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b289956-d3ba-4c97-b720-a1299016f50c · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f605c85a-e8f9-4e88-8dd3-3541048d8410 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering VideoChat: Chat-Centric Video Understanding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2712700-82b0-4fc0-8d28-6e25ba04c1da · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bca4a85f-1150-4670-9679-34d68d7426aa · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Unmasked teacher: Towards training-efficient video foundation models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 96801e0e-8f50-450f-b028-6908cc30c484 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8446f119-07ab-4432-b5aa-d7509bc90b28 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Tgif: A new dataset and benchmark on animated gif description
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7821931-0fb9-40ea-bf10-beb77dd26734 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Invariant grounding for video question answering
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 9e6bce0e-dd50-408a-aed7-b0a6e70f14f2 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Visual instruction tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70e1b3fc-56d0-42da-9e53-68b7fae9adba · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Image Captioning in news report scenario
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1a3515d-bb66-45f2-ae1a-6a6050f1e808 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b85211e8-21e5-4408-9977-0f9a05d48c3c · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Language Models are Few-Shot Learners
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f8ce990-2b20-43c6-b479-97139a1ae00f · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering ClipCap: CLIP Prefix for Image Captioning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e77a4b0-b44f-41a9-a51e-b03d7320effb · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Introducing chatgpt
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 42b513cb-85b4-4f4f-9b4a-7939a07d9b88 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering DINOv2: Learning Robust Visual Features without Supervision
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5f55249-37c0-45a3-9733-bd3b24ec8439 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering The language instinct: How the mind creates language
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3d4344dc-e917-48a8-9db2-96f9fb52e066 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Streaming Long Video Understanding with Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0f7bcc5-b8db-4543-96b7-a2f3ec1cc135 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Learning transferable visual models from natural language supervi- sion
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4cf9ec3-f3d7-4196-b36c-00a50127b88e · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Understanding Long Videos with Multimodal Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c47738b6-95eb-42db-ba9f-0139b38ed4b8 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Timechat: A time-sensitive multimodal large lan- guage model for long video understanding
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 129f7344-ec23-4378-b47d-8a072dd77879 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Fusecap: Leveraging large language mod- els for enriched fused image captions
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8380ebfa-8481-47a7-ab0d-9282dc5caa99 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Talking about large language models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 77f71f1d-843c-4dfd-b58e-b38000e6fa67 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Moviechat: From dense token to sparse memory for long video understanding
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation bd70c989-4d89-496d-8f9a-fa1479aa1f85 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering DreamFrame: Enhancing Video Understanding via Automatically Generated QA and Style-Consistent Keyframes
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24f70b62-e83e-48ee-a573-1d3d39749a8c · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Koala: Key frame-conditioned long video-llm
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b6df8fd2-1875-498d-968d-5377f094bf1d · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Galactica: A Large Language Model for Science
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f02d7a7-a679-4bfb-ab42-cabd6dd4f453 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering LLaMA: Open and Efficient Foundation Language Models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85b72e60-54db-4ba4-a526-492fda7f5264 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bc027d8-1995-4db6-b25b-3a116ff609c0 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Computing machinery and intelligence
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f2c664a1-9aac-49a4-9801-7c29e34152d5 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Attention is all you need
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d70177a6-aa19-42c0-9f6a-30dd8af9135f · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27daf447-93fd-489f-96da-22433a18bbbc · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering GIT: A Generative Image-to-text Transformer for Vision and Language
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36d93877-eaca-47ca-95b2-7fefe806919c · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 168359ec-92d4-4ff7-be2d-8a08834f87b4 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59c53a00-e26c-4fc3-ad06-d19b9a5aba0a · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering A large cross- modal video retrieval dataset with reading comprehension
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1bce6f61-334e-4adc-b45a-892d1cfe4f3d · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Next-qa: Next phase of question-answering to explaining temporal actions
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 16e2c91b-832e-470d-8db2-dcfbb173dadd · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Video as conditional graph hierarchy for multi-granular question answering
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e4ffe5f4-c9d2-42ec-a1d7-bcd935e56afb · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Video question answer- ing via gradually refined attention over appearance and mo- tion
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ac0a1a6-b20a-4c90-a3ac-a3664442ae4e · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering mplug-2: A modularized multi-modal foundation model across text, image and video
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2645119b-22bc-472e-8c14-06d5518652da · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb13a480-ead0-4234-ab84-56823a11744e · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering VideoCoCa: Video-Text Modeling with Zero-Shot Transfer from Contrastive Captioners
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df907a64-66ec-4748-a479-99be86623005 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Just ask: Learning to answer ques- tions from millions of narrated videos
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e89c310c-625b-48ae-af7f-6b72aac62960 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Zero-shot video question answering via frozen bidirectional language models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c8cbdc7d-43b5-4171-a8d2-750e90176b8d · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7a8556c-9426-4a4b-a846-eb87ab881c96 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Hitea: Hierarchical temporal- aware video-language pre-training
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5613c14b-5b87-482b-88e4-6f3b69eb3f1d · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering CLEVRER: CoLlision Events for Video REpresentation and Reasoning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation add73827-01e9-40e2-a263-df818fe155c7 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Self-chained image-language model for video localization and question answering
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1e15c84f-a3f7-4480-a4c2-9413a1cb7b0d · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Activitynet-qa: A dataset for understanding complex web videos via question answering
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 1dde4ae7-c105-4d57-9d7d-fbfb707794a0 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2b53762-26be-46d9-bf18-8476c458ca29 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering Streaming dense video captioning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8732139d-07c6-4e1d-8a75-29f261e19278 · outbound
FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cb02e7a-cd0c-4d8a-aa94-689115e96931 · inbound
Video-MMLU: A Massive Multi-Discipline Lecture Understanding Benchmark FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e828d07-d481-4c27-b553-2541daef71b7 · inbound
Prompt-Adapter Context Routing for Parameter-Efficient Multi-Shot Long Video Extrapolation FocusChat: Text-guided Long Video Understanding via Spatiotemporal Information Filtering
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.