Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T18:25:21.621268Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 84 of 84 outbound references and 2 inbound Pith citation observations for arXiv:2603.01400.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-15T18:25:21.621268Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-13T05:49:36.807884Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-13T05:52:22.302025Z
84 of 84 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7c10f71f-402b-4d28-871f-b86f7a71bebe · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1253536a-4a15-4c1e-8aee-7b006f87ad0f · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1055f925-b328-461c-b4ae-ae9e7c4202ec · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f8cb1e64-111b-4df6-b4c1-364fc4cc487a · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Hired: Attention-guided token dropping for efficient inference of high-resolution vision-language models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1974daac-f804-45c3-bfe7-fdf8eed4ac00 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Qwen2.5-VL Technical Report
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 69462350-8708-4509-97f6-41300568aaa4 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Token Merging: Your ViT But Faster
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 307b969c-1c54-4663-83d5-2209560690f9 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models AuroraCap: Efficient, Performant Video Detailed Captioning and a New Benchmark
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ba2e3575-e5bf-49a5-ac5c-d197410f1676 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Sharegpt4video: Improving video understand- ing and generation with better captions
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd15f5b5-a452-4350-a714-280ea226f9e8 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4e4871fc-92c5-47fb-bab8-9107a70e4b50 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites.Science China Information Sciences, 67(12):220101
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5df7632b-e571-4beb-bf7b-51967f4adc3a · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ee3ff07b-972b-45f9-bd2a-785c6d92e094 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.See https://vicuna
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cd2cbf18-b934-4781-a422-e2a96b0501eb · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Sinkhorn distances: Lightspeed computation of optimal transport
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e9f6adb-3187-4cd9-92eb-39f2b6bde528 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 46372f6a-4618-4cca-b194-81b6f406c0a5 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Video-mme: The first-ever comprehensive evaluation benchmark of multi-modal llms in video analysis
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2b099df-6088-4428-8865-d4a635427214 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 352f6555-5169-49e5-ab23-c28f34d52a3b · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models PruneVid: Visual Token Pruning for Efficient Video Large Language Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dff2baec-d180-4d96-9c24-da12c23c0a9b · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Chat-univi: Unified visual representation em- powers large language models with image and video under- standing
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4478dee3-eee3-4428-a230-199ad501ce69 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Sparsevila: Decoupling visual sparsity for efficient vlm inference
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3dd5e7e3-2fbb-44dd-af9b-fa21596f9abb · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models arXiv preprint arXiv:2505.18227 , year=
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 428ecb9f-3562-4fb6-8fa0-dfa92f52b5a5 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Lmms-eval: Accelerating the develop- ment of large multimoal models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8865f885-a212-4fcb-8ffe-17891dd57711 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 686986fc-58d7-4137-9653-ee300307b388 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Expansion and shrinkage of localization for weakly- supervised semantic segmentation.NeurIPS, 35:16037– 16051
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a2ef54d3-0c89-4a57-8c38-d213e13a7912 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 21ee01fe-f803-4313-ba9f-37ff6aef794d · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Cross-modal and uncertainty-aware agglomeration for open- vocabulary 3d scene understanding
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dfa922e3-a8f2-4bc4-9a88-33451116a8f3 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Orthogonal Projection Subspace to Aggregate Online Prior-knowledge for Continual Test-time Adaptation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b81222d6-1c0c-456b-a895-5978774aac55 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models VideoChat: Chat-Centric Video Understanding
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3c1fd37-8662-4bbe-95d9-8e16c98864c8 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Mvbench: A comprehensive multi-modal video understand- ing benchmark
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0ca99f3f-3ac6-469b-b433-50a8ad459118 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Llama-vid: An image is worth 2 tokens in large language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1cc47a42-5cb6-4f36-90de-b608a721c1cc · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50387c75-1d4d-413c-9df0-f1c51e32d1cc · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Vila: On pre-training for visual language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 97643793-6fb5-42a3-85e4-f8ab4c8720ae · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Improved baselines with visual instruction tuning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 333756eb-327e-4182-94d1-986f82bd549f · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Llavanext: Improved reasoning, ocr, and world knowledge
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c04c2190-6cf4-4248-a223-772acb2fbd56 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Multi-Stage Vision Token Dropping: Towards Efficient Multimodal Large Language Model
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f78afe71-546c-4bf0-832f-bfe4925584b2 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Less: Label-efficient and single-stage referring 3d instance segmentation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd783566-e8cb-4ea9-8c0f-855c4dcb10c8 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Hybrid-level instruction injection for video token com- pression in multi-modal large language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4b253035-c0cc-4c9a-ad00-75b558830e2b · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Nvila: Efficient frontier visual language models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 119419f0-a3c9-4320-aa07-95ed7de6e61f · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8d0dbef-36dc-4286-85f0-503ff62da652 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Egoschema: A diagnostic benchmark for very long- form video language understanding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f2727595-b486-495a-9b65-2650249c2315 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Perla: Perceptive 3d language assistant
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 08caf57a-46e1-4c13-a6bc-39a19e2ebb52 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models M ´emoire sur la th ´eorie des d ´eblais et des remblais.Mem
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 87e95e8c-7ad4-4f95-af23-b688b1630d64 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models T2td: Text-3d generation model based on prior knowledge guidance.IEEE Transactions on Pattern Analysis and Machine Intelligence, 47(1):172–189
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fad83d17-fa96-4fb7-add5-f142e4908837 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Learn- ing transferable visual models from natural language super- vision
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1ae580b-a2b5-4911-9389-01680269bb93 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Llava-prumerge: Adaptive token reduction for efficient large multimodal models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd723006-e8a7-4702-afa8-a3c10793bff8 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models arXiv preprint arXiv:2505.21334 , year=
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6373ec4-0774-4ee1-8c9a-40bb5e22f222 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models TempMe: Video Temporal Token Merging for Efficient Text-Video Retrieval
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6bbf309f-ade3-41ff-a155-ff5a2fe005b9 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Fastvid: Dynamic density pruning for fast video large language models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9bc14500-af25-4c75-9f31-cda3689159b6 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bfa32464-5aad-4bc0-9041-c8da6c5b5c72 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Moviechat: From dense token to sparse memory for long video understanding
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 989474e8-b56d-4783-82f1-5a2c5cb5ac92 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models TokenCarve: Information-Preserving Visual Token Compression in Multimodal Large Language Models
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7588e5c2-9c25-4a18-8251-fd0d9493dfd5 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Dycoke: Dynamic compression of tokens for fast video large language models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 52a2c208-bf56-4fd8-a039-25a8e104e1ea · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Stanford alpaca: An instruction-following llama model
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13e0de41-33bf-4729-85de-bf7be3322db2 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Gemini: A Family of Highly Capable Multimodal Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 27654a1b-7146-4301-a713-b381a7990ea8 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 31b238ec-1cb4-4325-8768-9d21905c70d2 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Introduction to optimal transport.Notes of Course at University of Cambridge, 3
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d5e2a7a3-8853-4cc8-93b2-0d7d378f2015 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Springer
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 41ed44ae-c1df-4774-9b91-4580736da2aa · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Ross3d: Re- constructive visual instruction tuning with 3d-awareness
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5763ab29-2d37-442c-86c7-bfe7cc13504d · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e54f96a9-a8cd-4e1e-8374-b0af4b18f2b0 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Uvmap-id: A controllable and personalized uv map generative model
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fa469ff6-c5a0-423b-a471-560860c4103d · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Stop Looking for Important Tokens in Multimodal Language Models: Duplication Matters More
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 936eb09b-3111-47d2-a29e-0e6f8d874e28 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Longvlm: Efficient long video understand- ing via large language models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0233a267-69d2-4fe4-8eb5-4edf7cdeec25 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Visual ChatGPT: Talking, Drawing and Editing with Visual Foundation Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5f3a305-0905-4ffe-9d69-5e604be1e28e · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Longvideobench: A benchmark for long-context interleaved video-language understanding
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2037e54b-31c8-416e-b5fc-ae29cfee9e3a · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc8c9f21-d6f1-439f-9ab8-4caacc9ff7d8 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Conical visual concentration for efficient large vision-language models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6c7c6dd1-3510-4128-a3f6-85ee4a6f7e05 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models PLLaVA : Parameter-free LLaVA Extension from Images to Videos for Video Dense Captioning
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 910a5ad1-c5bd-4f2a-bf96-d8646a932163 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Topv: Compatible token pruning with infer- ence time optimization for fast and low-memory multimodal vision language model
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6bb79c97-e5cc-467c-9eff-f274ef67b8d8 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Visionzip: Longer is better but not necessary in vision language models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9b795a22-c186-4d82-9ceb-07e2ef51ca43 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Atp-llava: Adaptive token pruning for large vision language models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 497c6d8d-8b7b-422b-a3cc-707d95f77df8 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Video question answering with prior knowledge and object-sensitive learning.IEEE Transactions on Image Processing, 31:5936–5948
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19451f1f-e827-4dee-9245-d8da96bf0313 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Sigmoid loss for language image pre-training
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7accfd02-4002-4099-bac7-57c4f08f4157 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models arXiv preprint arXiv:2505.22654 , year=
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1fa2b1de-1857-4a5b-ae05-2bf9db620fb5 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation da801dc0-aab0-4ce9-ad44-52e5365121a7 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Omnicharacter: Towards immersive role- playing agents with seamless speech-language personality interaction
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d6442c55-87c2-4295-8a0d-5f0a1c26eeb9 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Text-video re- trieval with global-local semantic consistent learning.IEEE Transactions on Image Processing
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0360cb4f-239e-43ce-a38b-344b429992d2 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Lmms-eval: Re- ality check on the evaluation of large multimodal models
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d743b53-706a-4d33-b0e2-5707633b4264 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models [cls] attention is all you need for training-free visual token pruning: Make vlm inference faster.arXiv e- prints, pages arXiv–2412
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d510b01-19d4-4d2f-8829-9cf4b29286e2 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3630be8a-844e-43a2-a638-af0d7ff4f7fd · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Llava- next: A strong zero-shot video understanding model
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 29b998a0-8613-4a1a-a67d-e000e269443c · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32f9a511-da57-4916-ac08-9d42fb29f604 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Video-3d llm: Learning position-aware video representation for 3d scene understanding
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf521ba6-1604-4758-b59a-e5446d40d410 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models LLaVA-3D: A Simple yet Effective Pathway to Empowering LMMs with 3D-awareness
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0c57d810-ae1e-4d7e-81fc-2223d5a8a1ef · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 07e72e4f-10d8-46c3-b257-6b3bfdb52187 · outbound
Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models Apollo: An explo- ration of video understanding in large multimodal models
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4204f00-7efd-4621-a236-4d9caf0f85ab · inbound
PoInit-of-View: Poisoning Initialization of Views Transfers Across Multiple 3D Reconstruction Systems Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e38e845a-639f-486f-8083-1956331cd973 · inbound
OTT-Vid: Optimal Transport Temporal Token Compression for Video Large Language Models Token Reduction via Local and Global Contexts Optimization for Efficient Video Large Language Models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.