Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T07:42:22.478647Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 100 of 147 outbound references and 43 inbound Pith citation observations for arXiv:2412.04468.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-23T07:42:22.478647Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T04:28:26.603246Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-04T16:39:58.306201Z
100 of 147 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c658f817-5680-4c89-9a34-5300500b094f · outbound
NVILA: Efficient Frontier Visual Language Models Visual Instruction Tuning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7e402fbf-49ce-4bb6-bc8d-d1fa3b554776 · outbound
NVILA: Efficient Frontier Visual Language Models VILA: On Pre- training for Visual Language Models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 21618f3b-b038-47eb-bf36-c96871706eaf · outbound
NVILA: Efficient Frontier Visual Language Models InternVL: Scal- ing up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation de0a1d31-fe8c-4354-aea6-046e16d7fbde · outbound
NVILA: Efficient Frontier Visual Language Models LLaVA-OneVision: Easy Visual Task Transfer
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation eb365286-d69a-4951-ae58-926ac93b665f · outbound
NVILA: Efficient Frontier Visual Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d9835b24-3200-4f31-b79c-68f7dd2c26a0 · outbound
NVILA: Efficient Frontier Visual Language Models RT-1: Robotics Transformer for Real-World Control at Scale
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 07340e2c-05a6-4030-b44a-9a3512e88a75 · outbound
NVILA: Efficient Frontier Visual Language Models NaVid: Video- based VLM Plans the Next Step for Vision-and- Language Navigation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fac3d417-f91f-4a1e-a41c-8d39f05ae5cc · outbound
NVILA: Efficient Frontier Visual Language Models NaVILA: Legged Robot Vision-Language-Action Model for Navigation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2c051837-21f5-4308-9aa1-12d9722b9c3e · outbound
NVILA: Efficient Frontier Visual Language Models DriveVLM: The Con- vergence of Autonomous Driving and Large Vision- Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 065bccfa-ae10-4737-8662-e41571af98e6 · outbound
NVILA: Efficient Frontier Visual Language Models Capabilities of Gemini Models in Medicine
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation de87e671-2e4d-4750-a93a-b575a91f387c · outbound
NVILA: Efficient Frontier Visual Language Models VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f0305393-6f15-4e35-9c92-1e8b21d1ea90 · outbound
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f436c905-dcd4-44ca-a083-c9d1fc291fdf · outbound
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0e945f66-600a-4f9c-ab60-a6e9bf05e738 · outbound
NVILA: Efficient Frontier Visual Language Models Sigmoid Loss for Language Image Pre-Training
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e93f96ae-481e-4e62-a3b8-71c89280d9a2 · outbound
NVILA: Efficient Frontier Visual Language Models Qwen2 Technical Report
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 72e248a8-940f-4bcb-a4a7-fa82bbfa9974 · outbound
NVILA: Efficient Frontier Visual Language Models When Do We Not Need Larger Vision Models? InEuropean Conference on Com- puter Vision (ECCV)
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation dfc1914d-df12-4c14-8922-a08385485d80 · outbound
NVILA: Efficient Frontier Visual Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c392ca64-216d-42c5-92cc-b81534e786c0 · outbound
NVILA: Efficient Frontier Visual Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ea14d6be-01bf-48f9-b97d-92803853f4b6 · outbound
NVILA: Efficient Frontier Visual Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a7ac950e-dd95-4f91-bdc0-6b0feafac78c · outbound
NVILA: Efficient Frontier Visual Language Models Tem- poral Segment Networks: Towards Good Practices for Deep Action Recognition
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 838def6e-7b0d-40da-a2fb-3d8bc4eaccd3 · outbound
NVILA: Efficient Frontier Visual Language Models Cambrian-1: A Fully Open, Vision- Centric Exploration of Multimodal LLMs
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cf2cddb9-dd6a-4ef2-b297-79d621f7bdb9 · outbound
NVILA: Efficient Frontier Visual Language Models What Matters When Building Vision-Language Models? In Conference on Neural Information Processing Systems (NeurIPS)
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f5600f71-53cf-473b-b204-887322d50fef · outbound
NVILA: Efficient Frontier Visual Language Models Selec- tion via Proxy: Efficient Data Selection for Deep Learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1163e972-deef-40b9-9123-c9e0e6067e82 · outbound
NVILA: Efficient Frontier Visual Language Models Demystifying CLIP Data
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0e70f842-e46f-40ce-ba63-ef807efa0db2 · outbound
NVILA: Efficient Frontier Visual Language Models SemDeDup: Data-efficient learning at web-scale through semantic deduplication
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 844471cb-a865-48c8-8d4b-d4904b48406a · outbound
NVILA: Efficient Frontier Visual Language Models D4: Improving LLM Pretraining via Document De-Duplication and Diversification
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5d6076ce-9d5b-442e-af8a-cd03b61d2312 · outbound
NVILA: Efficient Frontier Visual Language Models LESS: Select- ing Influential Data for Targeted Instruction Tuning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 95b2051f-5b18-4279-8f96-93f0b70a0724 · outbound
NVILA: Efficient Frontier Visual Language Models Data Selection via Optimal Control for Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3ec15776-1f00-489e-bcf4-bdb56a9bca33 · outbound
NVILA: Efficient Frontier Visual Language Models MiniPLM: Knowledge Distillation for Pre-Training Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b688c074-fbc0-45fe-a1fa-7b85d8d9a822 · outbound
NVILA: Efficient Frontier Visual Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 81739734-a2c9-4894-82f2-ef951d1da7b5 · outbound
NVILA: Efficient Frontier Visual Language Models Mixed Precision Training
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 50250526-4ebc-48ec-b0e4-66fe55ab60a8 · outbound
NVILA: Efficient Frontier Visual Language Models A Study of BFLOAT16 for Deep Learning Training
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 763e5068-7994-42db-ac48-6d78988c0002 · outbound
NVILA: Efficient Frontier Visual Language Models FP8-LM: Training FP8 Large Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8115b9a9-7792-413d-b96d-27780fc30645 · outbound
NVILA: Efficient Frontier Visual Language Models COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 3a9ec0b3-8fb2-4a21-9f76-98ed7625c5df · outbound
NVILA: Efficient Frontier Visual Language Models Liger Kernel: Efficient Triton Kernels for LLM Training
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6399bd8d-8383-441f-8e6a-2203a2c28312 · outbound
NVILA: Efficient Frontier Visual Language Models Android in the Zoo: Chain-of-Action-Thought for GUI Agents
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2829ea84-aa87-4322-a0bc-44a063cfa7c9 · outbound
NVILA: Efficient Frontier Visual Language Models ALFRED: A 14 NVILA: Efficient Frontier Visual Language Models Benchmark for Interpreting Grounded Instructions for Everyday Tasks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 94f0944f-bc99-40b0-aa59-d3f6a4581304 · outbound
NVILA: Efficient Frontier Visual Language Models nuScenes: A multimodal dataset for au- tonomous driving
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c3b1dec3-33db-45ab-9e43-2308ea293ca0 · outbound
NVILA: Efficient Frontier Visual Language Models PathVQA: 30000+ Questions for Medical Visual Question Answering
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9599a630-956d-4f7c-96e2-5c193a50c1c5 · outbound
NVILA: Efficient Frontier Visual Language Models Widget Captioning: Generat- ing Natural Language Description for Mobile User Interface Elements
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation df564449-1736-41d3-b567-f5585c43199e · outbound
NVILA: Efficient Frontier Visual Language Models AWQ: Activation-Aware Weight Quantization for On-Device LLM Compression and Acceleration
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4e818683-0073-4a39-bf6f-04a812287ba4 · outbound
NVILA: Efficient Frontier Visual Language Models PyTorch: An Imperative Style, High-Performance Deep Learn- ing Library
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 362f27c4-387b-4acc-8d88-8844a1090d22 · outbound
NVILA: Efficient Frontier Visual Language Models Unresolved cited work
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7299f300-cadb-42a6-b99d-47f60602afe6 · outbound
NVILA: Efficient Frontier Visual Language Models Transform- ers: State-of-the-Art Natural Language Processing
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bdf9982e-289a-404b-86f0-a773dde55f94 · outbound
NVILA: Efficient Frontier Visual Language Models DeepSpeed: System Optimiza- tions Enable Training Deep Learning Models with Over 100 Billion Parameters
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ebea230e-88ce-43fa-b5ef-82970b2c7fd7 · outbound
NVILA: Efficient Frontier Visual Language Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1699b915-1da3-4822-af0a-c2cb67256e73 · outbound
NVILA: Efficient Frontier Visual Language Models A Diagram is Worth a Dozen Images
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4e981fcc-98c0-4420-8116-b1b00f50cd7a · outbound
NVILA: Efficient Frontier Visual Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 665a079b-3d1f-4a86-bddc-bad24504bb99 · outbound
NVILA: Efficient Frontier Visual Language Models DocVQA: A Dataset for VQA on Document Images
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1025ed41-a7cd-4606-a99f-fcf192a0416c · outbound
NVILA: Efficient Frontier Visual Language Models InfographicVQA
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 23ecab5f-374f-4d4e-9d6f-7f423bb37078 · outbound
NVILA: Efficient Frontier Visual Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e89a03d8-0081-4be5-a23f-568cb384996c · outbound
NVILA: Efficient Frontier Visual Language Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Ex- pert AGI
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 888533a0-943d-4b67-a934-79718fc8fb21 · outbound
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c3d8f52c-a31e-422b-92a7-40587bf421ca · outbound
NVILA: Efficient Frontier Visual Language Models SEED-Bench: Bench- marking Multimodal LLMs with Generative Com- prehension
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c3c9bccb-336b-44eb-b7ab-d0bfee4ba589 · outbound
NVILA: Efficient Frontier Visual Language Models Towards VQA Models That Can Read
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation df3b8aaf-47a1-4696-bf74-5d56706a41d0 · outbound
NVILA: Efficient Frontier Visual Language Models Making the V in VQA Matter: Elevating the Role of Image Un- derstanding in Visual Question Answering
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a5d6d94b-0c36-40cd-b6ad-05d9f632e7d8 · outbound
NVILA: Efficient Frontier Visual Language Models ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6d4e6973-97bc-4567-a4d3-cfd190ff6f8f · outbound
NVILA: Efficient Frontier Visual Language Models MLVU: Benchmarking Multi-task Long Video Understanding
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 055794bc-c8ac-499c-8a8b-b2327ebca634 · outbound
NVILA: Efficient Frontier Visual Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6c74f650-7f1a-489a-9a26-33bb88593b47 · outbound
NVILA: Efficient Frontier Visual Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1cf91e4a-058b-46e0-825a-d4ae68ca3742 · outbound
NVILA: Efficient Frontier Visual Language Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e033c718-bf9c-44dd-9d21-9a460348fedc · outbound
NVILA: Efficient Frontier Visual Language Models LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e0925f8d-efec-4110-8251-c6b1daa0f532 · outbound
NVILA: Efficient Frontier Visual Language Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation edab20d8-b0ce-4df3-a0ff-1c2ad6558c6e · outbound
NVILA: Efficient Frontier Visual Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c5ee4a49-5d12-45af-8124-3dee8cec76f9 · outbound
NVILA: Efficient Frontier Visual Language Models Beyond the Nav- Graph: Vision-and-Language Navigation in Con- tinuous Environments
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 89292ad6-ecc1-4c16-a690-03ea6392567f · outbound
NVILA: Efficient Frontier Visual Language Models Vision- and-Language Navigation: Interpreting visually- grounded navigation instructions in real environ- ments
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 77b1c0f1-b55b-4e80-977e-0c679e5b3e33 · outbound
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation da30d0df-3a48-4caa-b0c8-7788deffc3d8 · outbound
NVILA: Efficient Frontier Visual Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation edd9bf7d-9161-4a80-b075-6286672d5cb8 · outbound
NVILA: Efficient Frontier Visual Language Models Gemini: A Family of Highly Capable Multimodal Models
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ddb48495-72ea-4dd7-a51c-653c855f8292 · outbound
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7685e922-b9cd-486f-8d3f-53beeaf886d3 · outbound
NVILA: Efficient Frontier Visual Language Models MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1c284017-9be1-4db1-aa96-a17c583a5a8c · outbound
NVILA: Efficient Frontier Visual Language Models NeMo: a toolkit for building AI applications using Neural Modules
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7dfa2d0c-36c5-40a7-8f07-00a77ad85d95 · outbound
NVILA: Efficient Frontier Visual Language Models VILA$^2$: VILA Augmented VILA
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 16dc2d83-7125-423b-8fce-4acc5191ded3 · outbound
NVILA: Efficient Frontier Visual Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 472ca8b3-5aa7-484a-9f42-e3db9cad979b · outbound
NVILA: Efficient Frontier Visual Language Models NVLM: Open Frontier-Class Multimodal LLMs
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aae3c11c-6695-43ab-8603-a73b3decfd54 · outbound
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation be39da09-02f6-4bf7-8458-0dd31891e10a · outbound
NVILA: Efficient Frontier Visual Language Models Token Merging: Your ViT But Faster
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a25880f7-17e2-4985-8864-df67bdb5350b · outbound
NVILA: Efficient Frontier Visual Language Models An Image is Worth 1/2 Tokens After Layer 2: Plug- and-Play Inference Acceleration for Large Vision- Language Models
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1609cef4-edcb-4880-a0b6-619850ec017e · outbound
NVILA: Efficient Frontier Visual Language Models PYRA: Parallel Yielding Re-Activation for Training-Inference Effi- cient Task Adaptation
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ad59abf5-8b18-4838-b771-f17481662b7a · outbound
NVILA: Efficient Frontier Visual Language Models vid-TLDR: Train- ing Free Token Merging for Light-Weight Video Transformer
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e21b41aa-b298-44f2-ad26-40da19a210df · outbound
NVILA: Efficient Frontier Visual Language Models Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1574aa84-21a1-442e-9207-b8afb5df4826 · outbound
NVILA: Efficient Frontier Visual Language Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6760d202-3531-4459-8452-31a5f500b91b · outbound
NVILA: Efficient Frontier Visual Language Models SpatialRGPT: Grounded Spatial Reason- ing in Vision Language Models
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e77fef2b-922b-4ab1-8c30-6d2ed0685a44 · outbound
NVILA: Efficient Frontier Visual Language Models DoReMi: Op- timizing Data Mixtures Speeds Up Language Model Pretraining
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 976e3530-1f39-4120-b463-541ccb795faf · outbound
NVILA: Efficient Frontier Visual Language Models MoDS: Model-oriented Data Selection for Instruction Tuning
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c4b9be6e-0317-4925-8dc5-709c656c726e · outbound
NVILA: Efficient Frontier Visual Language Models Scaling FP8 training to trillion-token LLMs
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 14c8b524-7342-4caa-a142-b1975c554440 · outbound
NVILA: Efficient Frontier Visual Language Models FP8 Formats for Deep Learning
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ae84b236-cebe-4939-931a-fc09e2da170b · outbound
NVILA: Efficient Frontier Visual Language Models Com- pact Language Models via Pruning and Knowledge Distillation
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a9bc8201-24ed-4ba1-9593-5b6c2256b0e8 · outbound
NVILA: Efficient Frontier Visual Language Models Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation be766e94-660f-42f8-8fca-43666945ccc8 · outbound
NVILA: Efficient Frontier Visual Language Models GPTQ: Accurate Post-Training Quan- tization for Generative Pre-Trained Transformers
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 60f5bdd0-3df4-49b6-a5f3-6ec1bc50dc64 · outbound
NVILA: Efficient Frontier Visual Language Models LoRA: Low-Rank Adaptation of Large Language Models
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0c4f7e19-de2c-45c7-846b-7172114b5793 · outbound
NVILA: Efficient Frontier Visual Language Models DoRA: Weight- Decomposed Low-Rank Adaptation
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8a622607-903e-4f87-a51b-ed3b40ddfd4f · outbound
NVILA: Efficient Frontier Visual Language Models QLoRA: Efficient Finetun- ing of Quantized LLMs
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6d67f1e4-b2a7-4c76-8162-f911cb025cdd · outbound
NVILA: Efficient Frontier Visual Language Models GaLore: Memory-Efficient LLM Training by Gradi- ent Low-Rank Projection
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 12465fea-7643-4fdb-a10f-7759182e5100 · outbound
NVILA: Efficient Frontier Visual Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2ca4f020-57c9-4508-95f9-499a99c5fe63 · outbound
NVILA: Efficient Frontier Visual Language Models Building and better understanding vision-language models: insights and future directions
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f58c8bef-486f-48b2-a869-2c475a052ca8 · outbound
NVILA: Efficient Frontier Visual Language Models PDF Associ- ation Dataset (PDFA)
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 68029f20-cedd-4d18-b4ab-e4e25e70a08d · outbound
NVILA: Efficient Frontier Visual Language Models ICDAR 2019 Competition on Large-Scale Street View Text with Partial Labeling
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fba87277-fe75-4583-95ac-b8fe163a77fb · outbound
NVILA: Efficient Frontier Visual Language Models ICDAR 2019 Robust Reading Challenge on Arbitrary-Shaped Text
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5a6750c8-025c-4166-bb09-d605f1438115 · outbound
NVILA: Efficient Frontier Visual Language Models COYO-700M: Image-Text Pair Dataset
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5b4a343d-7fb3-42fe-a459-72d59a9ef5db · inbound
MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models NVILA: Efficient Frontier Visual Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e86d5a3a-83de-45de-8fdc-b8cae8194041 · inbound
OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning NVILA: Efficient Frontier Visual Language Models
Reference 147
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation cbd5ebb6-c646-462d-9ea0-bfbe0c282b7a · inbound
FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models NVILA: Efficient Frontier Visual Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbb8d4ec-71b1-42a0-8a5e-08805a4c0a71 · inbound
ViDDAR: Vision Language Model-Based Task-Detrimental Content Detection for Augmented Reality NVILA: Efficient Frontier Visual Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 643ac8e0-6a44-48cd-b137-d9331bdd8c8b · inbound
VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding NVILA: Efficient Frontier Visual Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d33a54a7-9b41-4224-8a20-4010fe75f9eb · inbound
Temporal Preference Optimization for Long-Form Video Understanding NVILA: Efficient Frontier Visual Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 241e713a-6af2-45eb-8fb4-b577a8a5af47 · inbound
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models NVILA: Efficient Frontier Visual Language Models
Reference 199
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eedac329-4758-4a4c-8b3b-004a336318fa · inbound
SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer NVILA: Efficient Frontier Visual Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a066575-8af0-4f62-a543-226beb93ca90 · inbound
Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models NVILA: Efficient Frontier Visual Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6297378-2b6d-428a-b06c-488d05e77eee · inbound
LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models NVILA: Efficient Frontier Visual Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3c83d3d-f68a-4096-a917-9d2541186504 · inbound
WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs NVILA: Efficient Frontier Visual Language Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b6368733-9fef-477b-b20f-a6c9b9c6ef7a · inbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation NVILA: Efficient Frontier Visual Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8dab4d0-57e2-4045-bbef-6ea7cf7dc34d · inbound
Training-Free Reasoning and Reflection in MLLMs NVILA: Efficient Frontier Visual Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65a724fb-6abd-4279-8576-c22f418abf0d · inbound
Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning NVILA: Efficient Frontier Visual Language Models
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b46feefe-1245-44fc-8ce9-c9a71c63c07d · inbound
LaViDa: A Large Diffusion Language Model for Multimodal Understanding NVILA: Efficient Frontier Visual Language Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef6700ac-f5f4-4846-b60a-0cefd3575d1f · inbound
FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities NVILA: Efficient Frontier Visual Language Models
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a9d3f84-1937-4239-94cd-d89cce996e5c · inbound
ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models NVILA: Efficient Frontier Visual Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8f94370-a633-46b5-aad4-a98a9265096e · inbound
VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation NVILA: Efficient Frontier Visual Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df08e76c-feda-4625-aff8-133bf071001d · inbound
EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models NVILA: Efficient Frontier Visual Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53815c49-612f-4fdc-ace2-72e165abdbdc · inbound
FlexSelect: Flexible Token Selection for Efficient Long Video Understanding NVILA: Efficient Frontier Visual Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc85cb85-975e-4319-8afc-377612f6b484 · inbound
MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping NVILA: Efficient Frontier Visual Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54f3858b-f0ae-4f61-9f1b-9fe9c8f76c74 · inbound
AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation NVILA: Efficient Frontier Visual Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56673960-7564-4da5-a3a2-2610c81da451 · inbound
MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos NVILA: Efficient Frontier Visual Language Models
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e1d889c-7dd5-42e2-af56-33384038d571 · inbound
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? NVILA: Efficient Frontier Visual Language Models
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3b33f33-822a-4f75-b31c-5aaa08be4c8f · inbound
EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos NVILA: Efficient Frontier Visual Language Models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d9fd4ac1-42a1-4395-ae53-7f3282d22980 · inbound
ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning NVILA: Efficient Frontier Visual Language Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 94e994f1-c11d-493c-8020-9568fcb97295 · inbound
Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring NVILA: Efficient Frontier Visual Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcae26e5-9d74-4f1b-b725-9d1175c24896 · inbound
MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning NVILA: Efficient Frontier Visual Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae0a074e-4087-4795-8e9b-af2c3d3bd5ea · inbound
HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes NVILA: Efficient Frontier Visual Language Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c91c6ae7-578e-425a-949f-7cf55e871954 · inbound
Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation NVILA: Efficient Frontier Visual Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 19fd7d4d-a3e7-43ca-ab85-6c1d59da13d0 · inbound
DODO: Discrete OCR Diffusion Models NVILA: Efficient Frontier Visual Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 146ecfb9-4eba-431d-b9dc-8014c112f4e7 · inbound
XNote: Benchmarking Automated Community Notes Generation for Image-based Contextual Deception NVILA: Efficient Frontier Visual Language Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2b981b61-9835-4d9f-ad5f-fac46f34d757 · inbound
A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning NVILA: Efficient Frontier Visual Language Models
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5d3ba47f-3898-46ba-a810-ee075bf51136 · inbound
OProver: A Unified Framework for Agentic Formal Theorem Proving NVILA: Efficient Frontier Visual Language Models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 827d637a-0d04-46af-bae8-fcc3215e2e08 · inbound
StableVLA: Towards Robust Vision-Language-Action Models without Extra Data NVILA: Efficient Frontier Visual Language Models
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e9b193be-822e-4e95-b0ae-5d5df97a8910 · inbound
Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding NVILA: Efficient Frontier Visual Language Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e9047406-2dbb-4c90-889d-46404d94c104 · inbound
ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs NVILA: Efficient Frontier Visual Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f746aa27-d10a-4bf0-87c5-abcb2ad6eb0f · inbound
Worth Remembering: Surprise-Gated Robot Episodic Memory NVILA: Efficient Frontier Visual Language Models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1cc176a4-3e76-4d97-9237-96526862d2ee · inbound
video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding NVILA: Efficient Frontier Visual Language Models
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 89826ebe-5819-4ffd-a570-9dbf1521d10b · inbound
RADIO1D: Elastic Representations for Condensed Vision Modeling NVILA: Efficient Frontier Visual Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d17f9179-1967-46e3-a4f5-ef6e5753af6e · inbound
QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding NVILA: Efficient Frontier Visual Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea0d2259-c377-4a65-9e2d-3c4fe634035e · inbound
EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization NVILA: Efficient Frontier Visual Language Models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db80bc62-7e7f-4b6a-b90a-06ce968b1f47 · inbound
Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles NVILA: Efficient Frontier Visual Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.