Pith. sign in

Paper Citation Record · LEDGER

NVILA: Efficient Frontier Visual Language Models

As of 14 August 2026, this Paper Citation Record lists 100 of 147 outbound references and 43 inbound Pith citation observations for arXiv:2412.04468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04468 v3

Coverage vector

measured 100 of 147 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-23T07:42:22.478647Z

measured 143 of 143 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 43 of 43 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T04:28:26.603246Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T16:39:58.306201Z

Reference resolution

100 of 147 outbound references displayed

  • verified exact36
  • verified fuzzy60
  • unresolved1
  • parse uncertain2
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c658f817-5680-4c89-9a34-5300500b094f · outbound

This paper cites Visual Instruction Tuning.

NVILA: Efficient Frontier Visual Language Models Visual Instruction Tuning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.961908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:afdb5c551e52821f525b245966f1ab534ffe5f71ed4f6b5ac103b4b274f17b56

Observation 7e402fbf-49ce-4bb6-bc8d-d1fa3b554776 · outbound

This paper cites VILA: On Pre- training for Visual Language Models.

NVILA: Efficient Frontier Visual Language Models VILA: On Pre- training for Visual Language Models

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.138115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:6e5cf039cb0ddf7769c2b954286fd9848fceeda391ae7b6665311adadd1a6a4d

Observation 21618f3b-b038-47eb-bf36-c96871706eaf · outbound

This paper cites InternVL: Scal- ing up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

NVILA: Efficient Frontier Visual Language Models InternVL: Scal- ing up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.149747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:d16aa8f757c07ea6a6742f70f8b68d3e6c1e7726e55f338d4eec02fdf646e193

Observation de0a1d31-fe8c-4354-aea6-046e16d7fbde · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

NVILA: Efficient Frontier Visual Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.116592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:77cbd5ac7236422ee6ef695ef70f4858719779a997db48b278433921f0634289

Observation eb365286-d69a-4951-ae58-926ac93b665f · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

NVILA: Efficient Frontier Visual Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.034094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:1c36a6c69a6ba7998e4a6977387ad57499fad36be51ccee74127d4f3320729db

Observation d9835b24-3200-4f31-b79c-68f7dd2c26a0 · outbound

This paper cites RT-1: Robotics Transformer for Real-World Control at Scale.

NVILA: Efficient Frontier Visual Language Models RT-1: Robotics Transformer for Real-World Control at Scale

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.179860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:6f4c9c46f475abea9407ddb5b8d8fd05f3ab725324d5c659fa3cdf7d309242c2

Observation 07340e2c-05a6-4030-b44a-9a3512e88a75 · outbound

This paper cites NaVid: Video- based VLM Plans the Next Step for Vision-and- Language Navigation.

NVILA: Efficient Frontier Visual Language Models NaVid: Video- based VLM Plans the Next Step for Vision-and- Language Navigation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.183731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:ca8c37532c0383bb82fef61eceffb07fda1accf1c548e9a1990daed151b66185

Observation fac3d417-f91f-4a1e-a41c-8d39f05ae5cc · outbound

This paper cites NaVILA: Legged Robot Vision-Language-Action Model for Navigation.

NVILA: Efficient Frontier Visual Language Models NaVILA: Legged Robot Vision-Language-Action Model for Navigation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.167538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:a7d1805f613f586774c43e943a2f9d160c2e51d1dad6ff4deaffcd5ea2449cd2

Observation 2c051837-21f5-4308-9aa1-12d9722b9c3e · outbound

This paper cites DriveVLM: The Con- vergence of Autonomous Driving and Large Vision- Language Models.

NVILA: Efficient Frontier Visual Language Models DriveVLM: The Con- vergence of Autonomous Driving and Large Vision- Language Models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.164564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:68621a74c9dccea0a59a64782709539fa6a4dc377bac12f9dd39a86dadc3c675

Observation 065bccfa-ae10-4737-8662-e41571af98e6 · outbound

This paper cites Capabilities of Gemini Models in Medicine.

NVILA: Efficient Frontier Visual Language Models Capabilities of Gemini Models in Medicine

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.173481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:796c7c2bebea59344a820ec1ff47d45d0c2c01f3973bef3fc8556fa20630ff3c

Observation de87e671-2e4d-4750-a93a-b575a91f387c · outbound

This paper cites VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge.

NVILA: Efficient Frontier Visual Language Models VILA-M3: Enhancing Vision-Language Models with Medical Expert Knowledge

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.097491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:1e4d9b3fddc272b1c15f443f246865047e1668204a1a06f717e036fd6aa14bcb

Observation f0305393-6f15-4e35-9c92-1e8b21d1ea90 · outbound

This paper cites GPT-4o.

NVILA: Efficient Frontier Visual Language Models GPT-4o

Reference 12

Resolution
parse uncertain
raw_fallback, observed 2026-05-23T07:42:44.998704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:2879b9b5a3a1480229cddeaf233632fbe6c4d88c6f4c685eaf92d6cc208d94f1

Observation f436c905-dcd4-44ca-a083-c9d1fc291fdf · outbound

This paper cites Claude 3.5.

NVILA: Efficient Frontier Visual Language Models Claude 3.5

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.002617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:8778873bde960011bfea9d1f95084a3ffe61e089eca5d82cd296e043ddbdbec3

Observation 0e945f66-600a-4f9c-ab60-a6e9bf05e738 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

NVILA: Efficient Frontier Visual Language Models Sigmoid Loss for Language Image Pre-Training

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.006550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:44d83d98a4fed6379a22860331dc746e825c5a6017c0dd7b03bed307d180ef1a

Observation e93f96ae-481e-4e62-a3b8-71c89280d9a2 · outbound

This paper cites Qwen2 Technical Report.

NVILA: Efficient Frontier Visual Language Models Qwen2 Technical Report

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:42.974053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:96b7c2f942aab4b0b7f4daa54f2603485c9251f0811b08e6cdae818f286016a4

Observation 72e248a8-940f-4bcb-a4a7-fa82bbfa9974 · outbound

This paper cites When Do We Not Need Larger Vision Models? InEuropean Conference on Com- puter Vision (ECCV).

NVILA: Efficient Frontier Visual Language Models When Do We Not Need Larger Vision Models? InEuropean Conference on Com- puter Vision (ECCV)

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.021123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:5199df36b51acc4643b4f4e49222b4a3e097b91bdba43e57cc293efa35d8a61a

Observation dfc1914d-df12-4c14-8922-a08385485d80 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

NVILA: Efficient Frontier Visual Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.204294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:f3642dd793a994544fad7f99be13500add42aafde812c7d9bed7eeac5014b16c

Observation c392ca64-216d-42c5-92cc-b81534e786c0 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

NVILA: Efficient Frontier Visual Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.085376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:1d1cd6735e98dd07123f22f24064c0e6ed65f5d6396c12e539267c883ca8c1a4

Observation ea14d6be-01bf-48f9-b97d-92803853f4b6 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

NVILA: Efficient Frontier Visual Language Models LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.247372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:625b877d91d66a84fe447305c47f555a77fedd4cb325383a31f082731387ea96

Observation a7ac950e-dd95-4f91-bdc0-6b0feafac78c · outbound

This paper cites Tem- poral Segment Networks: Towards Good Practices for Deep Action Recognition.

NVILA: Efficient Frontier Visual Language Models Tem- poral Segment Networks: Towards Good Practices for Deep Action Recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.949356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:401554299ff799bf155d9507883e1eecfd9fda24c3e0eef2275b90e98f9d8abc

Observation 838def6e-7b0d-40da-a2fb-3d8bc4eaccd3 · outbound

This paper cites Cambrian-1: A Fully Open, Vision- Centric Exploration of Multimodal LLMs.

NVILA: Efficient Frontier Visual Language Models Cambrian-1: A Fully Open, Vision- Centric Exploration of Multimodal LLMs

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.945210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:7adacfaa2fd7d752fceea6a31213d9ce9de8e6cc3e80af01408bc6e689f8ba0e

Observation cf2cddb9-dd6a-4ef2-b297-79d621f7bdb9 · outbound

This paper cites What Matters When Building Vision-Language Models? In Conference on Neural Information Processing Systems (NeurIPS).

NVILA: Efficient Frontier Visual Language Models What Matters When Building Vision-Language Models? In Conference on Neural Information Processing Systems (NeurIPS)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.953902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:50d0b3d2bd912cdc9ab5fa9f4bae82355301f4f6e5e31d8ef8ceac64c2c98001

Observation f5600f71-53cf-473b-b204-887322d50fef · outbound

This paper cites Selec- tion via Proxy: Efficient Data Selection for Deep Learning.

NVILA: Efficient Frontier Visual Language Models Selec- tion via Proxy: Efficient Data Selection for Deep Learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.982503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:b0a5ed39db375fb235f8f2224b6d33da3865a872630601ea08fea62f85479043

Observation 1163e972-deef-40b9-9123-c9e0e6067e82 · outbound

This paper cites Demystifying CLIP Data.

NVILA: Efficient Frontier Visual Language Models Demystifying CLIP Data

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.937555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:9111beda670bb28b720e5bdb2bd5acd4b2c5314d06208368fe5b7587651c3bc8

Observation 0e70f842-e46f-40ce-ba63-ef807efa0db2 · outbound

This paper cites SemDeDup: Data-efficient learning at web-scale through semantic deduplication.

NVILA: Efficient Frontier Visual Language Models SemDeDup: Data-efficient learning at web-scale through semantic deduplication

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.028054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:72102be737845f1c40b84c60e6a06b0001033f6b2a3d0e459f6b41d12fba5e91

Observation 844471cb-a865-48c8-8d4b-d4904b48406a · outbound

This paper cites D4: Improving LLM Pretraining via Document De-Duplication and Diversification.

NVILA: Efficient Frontier Visual Language Models D4: Improving LLM Pretraining via Document De-Duplication and Diversification

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.933854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:a16e72d330722739fcaaee5330b03164bba03f9f11c6b7f15826dfa243d3c7c8

Observation 5d6076ce-9d5b-442e-af8a-cd03b61d2312 · outbound

This paper cites LESS: Select- ing Influential Data for Targeted Instruction Tuning.

NVILA: Efficient Frontier Visual Language Models LESS: Select- ing Influential Data for Targeted Instruction Tuning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.929763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:4988c5d860591797ca6418c5ab44acf581e018c1d801802d2fb355546f88e48d

Observation 95b2051f-5b18-4279-8f96-93f0b70a0724 · outbound

This paper cites Data Selection via Optimal Control for Language Models.

NVILA: Efficient Frontier Visual Language Models Data Selection via Optimal Control for Language Models

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.067445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:893e7281376a54eed20408cf578f8405cedf3758c493b49da4becc27baa9bf70

Observation 3ec15776-1f00-489e-bcf4-bdb56a9bca33 · outbound

This paper cites MiniPLM: Knowledge Distillation for Pre-Training Language Models.

NVILA: Efficient Frontier Visual Language Models MiniPLM: Knowledge Distillation for Pre-Training Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.229445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:ed2d8fd8a16f39d98df508bda31742a676a5069a701dd3a8022f54455692ce14

Observation b688c074-fbc0-45fe-a1fa-7b85d8d9a822 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

NVILA: Efficient Frontier Visual Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.122814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:ac02a5fee614a44fe24babac27b53623c8dee2632cb7a1eabf17b17ab1db4b13

Observation 81739734-a2c9-4894-82f2-ef951d1da7b5 · outbound

This paper cites Mixed Precision Training.

NVILA: Efficient Frontier Visual Language Models Mixed Precision Training

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.990057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:cb4246cb3413e71bae2a0736cfaab2ad479bc953384a84425077f9f4ad1471c1

Observation 50250526-4ebc-48ec-b0e4-66fe55ab60a8 · outbound

This paper cites A Study of BFLOAT16 for Deep Learning Training.

NVILA: Efficient Frontier Visual Language Models A Study of BFLOAT16 for Deep Learning Training

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.241273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:4c9b8c266f6b887cb48bd2c8bbf498c9cd0047ef61b58ccff8a067ef2626124e

Observation 763e5068-7994-42db-ac48-6d78988c0002 · outbound

This paper cites FP8-LM: Training FP8 Large Language Models.

NVILA: Efficient Frontier Visual Language Models FP8-LM: Training FP8 Large Language Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.015165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:83865d66869653678857d00e7769b8ab16522ffd6c64262660ce49074134ea49

Observation 8115b9a9-7792-413d-b96d-27780fc30645 · outbound

This paper cites COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training.

NVILA: Efficient Frontier Visual Language Models COAT: Compressing Optimizer states and Activation for Memory-Efficient FP8 Training

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.059995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:f861d567ec0b81b7c6dc87a65df26d692f9fe96a16268324870db92bc5fa5e56

Observation 3a9ec0b3-8fb2-4a21-9f76-98ed7625c5df · outbound

This paper cites Liger Kernel: Efficient Triton Kernels for LLM Training.

NVILA: Efficient Frontier Visual Language Models Liger Kernel: Efficient Triton Kernels for LLM Training

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.002506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:5582b56fa43c398f361970a847f9ef198af7b622e0d53be6f476b5ec6549af3b

Observation 6399bd8d-8383-441f-8e6a-2203a2c28312 · outbound

This paper cites Android in the Zoo: Chain-of-Action-Thought for GUI Agents.

NVILA: Efficient Frontier Visual Language Models Android in the Zoo: Chain-of-Action-Thought for GUI Agents

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.906025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:6695a9b5d968eb9293b6be6cfbdbe3ce99140fa90f83a035fa1e1063f475008d

Observation 2829ea84-aa87-4322-a0bc-44a063cfa7c9 · outbound

This paper cites ALFRED: A 14 NVILA: Efficient Frontier Visual Language Models Benchmark for Interpreting Grounded Instructions for Everyday Tasks.

NVILA: Efficient Frontier Visual Language Models ALFRED: A 14 NVILA: Efficient Frontier Visual Language Models Benchmark for Interpreting Grounded Instructions for Everyday Tasks

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.911213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:6c81ffc7946f8eb72f86da6e557bf6f9ce0c94929e7d7ddf2d21497bc711921c

Observation 94f0944f-bc99-40b0-aa59-d3f6a4581304 · outbound

This paper cites nuScenes: A multimodal dataset for au- tonomous driving.

NVILA: Efficient Frontier Visual Language Models nuScenes: A multimodal dataset for au- tonomous driving

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.915070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:316495ee2c018f24e40c05904631f848b3ab26cd70347878b7b1611f941b646f

Observation c3b1dec3-33db-45ab-9e43-2308ea293ca0 · outbound

This paper cites PathVQA: 30000+ Questions for Medical Visual Question Answering.

NVILA: Efficient Frontier Visual Language Models PathVQA: 30000+ Questions for Medical Visual Question Answering

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.185714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:5d79f13f1b3a24be7a5ccf5999404ab94144808de859ead70abc74b25662087a

Observation 9599a630-956d-4f7c-96e2-5c193a50c1c5 · outbound

This paper cites Widget Captioning: Generat- ing Natural Language Description for Mobile User Interface Elements.

NVILA: Efficient Frontier Visual Language Models Widget Captioning: Generat- ing Natural Language Description for Mobile User Interface Elements

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.974323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:faa575751316c6f190a6652aa7cdf68c2177d667a6fb901a4bead50a16aa8fa4

Observation df564449-1736-41d3-b567-f5585c43199e · outbound

This paper cites AWQ: Activation-Aware Weight Quantization for On-Device LLM Compression and Acceleration.

NVILA: Efficient Frontier Visual Language Models AWQ: Activation-Aware Weight Quantization for On-Device LLM Compression and Acceleration

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.965633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:4b3a893a3daa1e13512d1ee20b65ed06b0561c6a77687e99e4ffdf15f18c2e72

Observation 4e818683-0073-4a39-bf6f-04a812287ba4 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learn- ing Library.

NVILA: Efficient Frontier Visual Language Models PyTorch: An Imperative Style, High-Performance Deep Learn- ing Library

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.978357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:a9e06fa52d665eaed2044981d5f5496e4113375f90c8333c0f7b33cef9efbed2

Observation 362f27c4-387b-4acc-8d88-8844a1090d22 · outbound

This paper cites an unresolved cited work.

NVILA: Efficient Frontier Visual Language Models Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-05-23T07:42:44.878933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:715d54a9f07e01ee515c7a0855b923500285a7ead1ee989b91535a0f229c4fd5

Observation 7299f300-cadb-42a6-b99d-47f60602afe6 · outbound

This paper cites Transform- ers: State-of-the-Art Natural Language Processing.

NVILA: Efficient Frontier Visual Language Models Transform- ers: State-of-the-Art Natural Language Processing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.969823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:c85df7ff73b2ed7168497bef0e009c1d3631a91c264f386df1c87760a2825389

Observation bdf9982e-289a-404b-86f0-a773dde55f94 · outbound

This paper cites DeepSpeed: System Optimiza- tions Enable Training Deep Learning Models with Over 100 Billion Parameters.

NVILA: Efficient Frontier Visual Language Models DeepSpeed: System Optimiza- tions Enable Training Deep Learning Models with Over 100 Billion Parameters

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.882790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:5ed550a96445ee11e3815ad8e66bd6b195d888f637b7f3a7671b591abd957c6c

Observation ebea230e-88ce-43fa-b5ef-82970b2c7fd7 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

NVILA: Efficient Frontier Visual Language Models FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.865826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:33600df6d21016bdfdd3f7a123e3e1c65449bc476375d4233501982a969f618b

Observation 1699b915-1da3-4822-af0a-c2cb67256e73 · outbound

This paper cites A Diagram is Worth a Dozen Images.

NVILA: Efficient Frontier Visual Language Models A Diagram is Worth a Dozen Images

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.862305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:f6c207e085ef0b44bd8df7eee0e9278040027084516567c9dee9c84d726e4eb1

Observation 4e981fcc-98c0-4420-8116-b1b00f50cd7a · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

NVILA: Efficient Frontier Visual Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.871799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:ba33772e43ba819332c648403deaea1f6f7d377f86b77a1c7ecbbb2a634da39a

Observation 665a079b-3d1f-4a86-bddc-bad24504bb99 · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

NVILA: Efficient Frontier Visual Language Models DocVQA: A Dataset for VQA on Document Images

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.875552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:096f6e5b38ef48c11e0403a5b831d48d7dc7b137b01e4445826fe0f3b1f8785c

Observation 1025ed41-a7cd-4606-a99f-fcf192a0416c · outbound

This paper cites InfographicVQA.

NVILA: Efficient Frontier Visual Language Models InfographicVQA

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.888114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:c2222f6eeb3a85c535ea664ab3dafc2574921d1ecde1a4aa492776d5babb1163

Observation 23ecab5f-374f-4d4e-9d6f-7f423bb37078 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

NVILA: Efficient Frontier Visual Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.891994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:74e0f33d43b7c874fd5b18bea84ad7dfedccafff9c9930e85d156218b9cffeb1

Observation e89a03d8-0081-4be5-a23f-568cb384996c · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Ex- pert AGI.

NVILA: Efficient Frontier Visual Language Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Ex- pert AGI

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.986285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:1724c80043df6e1369fcabe388987fe8565079f6074d9595bd4a205b6ea65f61

Observation 888533a0-943d-4b67-a934-79718fc8fb21 · outbound

This paper cites Grok-1.5.

NVILA: Efficient Frontier Visual Language Models Grok-1.5

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.841228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:2464a42fffebce13b55b9c97423783783412dd8dd1480bbdc946d946f09ac23e

Observation c3d8f52c-a31e-422b-92a7-40587bf421ca · outbound

This paper cites SEED-Bench: Bench- marking Multimodal LLMs with Generative Com- prehension.

NVILA: Efficient Frontier Visual Language Models SEED-Bench: Bench- marking Multimodal LLMs with Generative Com- prehension

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.800633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:768510e96192d47db8e1beb8886294e784596a5d842b67cb19c0926250e05c99

Observation c3c9bccb-336b-44eb-b7ab-d0bfee4ba589 · outbound

This paper cites Towards VQA Models That Can Read.

NVILA: Efficient Frontier Visual Language Models Towards VQA Models That Can Read

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.792688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:816be587ce949ea959220cd862fb7f0cb5780bde814d01d452f128e160535cad

Observation df3b8aaf-47a1-4696-bf74-5d56706a41d0 · outbound

This paper cites Making the V in VQA Matter: Elevating the Role of Image Un- derstanding in Visual Question Answering.

NVILA: Efficient Frontier Visual Language Models Making the V in VQA Matter: Elevating the Role of Image Un- derstanding in Visual Question Answering

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.786539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:8ac0211b041fd5004c08babb48277c9d1b100f7294198e133bd45ede279dbc95

Observation a5d6d94b-0c36-40cd-b6ad-05d9f632e7d8 · outbound

This paper cites ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering.

NVILA: Efficient Frontier Visual Language Models ActivityNet-QA: A Dataset for Understanding Complex Web Videos via Question Answering

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.796813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:32c659319f66bd9630788d6a9f23fd5447c230c65bd8493aab78bcf8f9a36012

Observation 6d4e6973-97bc-4567-a4d3-cfd190ff6f8f · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

NVILA: Efficient Frontier Visual Language Models MLVU: Benchmarking Multi-task Long Video Understanding

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.079511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:61f29a89a7e6fba152d599e2367cbf493b24f017d5ed9c685702e156bb360310

Observation 055794bc-c8ac-499c-8a8b-b2327ebca634 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

NVILA: Efficient Frontier Visual Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:44.845079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:c965506166318dbd625e37f3ce7aad4f318677379085c8739c5ad4a95bc00294

Observation 6c74f650-7f1a-489a-9a26-33bb88593b47 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

NVILA: Efficient Frontier Visual Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.253957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:ed11b23639e85fc522146e1380c6c2f6f6667208cf8c5af896039502b48c672d

Observation 1cf91e4a-058b-46e0-825a-d4ae68ca3742 · outbound

This paper cites Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding.

NVILA: Efficient Frontier Visual Language Models Video-XL: Extra-Long Vision Language Model for Hour-Scale Video Understanding

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:42.987016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:71aa89472d865e2cedb7d205b1d4ac4d9a28074d8d44b11bc643422bab15fa7a

Observation e033c718-bf9c-44dd-9d21-9a460348fedc · outbound

This paper cites LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding.

NVILA: Efficient Frontier Visual Language Models LongVU: Spatiotemporal Adaptive Compression for Long Video-Language Understanding

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.110273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:08f08ef22c3ea43528f7f16dfc37d3cb475c0480b1326766e65e7742d8fb6b7b

Observation e0925f8d-efec-4110-8251-c6b1daa0f532 · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

NVILA: Efficient Frontier Visual Language Models Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.235839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:55672742241254a2479d18d95ea4215947e07ccb2feaabc9ab80bc7ade7dbd4d

Observation edab20d8-b0ce-4df3-a0ff-1c2ad6558c6e · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

NVILA: Efficient Frontier Visual Language Models Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.191341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:d784457ca39d2e481d63900252b23668f6b6cba374d62434ba4882c62f4e315d

Observation c5ee4a49-5d12-45af-8124-3dee8cec76f9 · outbound

This paper cites Beyond the Nav- Graph: Vision-and-Language Navigation in Con- tinuous Environments.

NVILA: Efficient Frontier Visual Language Models Beyond the Nav- Graph: Vision-and-Language Navigation in Con- tinuous Environments

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.171916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:a0ac7cea6b1e30b598088c9631c8481dce904e45b9e3d701c1d30ba4bc10b682

Observation 89292ad6-ecc1-4c16-a690-03ea6392567f · outbound

This paper cites Vision- and-Language Navigation: Interpreting visually- grounded navigation instructions in real environ- ments.

NVILA: Efficient Frontier Visual Language Models Vision- and-Language Navigation: Interpreting visually- grounded navigation instructions in real environ- ments

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.175899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:63902814e310072b5ea0991b4241d1cc300e648675ca84a2d876c919716e86de

Observation 77b1c0f1-b55b-4e80-977e-0c679e5b3e33 · outbound

This paper cites GPT-4V.

NVILA: Efficient Frontier Visual Language Models GPT-4V

Reference 67

Resolution
parse uncertain
raw_fallback, observed 2026-05-23T07:42:45.160700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:698b75872d5ca2ca639c96d2c7a9f7cf08383be64b2180ca1d099af2e4d6587e

Observation da30d0df-3a48-4caa-b0c8-7788deffc3d8 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

NVILA: Efficient Frontier Visual Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.216467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:97c57859c1c26312e1e23649e75b14095f1da00f0bcf7b681654b0a67f641d82

Observation edd9bf7d-9161-4a80-b075-6286672d5cb8 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

NVILA: Efficient Frontier Visual Language Models Gemini: A Family of Highly Capable Multimodal Models

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.198281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:42e7052a5989c2aa07c98bb02bc6175a79da8376d49ef4774ca631a262c42d08

Observation ddb48495-72ea-4dd7-a51c-653c855f8292 · outbound

This paper cites Claude 3.

NVILA: Efficient Frontier Visual Language Models Claude 3

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.157283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:332b11283a8d4f63d06288cff1b5b6001a12e771ba5f9998b6dd397cfc12b387

Observation 7685e922-b9cd-486f-8d3f-53beeaf886d3 · outbound

This paper cites MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning.

NVILA: Efficient Frontier Visual Language Models MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.073686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:d77ac9d38604876e8282ea47ce8cbd21e6e06d374d700cd676722a504a895e09

Observation 1c284017-9be1-4db1-aa96-a17c583a5a8c · outbound

This paper cites NeMo: a toolkit for building AI applications using Neural Modules.

NVILA: Efficient Frontier Visual Language Models NeMo: a toolkit for building AI applications using Neural Modules

Reference 72

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T07:42:43.155511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:8a0c44e36f545a67b121b71f95b594745326c82c61b2d73bc6ea73ddd79193f3

Observation 7dfa2d0c-36c5-40a7-8f07-00a77ad85d95 · outbound

This paper cites VILA$^2$: VILA Augmented VILA.

NVILA: Efficient Frontier Visual Language Models VILA$^2$: VILA Augmented VILA

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.222941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:264492f7517f19cd67544af91d4dbeb211a95c698cfd0922f881b89605c5a876

Observation 16dc2d83-7125-423b-8fce-4acc5191ded3 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

NVILA: Efficient Frontier Visual Language Models Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.210588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:9ae7e5627e17aa9147cc8442f932fbe8b52669127addde89d3ab7fbca68c014c

Observation 472ca8b3-5aa7-484a-9f42-e3db9cad979b · outbound

This paper cites NVLM: Open Frontier-Class Multimodal LLMs.

NVILA: Efficient Frontier Visual Language Models NVLM: Open Frontier-Class Multimodal LLMs

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:42.980731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:359a93af50f3e4861ebc08e2a80021d20f40510c70d552afdd6ce42b55fbc252

Observation aae3c11c-6695-43ab-8603-a73b3decfd54 · outbound

This paper cites Llama 3.

NVILA: Efficient Frontier Visual Language Models Llama 3

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.130109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:90aa1b38d4e9d4f7aaa1a78aabbc160db535b687bba715f51b3fda477ec7d55f

Observation be39da09-02f6-4bf7-8458-0dd31891e10a · outbound

This paper cites Token Merging: Your ViT But Faster.

NVILA: Efficient Frontier Visual Language Models Token Merging: Your ViT But Faster

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.122333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:0b5fda42dfa0d1b1105e47e0a3919fea7ce056da2a7057afc01cfc3d674cbd00

Observation a25880f7-17e2-4985-8864-df67bdb5350b · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug- and-Play Inference Acceleration for Large Vision- Language Models.

NVILA: Efficient Frontier Visual Language Models An Image is Worth 1/2 Tokens After Layer 2: Plug- and-Play Inference Acceleration for Large Vision- Language Models

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.126198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:75c737d9adfbb2d3a3666e12dfdf52622040e2cd6b3b56d72a33d1688bcf1ffd

Observation 1609cef4-edcb-4880-a0b6-619850ec017e · outbound

This paper cites PYRA: Parallel Yielding Re-Activation for Training-Inference Effi- cient Task Adaptation.

NVILA: Efficient Frontier Visual Language Models PYRA: Parallel Yielding Re-Activation for Training-Inference Effi- cient Task Adaptation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.134265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:c60e7be52ec0490945ce10079a0102166bbdb99eca163111c0c790bc9886d9d2

Observation ad59abf5-8b18-4838-b771-f17481662b7a · outbound

This paper cites vid-TLDR: Train- ing Free Token Merging for Light-Weight Video Transformer.

NVILA: Efficient Frontier Visual Language Models vid-TLDR: Train- ing Free Token Merging for Light-Weight Video Transformer

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.142009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:29f96f8a6e3dd5cef60d52d9a53c7aa5c0bebce99a04d35fb1dd54deb10ee42b

Observation e21b41aa-b298-44f2-ad26-40da19a210df · outbound

This paper cites Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding.

NVILA: Efficient Frontier Visual Language Models Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.145747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:cbc26a02b6e70cd369a5e378b50b088aab672ba82e43a1ac0e917ca7be0372a3

Observation 1574aa84-21a1-442e-9207-b8afb5df4826 · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

NVILA: Efficient Frontier Visual Language Models SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.161469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:ed6752fff4ac70b6d88ae9442c785ca4c50fa112361ac14ac2da08dad1238c07

Observation 6760d202-3531-4459-8452-31a5f500b91b · outbound

This paper cites SpatialRGPT: Grounded Spatial Reason- ing in Vision Language Models.

NVILA: Efficient Frontier Visual Language Models SpatialRGPT: Grounded Spatial Reason- ing in Vision Language Models

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.113656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:fd93bab4bc0f95c274db355290af6e04f8be4f0761ff12eb10ec7597d397a850

Observation e77fef2b-922b-4ab1-8c30-6d2ed0685a44 · outbound

This paper cites DoReMi: Op- timizing Data Mixtures Speeds Up Language Model Pretraining.

NVILA: Efficient Frontier Visual Language Models DoReMi: Op- timizing Data Mixtures Speeds Up Language Model Pretraining

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.109875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:5d541ce385495fd7fc4c3a4429b6e61d505afe1520df2592b72d6ceec53d7dbc

Observation 976e3530-1f39-4120-b463-541ccb795faf · outbound

This paper cites MoDS: Model-oriented Data Selection for Instruction Tuning.

NVILA: Efficient Frontier Visual Language Models MoDS: Model-oriented Data Selection for Instruction Tuning

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.091735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:ce3dfbb618da938906ab2bc6777a8f7abb79fca2a85dda8a8231b4d3f9937159

Observation c4b9be6e-0317-4925-8dc5-709c656c726e · outbound

This paper cites Scaling FP8 training to trillion-token LLMs.

NVILA: Efficient Frontier Visual Language Models Scaling FP8 training to trillion-token LLMs

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.040130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:cd3d4c3eee187d33531c2992fefa7f32b10176a5f90f5d91b6ba43c3d1363985

Observation 14c8b524-7342-4caa-a142-b1975c554440 · outbound

This paper cites FP8 Formats for Deep Learning.

NVILA: Efficient Frontier Visual Language Models FP8 Formats for Deep Learning

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-23T07:42:43.142391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:304fd202d233526c5c99a848a25fa1bf82d785a13c6325b63edac04886077508

Observation ae84b236-cebe-4939-931a-fc09e2da170b · outbound

This paper cites Com- pact Language Models via Pruning and Knowledge Distillation.

NVILA: Efficient Frontier Visual Language Models Com- pact Language Models via Pruning and Knowledge Distillation

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.101342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:10f5f44239008c444d10c82243a19ddb4e104e10ae23b1593c96e6ad4bf3848d

Observation a9bc8201-24ed-4ba1-9593-5b6c2256b0e8 · outbound

This paper cites Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes.

NVILA: Efficient Frontier Visual Language Models Everybody Prune Now: Structured Pruning of LLMs with only Forward Passes

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.149362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:1336042951920b368712fab5953b0d0ed2ed33652848c42b5dd5803784a3cdb4

Observation be766e94-660f-42f8-8fca-43666945ccc8 · outbound

This paper cites GPTQ: Accurate Post-Training Quan- tization for Generative Pre-Trained Transformers.

NVILA: Efficient Frontier Visual Language Models GPTQ: Accurate Post-Training Quan- tization for Generative Pre-Trained Transformers

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.106108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:eb97a93bf58117f161c43d5e7c7dec746cffd1200a7153c6143a33d941d2bcc5

Observation 60f5bdd0-3df4-49b6-a5f3-6ec1bc50dc64 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

NVILA: Efficient Frontier Visual Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.118346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:03946b2bb6da9243491d3e2847185fa90d22ca06cff0e00a0add95e2c82ccc43

Observation 0c4f7e19-de2c-45c7-846b-7172114b5793 · outbound

This paper cites DoRA: Weight- Decomposed Low-Rank Adaptation.

NVILA: Efficient Frontier Visual Language Models DoRA: Weight- Decomposed Low-Rank Adaptation

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.153608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:63b3be34b12c1fec339f383f47968220c2a363be7d342944d43e7da91712b46d

Observation 8a622607-903e-4f87-a51b-ed3b40ddfd4f · outbound

This paper cites QLoRA: Efficient Finetun- ing of Quantized LLMs.

NVILA: Efficient Frontier Visual Language Models QLoRA: Efficient Finetun- ing of Quantized LLMs

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.187526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:6a4f6d2a7d444d6e660bfde86d5765b8538750b4dba32f48decce2cf9de9561a

Observation 6d67f1e4-b2a7-4c76-8162-f911cb025cdd · outbound

This paper cites GaLore: Memory-Efficient LLM Training by Gradi- ent Low-Rank Projection.

NVILA: Efficient Frontier Visual Language Models GaLore: Memory-Efficient LLM Training by Gradi- ent Low-Rank Projection

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.077516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:721ba834d0c5107aaff4cfd700d89f77552d40c512e4c1b17b23252005949fa0

Observation 12465fea-7643-4fdb-a10f-7759182e5100 · outbound

This paper cites ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models.

NVILA: Efficient Frontier Visual Language Models ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-23T22:20:22.000193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:4055379b03670b2d7b7601f4e61d42b57d80bbca0bcb8c13659dfa129feb4740

Observation 2ca4f020-57c9-4508-95f9-499a99c5fe63 · outbound

This paper cites Building and better understanding vision-language models: insights and future directions.

NVILA: Efficient Frontier Visual Language Models Building and better understanding vision-language models: insights and future directions

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:42:43.129247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:5139bcf445a52c7e47e11858485d2b8c174890176bfe4c3b5e8097e85c476318

Observation f58c8bef-486f-48b2-a869-2c475a052ca8 · outbound

This paper cites PDF Associ- ation Dataset (PDFA).

NVILA: Efficient Frontier Visual Language Models PDF Associ- ation Dataset (PDFA)

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.069065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:3b33ddfa504db122ef8ae8f87cf4239fc1a93eaef5cdfdd629fd2b630f9e58e0

Observation 68029f20-cedd-4d18-b4ab-e4e25e70a08d · outbound

This paper cites ICDAR 2019 Competition on Large-Scale Street View Text with Partial Labeling.

NVILA: Efficient Frontier Visual Language Models ICDAR 2019 Competition on Large-Scale Street View Text with Partial Labeling

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.073587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:e0cc5c0c556d5a807c89842c339915480779572da82581ba565d6d3dd6d369d5

Observation fba87277-fe75-4583-95ac-b8fe163a77fb · outbound

This paper cites ICDAR 2019 Robust Reading Challenge on Arbitrary-Shaped Text.

NVILA: Efficient Frontier Visual Language Models ICDAR 2019 Robust Reading Challenge on Arbitrary-Shaped Text

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.081203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:4c7f8064bd972116bb2626e86847a9ba3a99bd4c568040a02c490f194f28348f

Observation 5a6750c8-025c-4166-bb09-d605f1438115 · outbound

This paper cites COYO-700M: Image-Text Pair Dataset.

NVILA: Efficient Frontier Visual Language Models COYO-700M: Image-Text Pair Dataset

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-23T07:42:45.085385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-23T07:42:22.478647Z digest=sha256:09e1c51dec09b6b3926a3f9c4798f8b8d429166f923a946d252a20e1d58861ba

Pith citing papers

Observation 5b4a343d-7fb3-42fe-a459-72d59a9ef5db · inbound

MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models cites this paper.

MLAN: Language-Based Instruction Tuning Preserves and Transfers Knowledge in Multimodal Language Models NVILA: Efficient Frontier Visual Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T19:40:42.628347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:40:42.628347Z digest=sha256:eeba7eb0fddcbe803ad1ef0562f7f42483c3f18cb4a708c1fe2b4b533df4d6af

Observation e86d5a3a-83de-45de-8fdc-b8cae8194041 · inbound

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning cites this paper.

OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning NVILA: Efficient Frontier Visual Language Models

Reference 147

Resolution
verified exact
local_arxiv, observed 2026-05-17T20:33:26.913053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T20:33:26.613927Z digest=sha256:1b298e6f1edbf6824a6807ee0f4f4810eeeb804a4d446c489d968a23b312178a

Observation cbd5ebb6-c646-462d-9ea0-bfbe0c282b7a · inbound

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models cites this paper.

FrameFusion: Combining Similarity and Importance for Video Token Reduction on Large Vision Language Models NVILA: Efficient Frontier Visual Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:09:25.074549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:09:25.074549Z digest=sha256:be3fd60e885ed9407e1dd5ca4fdae0d5653d8f0a4a1a1d3ee218771916c592b4

Observation dbb8d4ec-71b1-42a0-8a5e-08805a4c0a71 · inbound

ViDDAR: Vision Language Model-Based Task-Detrimental Content Detection for Augmented Reality cites this paper.

ViDDAR: Vision Language Model-Based Task-Detrimental Content Detection for Augmented Reality NVILA: Efficient Frontier Visual Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T17:08:56.455169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T17:08:56.455169Z digest=sha256:edc4f40e6c3db9ae02c6d8f57887fba7f6d9613a8d3bd760e680f047b75bbc57

Observation 643ac8e0-6a44-48cd-b137-d9331bdd8c8b · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding NVILA: Efficient Frontier Visual Language Models

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-11T01:19:59.947092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:e0d2d332872db250d991cdd2c9581bba4e4391f4a09a325a764d03ebffcc79f9

Observation d33a54a7-9b41-4224-8a20-4010fe75f9eb · inbound

Temporal Preference Optimization for Long-Form Video Understanding cites this paper.

Temporal Preference Optimization for Long-Form Video Understanding NVILA: Efficient Frontier Visual Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-10T15:35:30.199392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T15:35:30.199392Z digest=sha256:7b3f940bdcfb776b6d7089c1fb5826da93dcec35d1955ac45e7a6484b3b79be5

Observation 241e713a-6af2-45eb-8fb4-b577a8a5af47 · inbound

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models cites this paper.

Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models NVILA: Efficient Frontier Visual Language Models

Reference 199

Resolution
unresolved
no resolver link, observed 2026-08-10T18:04:34.873401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:04:34.873401Z digest=sha256:bfcf6f637aab651b1c8eeea773ba09aab1cc2b02e0a466a9debdaf721171a8cb

Observation eedac329-4758-4a4c-8b3b-004a336318fa · inbound

SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer cites this paper.

SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer NVILA: Efficient Frontier Visual Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T23:37:12.567981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T23:37:12.567981Z digest=sha256:7c9fc0ad7e5ed146274a67a516e9093e1784f32d75dbda9b2701d11ac6eff072

Observation 3a066575-8af0-4f62-a543-226beb93ca90 · inbound

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models cites this paper.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models NVILA: Efficient Frontier Visual Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.906231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.906231Z digest=sha256:6c72eff543a5e7fb3a174bfee39b503190a4a0fc85f904dc3b8121aa23cfb307

Observation e6297378-2b6d-428a-b06c-488d05e77eee · inbound

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models cites this paper.

LV-XAttn: Distributed Cross-Attention for Long Visual Inputs in Multimodal Large Language Models NVILA: Efficient Frontier Visual Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T12:23:40.532423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T12:23:40.532423Z digest=sha256:e4db9bf56cb8e3f354c05e02c6c481c3f1a1c7ccf7e3abd455413885f291b1ae

Observation c3c83d3d-f68a-4096-a917-9d2541186504 · inbound

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs cites this paper.

WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs NVILA: Efficient Frontier Visual Language Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:53:26.341702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T05:53:26.066674Z digest=sha256:c83997c5b1f3d81f7ea71fb24c943be876e2a568f706b8c5076cae751fe51deb

Observation b6368733-9fef-477b-b20f-a6c9b9c6ef7a · inbound

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation cites this paper.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation NVILA: Efficient Frontier Visual Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.126661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.126661Z digest=sha256:4b5265bc982ff8acbd69df21bffe8a3adcb8113220c83a2a4cd19c942c511e14

Observation a8dab4d0-57e2-4045-bbef-6ea7cf7dc34d · inbound

Training-Free Reasoning and Reflection in MLLMs cites this paper.

Training-Free Reasoning and Reflection in MLLMs NVILA: Efficient Frontier Visual Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:13:56.772463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:13:56.772463Z digest=sha256:711fea7d4ce391b0e89bf4ee4b3501db95ae785f580fd786559ff029cc535a5d

Observation 65a724fb-6abd-4279-8576-c22f418abf0d · inbound

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning cites this paper.

Bridging the Dynamic Perception Gap: Training-Free Draft Chain-of-Thought for Dynamic Multimodal Spatial Reasoning NVILA: Efficient Frontier Visual Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:02:25.841387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:02:25.841387Z digest=sha256:3960d98adbe697167505899f096839bed1c6615366c1027e216774a4cda63077

Observation b46feefe-1245-44fc-8ce9-c9a71c63c07d · inbound

LaViDa: A Large Diffusion Language Model for Multimodal Understanding cites this paper.

LaViDa: A Large Diffusion Language Model for Multimodal Understanding NVILA: Efficient Frontier Visual Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:59:36.482744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:59:36.482744Z digest=sha256:04fb026f3ae8f04a02ba97c338f3fbe4010a8e3d2ca03fe1730ca67fc98011f1

Observation ef6700ac-f5f4-4846-b60a-0cefd3575d1f · inbound

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities cites this paper.

FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities NVILA: Efficient Frontier Visual Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T14:05:01.704476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:05:01.704476Z digest=sha256:a4bf93c96160be30315f2e910160333cf966a2dedc6bb8d87da23c84d021b933

Observation 8a9d3f84-1937-4239-94cd-d89cce996e5c · inbound

ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models cites this paper.

ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models NVILA: Efficient Frontier Visual Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:33:56.166173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:33:56.166173Z digest=sha256:12877820d99d45a82b616394200d33ca32b176d0cd7e851c0627ed6055d0824c

Observation b8f94370-a633-46b5-aad4-a98a9265096e · inbound

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation cites this paper.

VCapsBench: A Large-scale Fine-grained Benchmark for Video Caption Quality Evaluation NVILA: Efficient Frontier Visual Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:07.044736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:48:07.044736Z digest=sha256:7fde029ca9b3c1ecbbb533d488405340e497c4bd2095f128456e4cc75fb91c68

Observation df08e76c-feda-4625-aff8-133bf071001d · inbound

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models cites this paper.

EffiVLM-BENCH: A Comprehensive Benchmark for Evaluating Training-Free Acceleration in Large Vision-Language Models NVILA: Efficient Frontier Visual Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:09:06.218354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:09:06.218354Z digest=sha256:ee85a3da705a57303396db1987fa8851f7a3d52ed7e60da6698eb115d0baeb0c

Observation 53815c49-612f-4fdc-ace2-72e165abdbdc · inbound

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding cites this paper.

FlexSelect: Flexible Token Selection for Efficient Long Video Understanding NVILA: Efficient Frontier Visual Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:59:07.844726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:59:07.844726Z digest=sha256:8f21f42dbae22846c4d4e0d8fd34fd70cfd258d14c69173f66ea5c1447dc349a

Observation bc85cb85-975e-4319-8afc-377612f6b484 · inbound

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping cites this paper.

MINT: Multimodal Instruction Tuning with Multimodal Interaction Grouping NVILA: Efficient Frontier Visual Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:10.410916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:10.410916Z digest=sha256:d8d0312740f97a82a5ab98da94f8b93bdc97435c06cb2ed5bbf9a0d2bb34634e

Observation 54f3858b-f0ae-4f61-9f1b-9fe9c8f76c74 · inbound

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation cites this paper.

AnimeShooter: A Multi-Shot Animation Dataset for Reference-Guided Video Generation NVILA: Efficient Frontier Visual Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:12:10.219918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:12:10.219918Z digest=sha256:c55c9bd55de2fa8cbe7806d43c77164caf16e7b0c8caba0748bc03073f5e183a

Observation 56673960-7564-4da5-a3a2-2610c81da451 · inbound

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos cites this paper.

MMR-V: What's Left Unsaid? A Benchmark for Multimodal Deep Reasoning in Videos NVILA: Efficient Frontier Visual Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T10:52:35.682714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:52:35.682714Z digest=sha256:121a5a4b100bd753209281d7c62e53ad80b264c48865d02be038e8112e94eb38

Observation 4e1d889c-7dd5-42e2-af56-33384038d571 · inbound

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? cites this paper.

EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? NVILA: Efficient Frontier Visual Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T10:29:42.990740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:29:42.990740Z digest=sha256:224fecb4da376db76ddb7beac2bdf6290e1a6e52fcc0281c6a6fb199e3b4b5e6

Observation e3b33f33-822a-4f75-b31c-5aaa08be4c8f · inbound

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos cites this paper.

EgoVLA: Learning Vision-Language-Action Models from Egocentric Human Videos NVILA: Efficient Frontier Visual Language Models

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-21T04:32:58.776597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T04:32:58.733165Z digest=sha256:c948ea5ce221e671e9a5c449f7812080627c1642c757bd841db465165f409758

Observation d9fd4ac1-42a1-4395-ae53-7f3282d22980 · inbound

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning cites this paper.

ThinkAct: Vision-Language-Action Reasoning via Reinforced Visual Latent Planning NVILA: Efficient Frontier Visual Language Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-19T03:22:00.959761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-19T03:18:14.655384Z digest=sha256:9f2b092b5f6601ad8dfe15d5793d05ab6b0ba0859858da4135fc48284435bea2

Observation 94e994f1-c11d-493c-8020-9568fcb97295 · inbound

Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring cites this paper.

Exploring the Application of Visual Question Answering (VQA) for Classroom Activity Monitoring NVILA: Efficient Frontier Visual Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T11:49:10.455295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:49:10.455295Z digest=sha256:c617b6cec7e2a3cd2af9fb4b8e0532d13414c05837170a190ace741ee4afa999

Observation dcae26e5-9d74-4f1b-b725-9d1175c24896 · inbound

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning cites this paper.

MagicVL-2B: Empowering Vision-Language Models on Mobile Devices with Lightweight Visual Encoders via Curriculum Learning NVILA: Efficient Frontier Visual Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T05:36:19.512240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:36:19.512240Z digest=sha256:e3d47be7e74037364b1f96faf8f0c847ce09c7a9d4bbe39626b516e5f267a4de

Observation ae0a074e-4087-4795-8e9b-af2c3d3bd5ea · inbound

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes cites this paper.

HumanPCR: Probing MLLM Capabilities in Diverse Human-Centric Scenes NVILA: Efficient Frontier Visual Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-05T19:03:10.742012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:03:10.742012Z digest=sha256:16850438fb479bdb44f8e8db6f9740541e2141d64d1e5898995f4faa656558f4

Observation c91c6ae7-578e-425a-949f-7cf55e871954 · inbound

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation cites this paper.

Progress-Think: Semantic Progress Reasoning for Vision-Language Navigation NVILA: Efficient Frontier Visual Language Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-17T21:05:15.963769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-17T21:02:38.013115Z digest=sha256:45f3e6f1d88cd568de0938c12fc8748093a3439e516faab42c9752fa89a8c05e

Observation 19fd7d4d-a3e7-43ca-ab85-6c1d59da13d0 · inbound

DODO: Discrete OCR Diffusion Models cites this paper.

DODO: Discrete OCR Diffusion Models NVILA: Efficient Frontier Visual Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T22:27:42.280107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:27:42.280107Z digest=sha256:85178afc08ae7668bfd689f74a124a57e72b118da8add63e6a769e6b5655beb1

Observation 146ecfb9-4eba-431d-b9dc-8014c112f4e7 · inbound

XNote: Benchmarking Automated Community Notes Generation for Image-based Contextual Deception cites this paper.

XNote: Benchmarking Automated Community Notes Generation for Image-based Contextual Deception NVILA: Efficient Frontier Visual Language Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-21T10:40:00.517722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-21T10:39:09.209094Z digest=sha256:9e1bcf5799925e769b2af239015bd7ae0ea84d60626c3c172af47d19a1e9fd25

Observation 2b981b61-9835-4d9f-ad5f-fac46f34d757 · inbound

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning cites this paper.

A Progressive Training Strategy for Vision-Language Models to Counteract Spatio-Temporal Hallucinations in Embodied Reasoning NVILA: Efficient Frontier Visual Language Models

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-11T09:41:01.759826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T15:53:52.775348Z digest=sha256:b3ba37d43bb540fde32ac67a8538c6c184d8668053a9fe775393ff4db56c5e69

Observation 5d3ba47f-3898-46ba-a810-ee075bf51136 · inbound

OProver: A Unified Framework for Agentic Formal Theorem Proving cites this paper.

OProver: A Unified Framework for Agentic Formal Theorem Proving NVILA: Efficient Frontier Visual Language Models

Reference 63

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T14:48:23.383955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-20T14:43:46.517807Z digest=sha256:6b19e2e931a4133875684fec98df06cdfc2192f1238b8c1e9d85c12021716218

Observation 827d637a-0d04-46af-bae8-fcc3215e2e08 · inbound

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data cites this paper.

StableVLA: Towards Robust Vision-Language-Action Models without Extra Data NVILA: Efficient Frontier Visual Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-20T11:58:15.155933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T11:55:22.498055Z digest=sha256:d4660c53209a593f89c02c1c55e4cb943c2c85694ce776f070b510ff7cbe6efe

Observation e9b193be-822e-4e95-b0ae-5d5df97a8910 · inbound

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding cites this paper.

Enhancing Visual Token Representations for Video Large Language Models via Training-Free Spatial-Temporal Pooling and Gridding NVILA: Efficient Frontier Visual Language Models

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T06:21:10.873047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T06:16:32.972099Z digest=sha256:f650612125e3b71c5c592cdc5a6e467d833737b9214db4e4510c93a8356d4b29

Observation e9047406-2dbb-4c90-889d-46404d94c104 · inbound

ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs cites this paper.

ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs NVILA: Efficient Frontier Visual Language Models

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T06:21:10.806191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T06:16:38.589260Z digest=sha256:856b3fb49cc48cb8ec175680f80c2d3291c9e6fc2d456b4898968e0b9899e10c

Observation f746aa27-d10a-4bf0-87c5-abcb2ad6eb0f · inbound

Worth Remembering: Surprise-Gated Robot Episodic Memory cites this paper.

Worth Remembering: Surprise-Gated Robot Episodic Memory NVILA: Efficient Frontier Visual Language Models

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:56:34.558067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T09:37:15.301043Z digest=sha256:5b221fd6c99f751cab9eb630daf4968de68b2cd9fd2b1d13ae403c477d86f45a

Observation 1cc176a4-3e76-4d97-9237-96526862d2ee · inbound

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding cites this paper.

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding NVILA: Efficient Frontier Visual Language Models

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-04T16:39:58.307382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T00:19:26.153682Z digest=sha256:2854375243336d193bf6205ba008be17eeb1e86d8dfa6202dce5cda2a0ff84e2

Observation 89826ebe-5819-4ffd-a570-9dbf1521d10b · inbound

RADIO1D: Elastic Representations for Condensed Vision Modeling cites this paper.

RADIO1D: Elastic Representations for Condensed Vision Modeling NVILA: Efficient Frontier Visual Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:14d7f72a589f56b5a293b46a4a5f9d7950bce2971cd8664ba95fd1aed9266c04

Observation d17f9179-1967-46e3-a4f5-ef6e5753af6e · inbound

QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding cites this paper.

QSVideo: Query-Conditioned Semantic Temporal Retrieval for Video Understanding NVILA: Efficient Frontier Visual Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-11T17:18:41.284513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T17:18:41.284513Z digest=sha256:0443a04e31a5d082bdfebe4b8ca01b0a0427b832ebea9978e6dfb6ebe56ccfb4

Observation ea0d2259-c377-4a65-9e2d-3c4fe634035e · inbound

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization cites this paper.

EmoAgent-R1: Towards Multimodal Emotion Understanding with Reinforcement Learning-based Dynamic Agent Specialization NVILA: Efficient Frontier Visual Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T08:49:13.405920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:49:13.405920Z digest=sha256:298a88dbd72c50ecf4b67c25366e61bd8444ff67da1f92667b950e1b2b9adf28

Observation db80bc62-7e7f-4b6a-b90a-06ce968b1f47 · inbound

Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles cites this paper.

Distilling Vision-Language Models for Robust Traffic Sign Perception in Autonomous Vehicles NVILA: Efficient Frontier Visual Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T04:28:26.603246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:28:26.603246Z digest=sha256:6306eda0d4d5fcb2cc5b49aa09102183b16fe9a825d17c49b72ef9ca6db39989