Pith. sign in

Paper Citation Record · LEDGER

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference

As of 18 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 3 inbound Pith citation observations for arXiv:2412.11475.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11475 v3

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T14:56:25.274118Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:13:55.100419Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-16T10:13:25.893446Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ec0e5239-7dd6-4732-b339-de7ac23051de · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T14:56:25.183742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:56:25.183742Z digest=sha256:28d50e818f83a354e4ce66f1c1631ec1039282cc37e429d44ea472bc500b1c80

Observation 1f0d4402-6331-47f9-94f2-7efc9df631d6 · outbound

This paper cites Squid: Long Context as a New Modality for Energy-Efficient On-Device Language Models.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference Squid: Long Context as a New Modality for Energy-Efficient On-Device Language Models

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-11T14:56:25.479973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:56:25.202850Z digest=sha256:a572dcca594b6aaa8af72e9af8b9fd5df01c951e90667a382bc072c66a4070fa

Observation 6fbdb3aa-9b2e-474b-9c87-03e0955f30d8 · outbound

This paper cites Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T14:56:25.208369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:56:25.208369Z digest=sha256:f78127e6007f16467808221b939d0cb606a6f3bebad21e49291f4a28f641087e

Observation 63199eca-cfec-4be8-bdce-b47ddd994430 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T14:56:25.214844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:56:25.214844Z digest=sha256:beaa70d8b9798c1e6dbb9fab1b52b93ea8c5c305c94fb0a3facf01cca9f19ce9

Observation 496b8a2e-bd61-47f1-b964-64b98ae0a7ad · outbound

This paper cites Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference Ferret-UI 2: Mastering Universal User Interface Understanding Across Platforms

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T14:56:25.225612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:56:25.225612Z digest=sha256:a17769493018b4b8329be23ef2cde9ff78b3f04bd0e053e2260516178e5cd9ff

Observation f2352a00-6adc-4510-ae19-c616f3862b2f · outbound

This paper cites MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T14:56:25.236951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:56:25.236951Z digest=sha256:5f12cb0fc859992f857579a99fefb8de790f7d1136b7e53e750131d885de721c

Observation 4a1886ed-1fba-44c1-b721-05aada00d860 · outbound

This paper cites Introducing orion, our first true augmented reality glasses, 2024b.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference Introducing orion, our first true augmented reality glasses, 2024b

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:56:25.589922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:56:25.247829Z digest=sha256:cb43b0201f30d118da331d8e2a5f4a3822d22a748dccb2617d928673ff1a66eb

Observation ebb69450-b345-4078-a341-33f7d557fb7a · outbound

This paper cites MLC-LLM, 2023-2024.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference MLC-LLM, 2023-2024

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:56:25.573347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:56:25.253485Z digest=sha256:87e047fa661bdbf915412d4961b885bffbda5681cfa62de9108fd70d747e2481

Observation 3d6cf64b-e886-48b7-b655-d3e64e53dee7 · outbound

This paper cites Ollama, 2023-2024.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference Ollama, 2023-2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T14:56:25.553554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:56:25.258419Z digest=sha256:e09151695b4e6464c738bfad671c35665afc98484a3e503a8b9300b5cd51a71a

Observation 4b3f9571-f3a6-4321-87be-95806625a7a4 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference Gemma: Open Models Based on Gemini Research and Technology

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T14:56:25.264091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:56:25.264091Z digest=sha256:f281887f2706784dc34f68b4563df4307f44df35d540a1aac44afb2e8edf8745

Observation 0069e66d-5914-41fd-9b25-78e40bd0ce71 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T14:56:25.269637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:56:25.269637Z digest=sha256:c70fff138182d7b3d2311aa129d3abd01dbfaa189fdfc6c66b81d0f0ed8b658e

Observation ab87ea45-a88f-423b-b55c-b993cb534dd1 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference Sigmoid Loss for Language Image Pre-Training

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T14:56:25.274118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:56:25.274118Z digest=sha256:1114786fa473f4a97b8cfdb9e4a6a74199510cd8dd644cbe635f8d97227b330e

Observation d8a55622-171e-420c-9ed2-34191e0a8516 · outbound

This paper cites The Llama 3 Herd of Models.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference The Llama 3 Herd of Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T14:56:25.220199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:56:25.220199Z digest=sha256:31fb73d356a0d993ebc031c3f6005c2b3a9bc2e56f0b0e2823b477e3baaf9ef6

Observation 0c22894f-f28d-41a2-8d7c-f5986ea75aac · outbound

This paper cites an unresolved cited work.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-11T14:56:25.607247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-11T14:56:25.242212Z digest=sha256:ff764ccf9e06ddc5d60cea9eed15cd3c8429903bb76778e9a22335e13e5b861c

Observation c259c23c-8566-4ad1-bea6-68f281cee6dd · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference PaliGemma: A versatile 3B VLM for transfer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T14:56:25.190474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:56:25.190474Z digest=sha256:3ae00ed10ee154a1e03089fafdbfe959e4f5d58cff9ecc8eaf287da7936488bf

Observation d4fd3c01-5ede-4f6a-9ed0-b8e86d0f0e5b · outbound

This paper cites Octopus v2: On-device language model for super agent.

OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference Octopus v2: On-device language model for super agent

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T14:56:25.197362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:56:25.197362Z digest=sha256:f29b78bae3b77662d106f50a8adfc635016fe90a38c214a6dbd87a66befc2cba

Pith citing papers

Observation 55649044-730d-41c5-bbe2-e7f7f061fbba · inbound

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety cites this paper.

Towards a Multi-Agent Vision-Language System for Zero-Shot Novel Hazardous Object Detection for Autonomous Driving Safety OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:13:55.100419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:13:55.100419Z digest=sha256:30debea1a970c20610619bed688f63dbdfec2ed2ee624e7754f3be903c258b11

Observation 8febbffd-7988-46e7-b6ce-d4a41c792e9f · inbound

A Review of 3D Object Detection with Vision-Language Models cites this paper.

A Review of 3D Object Detection with Vision-Language Models OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference

Reference 2022

Resolution
metadata mismatch
local_arxiv, observed 2026-08-16T10:13:26.010057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T10:13:25.113330Z digest=sha256:acefebdff94564be04ca470c2598d417d5273e574ef5a495172293fe49e3710f

Observation 77ac84e2-5936-4016-a7af-df25475e5e91 · inbound

AutoNeural: Co-Designing Vision-Language Models for NPU Inference cites this paper.

AutoNeural: Co-Designing Vision-Language Models for NPU Inference OmniVLM: A Token-Compressed, Sub-Billion-Parameter Vision-Language Model for Efficient On-Device Inference

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T18:56:56.607192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:56:56.607192Z digest=sha256:bc3bf7bafce621e18d1e70502f9e734ab5ddbd893ae22221ba5549221b705705