Pith. sign in

Paper Citation Record · LEDGER

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models

As of 8 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2607.23373.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23373 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T00:01:39.867668Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved69
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 094bf45c-ea4c-4235-a8c0-cd09ff089442 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:32.624103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:32.624103Z digest=sha256:55bccae8c93d18adeeaa9499a454924231e5d25efd3dacd4ba92039b489b7e36

Observation f1d37e2b-ebad-45e0-a7aa-d5dc00248eb2 · outbound

This paper cites In: Proceedings of the AAAI Conference on Artificial In- telligence.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the AAAI Conference on Artificial In- telligence

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:32.725286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:32.725286Z digest=sha256:771211369d1c2d85b2f4657c62dad78b470d9563df77867fff94788d87a51e70

Observation 965803ed-75c0-4000-891e-cb6e51bad4ff · outbound

This paper cites Qwen3-VL Technical Report.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:32.858371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:32.858371Z digest=sha256:cf1270d8a0a6fa479a43c0b934633600032f7163ceb76fa83bc0ea71c7714955

Observation ff630cb5-78ae-4394-97c4-266bfe7f780e · outbound

This paper cites Qwen2.5-VL Technical Report.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:32.922639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:32.922639Z digest=sha256:904d108d9475e494725f403607e3f1a035662cc11f68a56a77485e9ce2627815

Observation 07347f26-0010-46dc-a137-c82a5de590c0 · outbound

This paper cites Longformer: The Long-Document Transformer.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Longformer: The Long-Document Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.003935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.003935Z digest=sha256:b86ca127cfdfcf3b47b10a80fa0f5f716133dcb6e31c0bb8fd2704d78e8137f2

Observation 5be5d2c5-d832-4de9-8fa8-ff1cbf161ede · outbound

This paper cites Perception Encoder: The best visual embeddings are not at the output of the network.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Perception Encoder: The best visual embeddings are not at the output of the network

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.119140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.119140Z digest=sha256:70f0959abf3c1b408daef582101dcc7f0f2f064335fc63ac856a62e558984043

Observation d24f56e4-f760-42ae-8421-ae9c615a4d8c · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.171167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.171167Z digest=sha256:f446b47c53d85c663f65bed366703f9e894f5e63f572224f6f66744f80117d2a

Observation 5354d74b-5494-41bc-a951-10a73d61707e · outbound

This paper cites In: The Thirty-ninth Annual Conference on Neural Information Processing Systems (2025).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: The Thirty-ninth Annual Conference on Neural Information Processing Systems (2025)

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.245005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.245005Z digest=sha256:87cc8cb92862ecdb1aff58a3992afbff931904b8dc7647fe1319504aff79b252

Observation d659ef2f-90cb-45e5-99d1-923c05e3f84b · outbound

This paper cites Matryoshka Multimodal Models.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Matryoshka Multimodal Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.328665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.328665Z digest=sha256:46eddd3e7c8bc256aa9f3219aeab35e0b6e512b03a58eed1769240ded773de1f

Observation b80b76b5-a54f-4290-88b2-59a35f89dfad · outbound

This paper cites In: ICLR (2025).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: ICLR (2025)

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.418212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.418212Z digest=sha256:b6a2a993a44cfc6c316f185fb28a843966afe8043d00aac9bdfe57678e75bd62

Observation cc0bd62c-9fc6-4c5a-ad42-aaed435e4fd3 · outbound

This paper cites an unresolved cited work.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.516741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.516741Z digest=sha256:62b200d0d3259ead8689192f2acb91a969bb26d0ac5a99eb9d16466e0a943c35

Observation f31595c2-338e-4243-870f-03797b5e21e3 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Generating Long Sequences with Sparse Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.623542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.623542Z digest=sha256:7d0a210aa3086fd31b8c0e48a6ed6cddb26f20d3446d61cce327c222d34a09ed

Observation ccfa5a24-c162-4fba-9837-a13979c35d7a · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.718457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.718457Z digest=sha256:8d0aca22db6c89e8a84823394fab79a87a833c7bbe91e44d8dd01f7fd8f8d97c

Observation 966b53dd-de11-41d0-ba08-6fe1fc8d6a5d · outbound

This paper cites In: Proceedings of the IEEE/CVF International Confer- ence on Computer Vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF International Confer- ence on Computer Vision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.803370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.803370Z digest=sha256:c906efd89808c7865a570095dbff30cccb2b49b0a8ea07f5e987a7f863c20bbe

Observation f8c973fc-007f-411c-ac44-3ca37dce4eef · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.864979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.864979Z digest=sha256:408b3233c22e35c3e0aaac780e854d9ab69941998bb4c09783794ff4675a5418

Observation 898f5330-0203-4f9f-8d82-d5b332a18204 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.911479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.911479Z digest=sha256:af500961a60e70c1f7c960bf05e8e6ba7af78f91c186de234be5fc5e5f94dec1

Observation 304e0e1a-e400-486c-bb2e-85c4149e3d25 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.050706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.050706Z digest=sha256:30082f23282ae2d7a47e30e86c7f5d4ca3c826e0da53c58c99f4a21f9e8ac499

Observation 8920c283-0f8c-49ef-a77c-6d7068b26b8e · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.113031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.113031Z digest=sha256:cb71326aabb4f6ba63030f575a0fdf961fd343aa6bc388944efc1d6a3efffb66

Observation c3fb02bf-5c23-40ce-91cf-149d5d25eb02 · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.211569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.211569Z digest=sha256:bb97badcca66eede286d1858ff491884e88baeb5e508606222cd73dbb2c67efb

Observation c9262c34-fb72-4534-aed5-7ebca176928d · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.326764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.326764Z digest=sha256:86225c7acd0a4527e35ff24e7d5eb22eddc9646280238f53cdda4188c75ec7ad

Observation 902cb00f-32af-4c09-9427-9dad7aedafb2 · outbound

This paper cites Matryoshka Query Transformer for Large Vision-Language Models.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Matryoshka Query Transformer for Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.400885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.400885Z digest=sha256:41f326478ca8d66416313dce28f80bf88b4c64b504efa7ce0c603b2f92a0ba06

Observation d6084cd6-9915-4075-8734-7891f403498d · outbound

This paper cites In: Proceedings of the IEEE/CVF confer- ence on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF confer- ence on computer vision and pattern recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.480312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.480312Z digest=sha256:8bb46969053c360c79f316eb596bd2a8be87f44a65a98dc332634478d66478f7

Observation 690679a4-1fea-4ca2-9699-b901c36c8feb · outbound

This paper cites In: European conference on computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: European conference on computer vision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.628057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.628057Z digest=sha256:06fa2187c07a881873910c45351952084a2fe1a3a42ef7e81e826d5fa7c8e605

Observation e9675575-2c75-4aa4-a505-0b5133cc54cd · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.711225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.711225Z digest=sha256:b1292bb61b712780e9573016945239d452713f1721ef7892d468148306f08608

Observation e62b6a09-3d5c-40f4-9451-00bf76278783 · outbound

This paper cites In: Proceedings of the 2023 conference on empirical methods in natural language processing.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the 2023 conference on empirical methods in natural language processing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.779557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.779557Z digest=sha256:a3cd6fa0756b757aeb3b09a6cd10415dc1161fb5aec35640755a96c7c88234f1

Observation 9313867b-121f-4fb1-8843-5667837e5615 · outbound

This paper cites an unresolved cited work.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.854510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.854510Z digest=sha256:4335dd968a9e3f679bae5fa3539b80b15865b9986f81be2b83c4b694a8fbf96a

Observation bf7eea37-43dc-4d33-9d9d-ab3ab76ff443 · outbound

This paper cites an unresolved cited work.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.941159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.941159Z digest=sha256:21abf5a0557079e9066100cbb0f44609d7bb727d7b4b611ad97c282147cb8f4c

Observation 1a5f5f1c-d61b-4d93-ba34-c416777594c5 · outbound

This paper cites Science China Information Sciences67(12), 220102 (2024).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Science China Information Sciences67(12), 220102 (2024)

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.043229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.043229Z digest=sha256:2657609bd9008e615f21cfa179d977cd9612ab936da77620f1c7205c9b025223

Observation 6b5350e6-8f23-4bca-9f1f-b7bb25cca713 · outbound

This paper cites an unresolved cited work.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.188413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.188413Z digest=sha256:c9f215d39e107f387c27707a8788c5f2843971457561702bcd7ba288cb0a64ef

Observation dac0d11d-68b8-4d53-a6df-1d698f3ad44e · outbound

This paper cites Advances in neural information processing systems35, 2507– 2521 (2022).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Advances in neural information processing systems35, 2507– 2521 (2022)

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.260051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.260051Z digest=sha256:4a7ba8a3753f5975728fb1d6b8ce2777e5eb19b503f425bf2365ffc950a5e4a7

Observation 518b8411-0511-451d-9c2f-3f7e386c7d83 · outbound

This paper cites In: Proceedings of the European conference on computer vision (ECCV).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the European conference on computer vision (ECCV)

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.361614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.361614Z digest=sha256:90fea6f1f1815c0957fb5c4e3da623c12bb45f45a773b6b22d3125c072a71e23

Observation 90f639cf-4016-478a-88b1-479bd5ad07ca · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models SmolVLM: Redefining small and efficient multimodal models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.403324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.403324Z digest=sha256:233942f4b2a3b79354c865abc3ecd3a8315464a7eecfa42fcfda9e8eb16c3316

Observation 0e022fb1-b882-4283-aa1d-3587102bf34d · outbound

This paper cites In: Findings of the association for computational linguistics: ACL 2022.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Findings of the association for computational linguistics: ACL 2022

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.480435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.480435Z digest=sha256:de61cf45a1bf8bad1a454729deb8a87e7f6a93618c27fe8e5204e9f44b802b68

Observation 1ef4d817-01bb-415c-b67f-34fc60606b83 · outbound

This paper cites In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.564805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.564805Z digest=sha256:6181b8105cc09af939d0186c387b7bae681c178e9fbc36d4b44393d26aecf2ef

Observation b0d7c3ed-df52-414c-b6dc-69d9332eef8e · outbound

This paper cites In: Proceedings of the IEEE/CVF winter conference on applications of computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF winter conference on applications of computer vision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.625122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.625122Z digest=sha256:ec4eb4973ed05c1ac08d53d8a1c00cad4470a47551526eb0a6af4888720f2caf

Observation efdc97a7-c351-4f85-8dea-eaeb47bfb5e6 · outbound

This paper cites Separable Self-attention for Mobile Vision Transformers.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Separable Self-attention for Mobile Vision Transformers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.791774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.791774Z digest=sha256:f5e595fcd30c988d591a28d23cc59461f7bdc25220faad3d3214dba6ff41cf6b

Observation fec8168e-c8aa-4f72-88ff-aadeb31d509b · outbound

This paper cites In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.933720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.933720Z digest=sha256:1f6905609ee8d2f511d07b585e0f785e64b6023ac0589dc9db7a53286e2e2db3

Observation 392e581b-4f1c-4446-b7fb-56f766bf4b95 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.053011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.053011Z digest=sha256:a207df99707828c6a666af88cd38102700222e2db8119d6a743476eb1799b1d2

Observation 66e7714e-842d-400b-a8c8-389fec46d6ea · outbound

This paper cites In: European conference on computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: European conference on computer vision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.212439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.212439Z digest=sha256:a52afa5aa426331668b74741129b11dd521b58235332b5199e3e321b2f1a38c0

Observation a9979fba-d67b-4110-9b91-9c616ac07627 · outbound

This paper cites Advances in neural information processing sys- tems32(2019).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Advances in neural information processing sys- tems32(2019)

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.327194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.327194Z digest=sha256:7db202444a0974c6b47ca5ecde8785a892947e761b5b0d437e7b61c585a2a34c

Observation 4f957f63-a273-4a61-8e57-1132151a4b4d · outbound

This paper cites In: International conference on machine learning.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: International conference on machine learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.383614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.383614Z digest=sha256:bc14d768c36e291952b204dd1625c1537e777107fe69a3dda05833baff1c35d1

Observation ac691506-1418-418f-8642-8f1a39d66ff0 · outbound

This paper cites In: Proceed- ings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceed- ings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.578277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.578277Z digest=sha256:98d38f83821856c66c54a8193ddc4b2441b4cd7bcdc553f021474355119f9d2e

Observation e08d5a83-4b32-47d6-99b5-eb49a7e53b41 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.708581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.708581Z digest=sha256:dc27116f046a956b97a3dac91ca2232df2745c7448364819f7b25b15a2984937

Observation 8a8ec369-3840-4719-8f29-940fd50b73b5 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.832203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.832203Z digest=sha256:ed567fd24249ba1ee882294220d5243da6bbdd6852d1d53614d7594eadf17c1a

Observation b761f5b3-a653-4cf2-a64b-2b9000002e1c · outbound

This paper cites In: European conference on computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: European conference on computer vision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.985781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.985781Z digest=sha256:b082ad7eabdfb0ff8460cc277d4e8f7a849dc0563f5485884bb719c89b7f46a4

Observation 02974f28-d28a-42dc-84c6-9b21e8e4d729 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.103468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.103468Z digest=sha256:7c4c19646bb5e1ad946295fcb1008452ca6c62734562f098d109a193c98ee299

Observation 0deded97-91c5-4093-8024-557daf45ee63 · outbound

This paper cites In: International conference on machine learning.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: International conference on machine learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.255686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.255686Z digest=sha256:d3177d7bd0ce7e8813de4f72188cad54d616639d2b6a8648df263a2c5df2110c

Observation 13e07a65-2d5a-469d-bf35-4e99ee468c0e · outbound

This paper cites TULIP: Towards Unified Language-Image Pretraining.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models TULIP: Towards Unified Language-Image Pretraining

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.402513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.402513Z digest=sha256:f20432cab06c734777d5f8780ffbf039f40c6372df02c7624d2243e5b969f370

Observation 26ef6db0-6940-46be-b7d9-cdf0d43f86c6 · outbound

This paper cites Patches Are All You Need?.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Patches Are All You Need?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.527536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.527536Z digest=sha256:cbb0e4935c045c0afb5e4af4c6c742647bc98a21af655e84be8339a30c327978

Observation 3d24a58a-acf8-4e8d-b015-bf246446a3bd · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.632360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.632360Z digest=sha256:fb166043d7df6f9c5b1a0fa0920b115ed156d63e0382cd4244c4e023b1b41d48

Observation 7cefa37e-9d15-48fe-97e3-bc8a10b72dc0 · outbound

This paper cites Advances in Neural Information Pro- cessing Systems36, 46830–46855 (2023).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Advances in Neural Information Pro- cessing Systems36, 46830–46855 (2023)

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.739735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.739735Z digest=sha256:21fec9083751209809142fc55908ee922c04bfd9e136bea56adc1263d974e5bf

Observation 43709f27-d5ee-443b-98f2-53beaa0abf43 · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.864270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.864270Z digest=sha256:fd92c8aec61d0cafefc2aa40b4737af31f81ef8471fda8938a28980ab12137cd

Observation 67ecff82-707c-409e-8f91-d081c975443c · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.978165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.978165Z digest=sha256:0e5bd393e366f119fe2edbf41088825c4fa07506a63ac02f5fd263a73223e473

Observation 4ff79c6b-ed9f-4880-a567-30c7ae0c945c · outbound

This paper cites Advances in neural information pro- cessing systems30(2017).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Advances in neural information pro- cessing systems30(2017)

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.050580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.050580Z digest=sha256:e166ddf8a09a926314068c330de37f6e54fe0dedb0504cf7a01bc32316773fb3

Observation f2cf568c-90f1-4973-80d3-e2ad1412f34a · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.158513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.158513Z digest=sha256:ba7f510a30b5e598bbdf70dd4120ccff274f76f9c25b84f3ad5acc40641d1214

Observation 215009d9-eed3-4b8f-a244-fb2b3fb76d13 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.278782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.278782Z digest=sha256:820a1ed1abaca56c028f2792d0c6f638ab3ce849bad17fc581a0645fbb5b05db

Observation 9e989e71-289b-4c80-acec-be3ad75f102b · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.405105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.405105Z digest=sha256:402c1b552d5c94e2c51f989fea8065859b96a937fe1f2fcaf5503133437347a6

Observation 73ff9c81-59b2-4b4f-a334-417a205cd2b2 · outbound

This paper cites CVPR (2025).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models CVPR (2025)

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.562942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.562942Z digest=sha256:d509dc5feb5f4e8278cc29c2818c93c47fd084e8200530e03af4809f480a3dcb

Observation 2d6969b0-9ef0-468c-9bac-7297c78604e5 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.737392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.737392Z digest=sha256:43d3836ddc7861bdf4963e2cd5a8d29395755ca6677e1e45d53e4b6b0db90a30

Observation 879cc3ac-12cd-49fd-a0d2-a935d2e7f9d5 · outbound

This paper cites In: The Thirteenth International Conference on Learning Representa- tions.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: The Thirteenth International Conference on Learning Representa- tions

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.805471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.805471Z digest=sha256:c1511dac8bf017fedf2a8dc2c3aa13ea9d994498a77731aa810bd6812e21345c

Observation db39c310-8faf-4e46-bb8d-fd63e5af1d47 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.927785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.927785Z digest=sha256:a0b2bfa618c7324362ac8ecc6988c0319380c94b23e2a8fa70e76ecb01211d9f

Observation 982fa184-a23d-49d3-9749-8c5c1ed3f9d4 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.027880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.027880Z digest=sha256:8de720e3b62265ccd1e4eb61515e95bc60b1c96f4f6d1cdace5511d8f52e21fd

Observation 4a7a6ce6-7135-4398-809d-52fc868c2d20 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pat- tern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pat- tern recognition

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.145738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.145738Z digest=sha256:7cdc75224e5239dc441384930335b7b3705283ce3a6a9e838c57b9c47878c333

Observation ca06b46d-6b45-4976-b46d-6285fe82513a · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.208088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.208088Z digest=sha256:45a15b3a9bf8bee2e5530ecc6fcde4fe15fbedc2bca7cabb7b2ceaa909012eb4

Observation 5a712220-6839-4948-a204-d866fc4a2017 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.326347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.326347Z digest=sha256:bdea12e2bdb46f5b4a2edce476edd5c60c3e0c6d588b2ab2a6f8ab1e3b1367e4

Observation 62f12636-3703-4b1f-aa19-017e96d3ab07 · outbound

This paper cites arXiv e-prints pp.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models arXiv e-prints pp

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.434396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.434396Z digest=sha256:65a724257eadb84e3988fc514c0d864ca2784204100091f3d1e56e6dd5e89744

Observation 399b288d-08f2-4733-8220-7de4df0a0dba · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.532976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.532976Z digest=sha256:328c6e295364301e6bad3b14138e150a5f897179501ebd4995f20da4e9e9632a

Observation 3814d9e8-3e61-434a-a671-c35ad689f654 · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.654463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.654463Z digest=sha256:ee4c53464579c5ae0491262d0a255ba1cec17260c1e5ee246e910c1ef1e64ba4

Observation f1d4c1ff-9071-4986-8938-1af28d583a6b · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.738511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.738511Z digest=sha256:ab9be426222a124b0d8ee24d0c678ee347330431ce000d85015fefd45bd226fb

Observation 8fe9eb04-1671-44ed-966d-39493cdc64f6 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 70

Resolution
malformed identifier
no resolver link, observed 2026-07-31T00:01:39.867668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.867668Z digest=sha256:dfdbb63a958d97a2c47db60a88f04abeda80dd4077d3d6a118f89a7b49f9aa0b

Pith citing papers

No inbound Pith citation observations are available.