Pith. sign in

Paper Citation Record · LEDGER

NanoVLMs: How small can we go and still make coherent Vision Language Models?

As of 23 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 2 inbound Pith citation observations for arXiv:2502.07838.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07838 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:35:05.203103Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-21T08:39:20.202572Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T08:39:53.356638Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact2
  • verified fuzzy6
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 71ab9b5c-603c-4583-ab3d-1199acd0d0ba · outbound

This paper cites Deep Learning using Rectified Linear Units (ReLU).

NanoVLMs: How small can we go and still make coherent Vision Language Models? Deep Learning using Rectified Linear Units (ReLU)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.020270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.020270Z digest=sha256:f74e5d90fa975ed7b9cbe48b54891791e07087f8ad1cd8384363dfb9671035d9

Observation b8ff9256-8223-44e1-81fe-b2a325854594 · outbound

This paper cites L., and Parikh, D.

NanoVLMs: How small can we go and still make coherent Vision Language Models? L., and Parikh, D

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.025112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.025112Z digest=sha256:0f14d16412cee5a7681e4299063a7371dc6956b086cf2dd61468ddd96eefe534

Observation b2f193c3-92e0-4cf1-8802-a265776c2a08 · outbound

This paper cites sPhinX: Sample Efficient Multilingual Instruction Fine-Tuning Through N-shot Guided Prompting.

NanoVLMs: How small can we go and still make coherent Vision Language Models? sPhinX: Sample Efficient Multilingual Instruction Fine-Tuning Through N-shot Guided Prompting

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.029759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.029759Z digest=sha256:c7cce005dcd979f1837f57aac404cf1d1d61b9dad3e303c4c35afe216c455b73

Observation 63a6ba0c-1583-47dc-a6f8-4357cc718012 · outbound

This paper cites Llama 3: Next-generation open-source language models, 2024.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Llama 3: Next-generation open-source language models, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:35:06.295163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:35:05.035068Z digest=sha256:86693fe2fbf4d62ffd49b00d3d1774dc729db9efae836c9a5cebfcd0ca2d74c1

Observation 0867f3ca-0589-4565-8da8-7dfd6f62c1f0 · outbound

This paper cites Multimodal Deep Learning.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Multimodal Deep Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.039509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.039509Z digest=sha256:a94404f105cd55c5a704418f9edf80bafc6f179be2197c08787a23a8b8d66bda

Observation cc374f1e-019e-4885-8ded-cd4e36732e6b · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Flamingo: a Visual Language Model for Few-Shot Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.044985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.044985Z digest=sha256:ba3f7230dd0d7873a1eb35b5b0b40f5cc1022c840e7cee559a5aacce685a51ae

Observation 41bf2048-cca2-4962-938a-ed8b5fc32f3f · outbound

This paper cites Layer Normalization.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Layer Normalization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.050657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.050657Z digest=sha256:eda9a9b9b79c054e4f8b573d9eb5e620b262f82c9f6bed9a4824f0afa9293a7a

Observation 762b3e11-3891-419b-8f55-6a9306fbeb64 · outbound

This paper cites Honeybee: Locality-enhanced Projector for Multimodal LLM.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Honeybee: Locality-enhanced Projector for Multimodal LLM

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.055236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.055236Z digest=sha256:f340f44dffeabf7f3e09900ed162899d30d6e9f2f56a07d83a52d773771365f1

Observation f137f531-b55f-4569-af54-2a5d362987e9 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.059823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.059823Z digest=sha256:d83fd995da8b4d341e594b1f19eabe1a0f10c2f388b1f4783b0d4ffe396f2e1d

Observation 63b43545-9a26-4b86-9c12-b70c4087ab7c · outbound

This paper cites TinyStories: How Small Can Language Models Be and Still Speak Coherent English?.

NanoVLMs: How small can we go and still make coherent Vision Language Models? TinyStories: How Small Can Language Models Be and Still Speak Coherent English?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.064299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.064299Z digest=sha256:b786219052fcdc1cdce078794a96a7d5126e543eec1b833899d4deeb0355d80d

Observation ccd4f1d9-8243-43bb-b9d4-7255dd799adb · outbound

This paper cites Vision-language pretraining: Bridging vision and language with transformers, 2023.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Vision-language pretraining: Bridging vision and language with transformers, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:35:06.279804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:35:05.068784Z digest=sha256:8dcfab9f7d9451353d4d9d602ffb267a6362149f742b0af58664258d442808ed

Observation d225a2d6-c379-4880-92a0-457c957231e0 · outbound

This paper cites Exploring the frontier of vision-language models: A survey of current methodologies and future directions, 2024.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Exploring the frontier of vision-language models: A survey of current methodologies and future directions, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.073217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.073217Z digest=sha256:1af62f8dd5099ab0f4b55ea22c0110182bd62ce6ee5c83685ec9149be82621a2

Observation 7d471db5-6a7a-4d2f-9df2-6cb1182a2437 · outbound

This paper cites VCoder: Versatile Vision Encoders for Multimodal Large Language Models.

NanoVLMs: How small can we go and still make coherent Vision Language Models? VCoder: Versatile Vision Encoders for Multimodal Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.078035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.078035Z digest=sha256:b67493db08a75a196f3915bfa937732884f7b9ce3250fa99d435d8b615ca6e16

Observation b4b542eb-68eb-4b8f-92f2-ed0cde7ced25 · outbound

This paper cites The illustrated gpt-2, 2019.

NanoVLMs: How small can we go and still make coherent Vision Language Models? The illustrated gpt-2, 2019

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:35:06.263411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:35:05.082807Z digest=sha256:fc9324553619a43ae3691417db23c2ab4bf667bc328abcf25f6507b60e69f505

Observation 59305fb5-9d11-45e6-aa55-6f55bdb4bcfd · outbound

This paper cites BRAVE: Broadening the visual encoding of vision-language models.

NanoVLMs: How small can we go and still make coherent Vision Language Models? BRAVE: Broadening the visual encoding of vision-language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.087875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.087875Z digest=sha256:c108fd8926864173acf1b6efd4c51479917dd29df2c72e294d7b58448edce6fa

Observation cac67ee3-c85f-4987-94ab-a04fdde58d7d · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

NanoVLMs: How small can we go and still make coherent Vision Language Models? BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.092936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.092936Z digest=sha256:9ffad14cef14701de04e3bad08e492a48c083f5844e0a253650804d51d450839

Observation c98324ae-7756-4264-8f3b-85c1f443b873 · outbound

This paper cites an unresolved cited work.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-08T13:35:06.248245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:35:05.098193Z digest=sha256:16a3af365f665ffb92b4a649283745be3fe90d526bd82a3cc2e73c17dc108901

Observation fef8db1f-2b17-45b7-9119-613ebd755233 · outbound

This paper cites Forgetful causal masking makes causal language models better few-shot learners.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Forgetful causal masking makes causal language models better few-shot learners

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:35:06.233533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:35:05.102625Z digest=sha256:be52b4d513ddb139e92faa3870175fc1951925f0748330455baf873ad51e03e8

Observation a54f062f-5abe-4ce2-9974-3d63de6ba1fc · outbound

This paper cites Visual Instruction Tuning.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Visual Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.106822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.106822Z digest=sha256:d61a45bcebdfb10338599784eea025bc8e9fa69a6a0be9acd1a6bb3fb24be363

Observation 5ca94c39-6def-411b-bc0b-9dc0efd34585 · outbound

This paper cites VividMed: Vision Language Model with Versatile Visual Grounding for Medicine.

NanoVLMs: How small can we go and still make coherent Vision Language Models? VividMed: Vision Language Model with Versatile Visual Grounding for Medicine

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.111491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.111491Z digest=sha256:35e4283de249c9e5fba530bc3ccc8968bc5d13ba1b1228ac32f38fd40b642e35

Observation 045d3b54-74c4-480c-9580-c4bed57dc7ce · outbound

This paper cites Chatgpt: A language model for conversational ai.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Chatgpt: A language model for conversational ai

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:35:06.218857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:35:05.115922Z digest=sha256:ceecda9ebdd3cbba7b689a13fe21602ea80d35bc921a74120fc53869923fa8b4

Observation 4dd829f7-8f3f-4d3d-aa6b-f424ae3e2575 · outbound

This paper cites GPT-4 Technical Report.

NanoVLMs: How small can we go and still make coherent Vision Language Models? GPT-4 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.120292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.120292Z digest=sha256:36b1de842d54b44a27c123157301c78e1ce8b47c9f59b574204e20f581bf0253

Observation 157ff30a-f100-4790-9e5e-f971d487f3b3 · outbound

This paper cites Gpt-4o system card, 2024.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Gpt-4o system card, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T13:35:06.203625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:35:05.124670Z digest=sha256:b35c167d88a002975a8da578a9b6d9674ce6a89b9cf5ad68b8401ed7ed92bfb6

Observation d10d0c34-a4a5-4459-80a1-56fd60c5fa1c · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.128799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.128799Z digest=sha256:b940c863d110a749a186d8d7a4f0b53242f332eb4c0360932cba2044569d9e58

Observation 75f42b9c-4689-459b-851b-e1dbb39ae201 · outbound

This paper cites A., Wang, L., Cervantes, C.

NanoVLMs: How small can we go and still make coherent Vision Language Models? A., Wang, L., Cervantes, C

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.133195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.133195Z digest=sha256:9cd59980fd34f1d7315a2cc482c6c8d9892ea3a563dd7927c8f9cb0905b68a08

Observation 747ca1a0-b68b-4c0d-b0a3-9d1406a98aa4 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Learning Transferable Visual Models From Natural Language Supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.137172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.137172Z digest=sha256:0caf4a81dfa8e899690dc3fabdcde622139f45ef7929717a0f26fd2d2da59d98

Observation 4b137d05-271b-4c18-89e7-ed5a907054f4 · outbound

This paper cites An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition.

NanoVLMs: How small can we go and still make coherent Vision Language Models? An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition

Reference 27

Resolution
verified exact
doi, observed 2026-08-08T13:35:05.246056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:35:05.141271Z digest=sha256:5c66468a5c82a5a92783917780238dc341bfe77da48082a974d8b8079d777256

Observation bd7a9c8b-04b3-46e6-acee-abd39ef0c37d · outbound

This paper cites Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.146089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.146089Z digest=sha256:85e61d5fd3a1198423a731b47877af55088768e78f4a76e042055ffecaa50d22

Observation 2c28946e-cf4c-4d21-9012-ab3a32c1664d · outbound

This paper cites Attention Is All You Need.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Attention Is All You Need

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.150365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.150365Z digest=sha256:c668c10237393009cc824de426f08bdccd3d92ffc1c9f957e835d7418e8eded5

Observation 5929d016-defc-4bbd-8c10-821905a6b605 · outbound

This paper cites Show and tell: A neural image caption generator.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Show and tell: A neural image caption generator

Reference 30

Resolution
verified exact
arxiv_id_nonexistent, observed 2026-08-08T13:35:05.663478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-08T13:35:05.154692Z digest=sha256:e3a671298d312930634070f893f702ab3c4da9700aacfd54522a4a23373e6fe6

Observation 3db7e792-b817-45ae-a83a-6bbb1ebea96c · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

NanoVLMs: How small can we go and still make coherent Vision Language Models? GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.159093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.159093Z digest=sha256:d71ead95c834849ff0555153d9bc1f57d41c45756088d0285fe43518d0370da4

Observation 460e4414-7861-4a94-b4e1-dae41a74b03a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.163450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.163450Z digest=sha256:05fe70e7f1da95f312145f03ea7e62ce023e93465df26ae0c18e2748d6865a98

Observation e4cab989-0a6d-4dfa-9ae4-4b98aa9cc0e1 · outbound

This paper cites Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.167918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.167918Z digest=sha256:e1a710efe12a6b65a7d4176d2c8473ed33e12bd59f37154a8b8a4578c3f83588

Observation 76bedbdf-2b86-4d7c-8a64-90f5ac9d5a2d · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

NanoVLMs: How small can we go and still make coherent Vision Language Models? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.172373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.172373Z digest=sha256:804d727c8177e06f745e2618132410dc2999dff883ab6014ac97011d6e5c3842

Observation 1ffe2807-6b8b-4513-a546-39fb773b7951 · outbound

This paper cites StableMask: Refining Causal Masking in Decoder-only Transformer.

NanoVLMs: How small can we go and still make coherent Vision Language Models? StableMask: Refining Causal Masking in Decoder-only Transformer

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.176959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.176959Z digest=sha256:65825e048d51b6e9281c93ec6369345f4238b6665c756d291dc1d2f24790e923

Observation c455f70b-8ff5-4f81-934d-8489ed68f59a · outbound

This paper cites TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones.

NanoVLMs: How small can we go and still make coherent Vision Language Models? TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.183485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.183485Z digest=sha256:e5e6ac9f3214242605a98e42d6efa3bd7532c6de57dfda3fa86562714117d01f

Observation 7d3ce9ec-023e-454b-88fb-89d1242cb157 · outbound

This paper cites A Survey of Large Language Models.

NanoVLMs: How small can we go and still make coherent Vision Language Models? A Survey of Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.188506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.188506Z digest=sha256:127eb7972538bcc5444c44822a6f65e0282b558bfa1a5057e0696d60a3491c17

Observation 01582bc1-595c-48a6-8f70-ee75b66b9177 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.193085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.193085Z digest=sha256:4d056dac959cab4e2c92d969ff8caa7ff1c94edbb0896675622355fbc19ab3cc

Observation b26d970d-eb97-4f57-9044-4e2cfbf26af6 · outbound

This paper cites Mini GPT -4: Enhancing vision-language understanding with advanced large language models.

NanoVLMs: How small can we go and still make coherent Vision Language Models? Mini GPT -4: Enhancing vision-language understanding with advanced large language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.198092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.198092Z digest=sha256:876f8c7345da2fbf1c4517abe7954d0d3af3394f1da8b16362408a2ff068a098

Observation 2b56ab03-731e-44ae-893b-9516474904e1 · outbound

This paper cites write newline.

NanoVLMs: How small can we go and still make coherent Vision Language Models? write newline

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T13:35:05.203103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T13:35:05.203103Z digest=sha256:41f621f7d5c6ad0f464cefae17779ddb7da3976e26dfd98c9e2395dfefb16265

Pith citing papers

Observation 3451ea62-21e4-4270-945a-86045c4c1d51 · inbound

Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models cites this paper.

Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models NanoVLMs: How small can we go and still make coherent Vision Language Models?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:58.510682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-10T16:27:58.757680Z digest=sha256:d19686ac8c152eb05dc8387b49717a4682a08c39e71f6b1896069842c3c0e198

Observation 3efb266f-0772-4aba-94f0-7737a9b03026 · inbound

Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models cites this paper.

Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models NanoVLMs: How small can we go and still make coherent Vision Language Models?

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:39:53.358666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-21T08:39:20.202572Z digest=sha256:7390c4caa4fde2d0b10e4c753cfbc1c8b47f5b9fd0e3199a642114bbc7de75e4