Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T13:35:05.203103Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 2 inbound Pith citation observations for arXiv:2502.07838.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T13:35:05.203103Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-21T08:39:20.202572Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T08:39:53.356638Z
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 71ab9b5c-603c-4583-ab3d-1199acd0d0ba · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Deep Learning using Rectified Linear Units (ReLU)
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8ff9256-8223-44e1-81fe-b2a325854594 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? L., and Parikh, D
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2f193c3-92e0-4cf1-8802-a265776c2a08 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? sPhinX: Sample Efficient Multilingual Instruction Fine-Tuning Through N-shot Guided Prompting
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63a6ba0c-1583-47dc-a6f8-4357cc718012 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Llama 3: Next-generation open-source language models, 2024
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 0867f3ca-0589-4565-8da8-7dfd6f62c1f0 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Multimodal Deep Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc374f1e-019e-4885-8ded-cd4e36732e6b · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Flamingo: a Visual Language Model for Few-Shot Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41bf2048-cca2-4962-938a-ed8b5fc32f3f · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Layer Normalization
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 762b3e11-3891-419b-8f55-6a9306fbeb64 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Honeybee: Locality-enhanced Projector for Multimodal LLM
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f137f531-b55f-4569-af54-2a5d362987e9 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Microsoft COCO Captions: Data Collection and Evaluation Server
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63b43545-9a26-4b86-9c12-b70c4087ab7c · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? TinyStories: How Small Can Language Models Be and Still Speak Coherent English?
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccd4f1d9-8243-43bb-b9d4-7255dd799adb · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Vision-language pretraining: Bridging vision and language with transformers, 2023
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d225a2d6-c379-4880-92a0-457c957231e0 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Exploring the frontier of vision-language models: A survey of current methodologies and future directions, 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d471db5-6a7a-4d2f-9df2-6cb1182a2437 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? VCoder: Versatile Vision Encoders for Multimodal Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4b542eb-68eb-4b8f-92f2-ed0cde7ced25 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? The illustrated gpt-2, 2019
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 59305fb5-9d11-45e6-aa55-6f55bdb4bcfd · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? BRAVE: Broadening the visual encoding of vision-language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cac67ee3-c85f-4987-94ab-a04fdde58d7d · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c98324ae-7756-4264-8f3b-85c1f443b873 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation fef8db1f-2b17-45b7-9119-613ebd755233 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Forgetful causal masking makes causal language models better few-shot learners
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation a54f062f-5abe-4ce2-9974-3d63de6ba1fc · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Visual Instruction Tuning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ca94c39-6def-411b-bc0b-9dc0efd34585 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? VividMed: Vision Language Model with Versatile Visual Grounding for Medicine
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 045d3b54-74c4-480c-9580-c4bed57dc7ce · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Chatgpt: A language model for conversational ai
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 4dd829f7-8f3f-4d3d-aa6b-f424ae3e2575 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? GPT-4 Technical Report
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 157ff30a-f100-4790-9e5e-f971d487f3b3 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Gpt-4o system card, 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d10d0c34-a4a5-4459-80a1-56fd60c5fa1c · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Kosmos-2: Grounding Multimodal Large Language Models to the World
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75f42b9c-4689-459b-851b-e1dbb39ae201 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? A., Wang, L., Cervantes, C
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 747ca1a0-b68b-4c0d-b0a3-9d1406a98aa4 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Learning Transferable Visual Models From Natural Language Supervision
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b137d05-271b-4c18-89e7-ed5a907054f4 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? An end-to-end trainable neural network for image-based sequence recognition and its application to scene text recognition
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation bd7a9c8b-04b3-46e6-acee-abd39ef0c37d · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Dragonfly: Multi-Resolution Zoom-In Encoding Enhances Vision-Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c28946e-cf4c-4d21-9012-ab3a32c1664d · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Attention Is All You Need
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5929d016-defc-4bbd-8c10-821905a6b605 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Show and tell: A neural image caption generator
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3db7e792-b817-45ae-a83a-6bbb1ebea96c · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? GIT: A Generative Image-to-text Transformer for Vision and Language
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 460e4414-7861-4a94-b4e1-dae41a74b03a · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4cab989-0a6d-4dfa-9ae4-4b98aa9cc0e1 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Efficient Vision-Language Models by Summarizing Visual Tokens into Compact Registers
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76bedbdf-2b86-4d7c-8a64-90f5ac9d5a2d · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ffe2807-6b8b-4513-a546-39fb773b7951 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? StableMask: Refining Causal Masking in Decoder-only Transformer
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c455f70b-8ff5-4f81-934d-8489ed68f59a · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? TinyGPT-V: Efficient Multimodal Large Language Model via Small Backbones
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d3ce9ec-023e-454b-88fb-89d1242cb157 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? A Survey of Large Language Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01582bc1-595c-48a6-8f70-ee75b66b9177 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b26d970d-eb97-4f57-9044-4e2cfbf26af6 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? Mini GPT -4: Enhancing vision-language understanding with advanced large language models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b56ab03-731e-44ae-893b-9516474904e1 · outbound
NanoVLMs: How small can we go and still make coherent Vision Language Models? write newline
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3451ea62-21e4-4270-945a-86045c4c1d51 · inbound
Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models NanoVLMs: How small can we go and still make coherent Vision Language Models?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 3efb266f-0772-4aba-94f0-7737a9b03026 · inbound
Beyond Attention Scores: SVD-Based Vision Token Pruning for Efficient Vision-Language Models NanoVLMs: How small can we go and still make coherent Vision Language Models?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.