Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:05:22.416317Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 2 inbound Pith citation observations for arXiv:2506.03643.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:05:22.416317Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-06-28T07:09:25.049534Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T07:06:44.335867Z
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5fa297b5-09a8-47db-b861-d400370d9dc8 · outbound
Images are Worth Variable Length of Representations GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45912b42-4cdf-44e2-9917-854b7273e822 · outbound
Images are Worth Variable Length of Representations Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d03f0524-d703-4c29-84d8-f74e3ed3d7c6 · outbound
Images are Worth Variable Length of Representations Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2f8e363-b706-42b5-8f02-f5da44c53879 · outbound
Images are Worth Variable Length of Representations Revisiting active perception.Autonomous Robots, 42:177–196, 2018
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fd3a7952-9f8a-4c5d-a704-96bcde14e1db · outbound
Images are Worth Variable Length of Representations Blur image detection using laplacian operator and open-cv
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f132ad6e-5e22-4f43-a69a-77c5872919e6 · outbound
Images are Worth Variable Length of Representations Statistical inference for probabilistic functions of finite state markov chains.The annals of mathematical statistics, 37(6):1554–1563, 1966
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 891de3fe-d2f9-477f-8142-191130b5ae5e · outbound
Images are Worth Variable Length of Representations Pythia: A suite for analyzing large language models across training and scaling
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 01c49392-716c-4588-a8b7-6ca5ececbaaf · outbound
Images are Worth Variable Length of Representations Token merging: Your vit but faster
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad4bbdee-535d-4a74-9173-da359f52b388 · outbound
Images are Worth Variable Length of Representations Food-101 – mining discriminative components with random forests
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2172ddcb-58ec-4afc-ac8b-c87d2e38167b · outbound
Images are Worth Variable Length of Representations Emerging properties in self-supervised vision transformers
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e44a12e4-71b8-4b62-b426-1cd75b084c70 · outbound
Images are Worth Variable Length of Representations Efficient large multi-modal models via visual context compression
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 01fa2fea-701b-42d8-a567-8bd7311bcea3 · outbound
Images are Worth Variable Length of Representations Review of image classification algorithms based on convolutional neural networks.Remote Sensing, 13(22):4712, 2021
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ca84c8ff-6e6b-486e-b2f8-78b8f3a4b496 · outbound
Images are Worth Variable Length of Representations An empirical study of smoothing techniques for language modeling
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8d482209-8cb5-4eab-a1d2-0e307aee191e · outbound
Images are Worth Variable Length of Representations Cimpoi, S
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e7ac09b-2df3-4c41-b5d0-2c3d7ccca8a1 · outbound
Images are Worth Variable Length of Representations An analysis of single layer networks in unsupervised feature learning aistats
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c9be1d94-02f7-47eb-b6b9-6470a20c06c5 · outbound
Images are Worth Variable Length of Representations Scaling up dataset distillation to imagenet-1k with constant memory
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e6443a4-945f-40b8-b154-11e647a0239c · outbound
Images are Worth Variable Length of Representations Top-down control of eye movements: Yarbus revisited.Visual Cognition, 17(6-7):790–811, 2009
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c76268b8-cb52-4d1e-a89b-0b55abb2c568 · outbound
Images are Worth Variable Length of Representations Imagenet: A large-scale hierarchical image database
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ff348b87-899f-4bd0-9660-cbc8c5a06b90 · outbound
Images are Worth Variable Length of Representations An image is worth 16x16 words: Transformers for image recognition at scale
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92471137-7e22-4321-8b0b-a8cba96c9d0f · outbound
Images are Worth Variable Length of Representations Adaptive length image tok- enization via recurrent allocation
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6b25f19e-82c0-4e3e-90c4-99087b80d5c5 · outbound
Images are Worth Variable Length of Representations Taming transformers for high-resolution image synthesis
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 95c113f2-cfef-4850-8774-b2cb5ebb92f9 · outbound
Images are Worth Variable Length of Representations Multimodal autoregressive pre-training of large vision encoders, 2024
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d7e007a9-dfd4-4daf-8305-60905aef446f · outbound
Images are Worth Variable Length of Representations Making the v in vqa matter: Elevating the role of image understanding in visual question answering, 2017
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6bd0c5f-092f-44b4-b79f-b4dcf9d75158 · outbound
Images are Worth Variable Length of Representations The Llama 3 Herd of Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe918958-fb65-46b5-ba89-513bb74ace0d · outbound
Images are Worth Variable Length of Representations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43299d14-b086-4b06-b686-53aa23a0cc53 · outbound
Images are Worth Variable Length of Representations A review of semantic segmentation using deep neural networks.International journal of multimedia information retrieval, 7:87–93, 2018
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 019de7a3-6c95-41e8-9e30-3ea473633ec8 · outbound
Images are Worth Variable Length of Representations A brief survey on semantic segmentation with deep learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48f972e5-dcc0-4816-9149-9f89f62bbbe1 · outbound
Images are Worth Variable Length of Representations Hierarchical cross-modal agent for robotics vision-and-language navigation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6bece6e8-d26d-4f67-b9c6-82763fa8dad1 · outbound
Images are Worth Variable Length of Representations Perceiver IO: A General Architecture for Structured Inputs & Outputs
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3001747-ebb7-4a85-b3d1-db8d77ccf7f0 · outbound
Images are Worth Variable Length of Representations Perceiver: General perception with iterative attention
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b001137-dc8a-4ab0-a6e0-3bccaf5394aa · outbound
Images are Worth Variable Length of Representations Auto-encoding variational bayes, 2013
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36a19078-aea1-4d25-ba63-b83344cfb3e3 · outbound
Images are Worth Variable Length of Representations Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41ee6020-27b0-4d1c-8767-75c13e86566b · outbound
Images are Worth Variable Length of Representations Learning multiple layers of features from tiny images
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77620b32-7059-4b66-a838-71e014c302c0 · outbound
Images are Worth Variable Length of Representations Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9380372-0304-4957-9e22-d96bf19a583d · outbound
Images are Worth Variable Length of Representations The roles of vision and eye movements in the control of activities of daily living.Perception, 28(11):1311–1328, 1999
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 42c60635-8e1e-4e20-972e-3a8ba85d4ffe · outbound
Images are Worth Variable Length of Representations Microsoft coco: Common objects in context
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f52a53eb-89ab-4977-b03b-1fc26cb8d4c5 · outbound
Images are Worth Variable Length of Representations Visual instruction tuning, 2023
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 96561ff4-5ad4-48de-8efe-babd7df27e49 · outbound
Images are Worth Variable Length of Representations A survey of image classification methods and techniques for improving classification performance.International journal of Remote sensing, 28(5):823–870, 2007
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4e3f4811-8715-4b05-a7b0-9bbde605c2c1 · outbound
Images are Worth Variable Length of Representations Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5692391d-aa2f-4c6c-a73f-199b11dd5a1a · outbound
Images are Worth Variable Length of Representations Fine-grained visual classification of aircraft, 2013
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00df95de-f42a-4d2f-9812-6f9301f61c5f · outbound
Images are Worth Variable Length of Representations Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7da9feef-d2de-41b9-a215-da20f4d4b611 · outbound
Images are Worth Variable Length of Representations Chartqa: A benchmark for question answering about charts with visual and logical reasoning, 2022
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b8df15e-85fb-43e6-bb6a-f7c9738b71e3 · outbound
Images are Worth Variable Length of Representations V Jawahar
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d92654c2-b6b8-4e24-9f1a-081cd8ab8de1 · outbound
Images are Worth Variable Length of Representations Unresolved cited work
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aaeaaffa-f164-4f98-8d82-5f7d2d7c1063 · outbound
Images are Worth Variable Length of Representations Stl-10, nov 2024
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2ae2af82-135c-4123-8a5c-2f48793d3f72 · outbound
Images are Worth Variable Length of Representations Learning transferable visual models from natural language supervision
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82b92334-15b0-45d2-94c7-f836d7634716 · outbound
Images are Worth Variable Length of Representations Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10ca220c-212c-444f-8b32-e6b1ef1576ec · outbound
Images are Worth Variable Length of Representations Dynamicvit: Efficient vision transformers with dynamic token sparsification.Advances in neural information processing systems, 34:13937–13949, 2021
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be09f3f-052c-41df-aa96-d8b96c5b0430 · outbound
Images are Worth Variable Length of Representations Generating diverse high-fidelity images with vq-vae-2
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88a6ba60-4198-4c95-bf35-0f7e44931193 · outbound
Images are Worth Variable Length of Representations High-resolution image synthesis with latent diffusion models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe8720a3-f6eb-49ed-81c7-91b2fcbd33f4 · outbound
Images are Worth Variable Length of Representations Towards vqa models that can read, 2019
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05203ea4-08d8-4c28-9b26-bfc328771026 · outbound
Images are Worth Variable Length of Representations Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning.Artificial intelligence, 112(1-2):181–211, 1999
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8866789d-7efd-46b3-91d9-5bbeb1ed0fcb · outbound
Images are Worth Variable Length of Representations Gemini: A Family of Highly Capable Multimodal Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e42155a-daaf-4746-9f3f-11801312b986 · outbound
Images are Worth Variable Length of Representations Neural discrete representation learning.Advances in neural information processing systems, 30, 2017
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d13295d0-1535-4bae-b745-3635609c0cdd · outbound
Images are Worth Variable Length of Representations Attention is all you need
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6c02abb-3fb2-480d-8926-60765e2584e4 · outbound
Images are Worth Variable Length of Representations Supervised hashing for image retrieval via image representation learning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 584a92dd-b964-41ec-8b3b-20017a04943d · outbound
Images are Worth Variable Length of Representations Ehinger, Aude Oliva, and Antonio Torralba
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bc20ad02-831e-4c8b-b824-d8e030d24534 · outbound
Images are Worth Variable Length of Representations A-vit: Adaptive tokens for efficient vision transformer
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7f27687-3010-43f3-bdbe-4f0cc2ef0e23 · outbound
Images are Worth Variable Length of Representations Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f1ff9be-878a-4111-aedb-0634abbdcec8 · outbound
Images are Worth Variable Length of Representations An image is worth 32 tokens for reconstruction and generation.Advances in Neural Information Processing Systems, 37:128940–128966, 2024
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88729760-bbb3-4dce-b971-1b0cccb86c0b · outbound
Images are Worth Variable Length of Representations Object detection with deep learning: A review.IEEE transactions on neural networks and learning systems, 30(11):3212–3232, 2019
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 73c305e8-6a6d-4ba0-b741-f6e23dfe44ca · outbound
Images are Worth Variable Length of Representations STOP” on a sign as “SHOP
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation acf9eb43-b6b6-45a5-8c17-8e97430f7680 · inbound
DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking Images are Worth Variable Length of Representations
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2b64c07e-428e-4fbb-9902-0b96192c6b33 · inbound
ChannelTok: Efficient Flexible-Length Vision Tokenization Images are Worth Variable Length of Representations
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.