Pith. sign in

Paper Citation Record · LEDGER

Images are Worth Variable Length of Representations

As of 19 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 2 inbound Pith citation observations for arXiv:2506.03643.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03643 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:05:22.416317Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T07:09:25.049534Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:06:44.335867Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved40
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5fa297b5-09a8-47db-b861-d400370d9dc8 · outbound

This paper cites GPT-4 Technical Report.

Images are Worth Variable Length of Representations GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.235943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.235943Z digest=sha256:34f6c53e8bfe4fd21c4f751ffd41cc94993cb2298119c2a024d6f9c2e9ee61e6

Observation 45912b42-4cdf-44e2-9917-854b7273e822 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023.

Images are Worth Variable Length of Representations Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.239610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.239610Z digest=sha256:70e7ef44f19289ffd9d658a10e46f8cbad2c1e59c738696da4200a647fb415c7

Observation d03f0524-d703-4c29-84d8-f74e3ed3d7c6 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022.

Images are Worth Variable Length of Representations Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.242682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.242682Z digest=sha256:0ea110ba6700d54ac76334529aef7717adabe939b1a6bce6776aa9b4efdfc6b6

Observation a2f8e363-b706-42b5-8f02-f5da44c53879 · outbound

This paper cites Revisiting active perception.Autonomous Robots, 42:177–196, 2018.

Images are Worth Variable Length of Representations Revisiting active perception.Autonomous Robots, 42:177–196, 2018

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.934151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.245814Z digest=sha256:47aafee4cb50556391388e4e51380a4f8b614107b13477f6a1b66a54890c5547

Observation fd3a7952-9f8a-4c5d-a704-96bcde14e1db · outbound

This paper cites Blur image detection using laplacian operator and open-cv.

Images are Worth Variable Length of Representations Blur image detection using laplacian operator and open-cv

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.924943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.248818Z digest=sha256:ee29f734d38caf33fefa36dcd3956a3278d1fd6121ff806d4a7cc9e17bd8eea9

Observation f132ad6e-5e22-4f43-a69a-77c5872919e6 · outbound

This paper cites Statistical inference for probabilistic functions of finite state markov chains.The annals of mathematical statistics, 37(6):1554–1563, 1966.

Images are Worth Variable Length of Representations Statistical inference for probabilistic functions of finite state markov chains.The annals of mathematical statistics, 37(6):1554–1563, 1966

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.914977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.251721Z digest=sha256:3efad992eda56667045cb29700f05d73962f6d5fc685d4f68d09806f46d7a10b

Observation 891de3fe-d2f9-477f-8142-191130b5ae5e · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Images are Worth Variable Length of Representations Pythia: A suite for analyzing large language models across training and scaling

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.904132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.255187Z digest=sha256:56ea4d627d8ddbfa60efbeff57b8ede09146074bf4c43dd65a4f4a9f7bb8e071

Observation 01c49392-716c-4588-a8b7-6ca5ececbaaf · outbound

This paper cites Token merging: Your vit but faster.

Images are Worth Variable Length of Representations Token merging: Your vit but faster

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.257886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.257886Z digest=sha256:2ca0fd6f7ce37314dc071270223e5734a494db94f88a96fb341133e8f761558e

Observation ad4bbdee-535d-4a74-9173-da359f52b388 · outbound

This paper cites Food-101 – mining discriminative components with random forests.

Images are Worth Variable Length of Representations Food-101 – mining discriminative components with random forests

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.260665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.260665Z digest=sha256:5a9420f575e06df1332a8e88d36fb467d9277690cad81af04c64412bfbee2710

Observation 2172ddcb-58ec-4afc-ac8b-c87d2e38167b · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Images are Worth Variable Length of Representations Emerging properties in self-supervised vision transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.263449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.263449Z digest=sha256:9807ebd2fe0854409b0d487e07c5da64956c1bb2e7902332e76b8e07bfff2a43

Observation e44a12e4-71b8-4b62-b426-1cd75b084c70 · outbound

This paper cites Efficient large multi-modal models via visual context compression.

Images are Worth Variable Length of Representations Efficient large multi-modal models via visual context compression

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.873209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.266244Z digest=sha256:dedf0ce9fa67aed592ba2f17901ec669dcde8dbd0c9e7dd15142c09ef176d088

Observation 01fa2fea-701b-42d8-a567-8bd7311bcea3 · outbound

This paper cites Review of image classification algorithms based on convolutional neural networks.Remote Sensing, 13(22):4712, 2021.

Images are Worth Variable Length of Representations Review of image classification algorithms based on convolutional neural networks.Remote Sensing, 13(22):4712, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.862922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.269137Z digest=sha256:5eb999019129dbd07089ca77c8cf76e15955aa74406b1db133623b367e91b3cb

Observation ca84c8ff-6e6b-486e-b2f8-78b8f3a4b496 · outbound

This paper cites An empirical study of smoothing techniques for language modeling.

Images are Worth Variable Length of Representations An empirical study of smoothing techniques for language modeling

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.852630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.271750Z digest=sha256:3d012305af1f6bab4de938b4fda197a2748b7917be92d059b428e019846d154b

Observation 8d482209-8cb5-4eab-a1d2-0e307aee191e · outbound

This paper cites Cimpoi, S.

Images are Worth Variable Length of Representations Cimpoi, S

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.274496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.274496Z digest=sha256:41196115a263d583e7a7acb6c680e1f33ec5ec0600ec4b035eadeb3743455c20

Observation 6e7ac09b-2df3-4c41-b5d0-2c3d7ccca8a1 · outbound

This paper cites An analysis of single layer networks in unsupervised feature learning aistats.

Images are Worth Variable Length of Representations An analysis of single layer networks in unsupervised feature learning aistats

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.836894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.277525Z digest=sha256:92ca9c85040f7835e909f7f0a705db9c72c91dafbfd5d68959ec62acfc81a37f

Observation c9be1d94-02f7-47eb-b6b9-6470a20c06c5 · outbound

This paper cites Scaling up dataset distillation to imagenet-1k with constant memory.

Images are Worth Variable Length of Representations Scaling up dataset distillation to imagenet-1k with constant memory

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.280475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.280475Z digest=sha256:86852c11592138eb7e2472150f6970403df26ccfaef6f070163e36c2c0dcc82b

Observation 9e6443a4-945f-40b8-b154-11e647a0239c · outbound

This paper cites Top-down control of eye movements: Yarbus revisited.Visual Cognition, 17(6-7):790–811, 2009.

Images are Worth Variable Length of Representations Top-down control of eye movements: Yarbus revisited.Visual Cognition, 17(6-7):790–811, 2009

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.820288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.283436Z digest=sha256:d499e7045b46b4fe501eb868677d02b846db650d57b5630c48770680892ebb80

Observation c76268b8-cb52-4d1e-a89b-0b55abb2c568 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Images are Worth Variable Length of Representations Imagenet: A large-scale hierarchical image database

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.809368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.286379Z digest=sha256:e761bd990e9a5443ac0f7bcc89a32012afea323ef5d21817c3404ffdea553411

Observation ff348b87-899f-4bd0-9660-cbc8c5a06b90 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Images are Worth Variable Length of Representations An image is worth 16x16 words: Transformers for image recognition at scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.289289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.289289Z digest=sha256:72f7008e9ca5496afb79255129cbb45ec106c8e0a1a743c4ccae00567fffa202

Observation 92471137-7e22-4321-8b0b-a8cba96c9d0f · outbound

This paper cites Adaptive length image tok- enization via recurrent allocation.

Images are Worth Variable Length of Representations Adaptive length image tok- enization via recurrent allocation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.791905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.292319Z digest=sha256:d6a821eb3c7990ca1aeb46d2411e1992db0e2c5876e54856a2043beff23f68d4

Observation 6b25f19e-82c0-4e3e-90c4-99087b80d5c5 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Images are Worth Variable Length of Representations Taming transformers for high-resolution image synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.295367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.295367Z digest=sha256:11bbc691e79812028766d005481f4a2efea833bdc1f3e284dbac41a9df8e8682

Observation 95c113f2-cfef-4850-8774-b2cb5ebb92f9 · outbound

This paper cites Multimodal autoregressive pre-training of large vision encoders, 2024.

Images are Worth Variable Length of Representations Multimodal autoregressive pre-training of large vision encoders, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.775879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.298398Z digest=sha256:18f6bd6bb52b636be0e6f7854227ac9a98299f9b7f2224718c90d0e551354852

Observation d7e007a9-dfd4-4daf-8305-60905aef446f · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering, 2017.

Images are Worth Variable Length of Representations Making the v in vqa matter: Elevating the role of image understanding in visual question answering, 2017

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.301186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.301186Z digest=sha256:db752f7543da5e17c3fcac5847c9847e6b722aa518c36584bbe989be15916594

Observation c6bd0c5f-092f-44b4-b79f-b4dcf9d75158 · outbound

This paper cites The Llama 3 Herd of Models.

Images are Worth Variable Length of Representations The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.304085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.304085Z digest=sha256:19d2f1623801dda50cf0ff7d1602c4f2111a01dc4e0ed627371fbd9f8f6088b1

Observation fe918958-fb65-46b5-ba89-513bb74ace0d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Images are Worth Variable Length of Representations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.307061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.307061Z digest=sha256:2e6c2c78742ae382a57fb97fe518dafd6edfad74899c160f77eafb887a7bf7e6

Observation 43299d14-b086-4b06-b686-53aa23a0cc53 · outbound

This paper cites A review of semantic segmentation using deep neural networks.International journal of multimedia information retrieval, 7:87–93, 2018.

Images are Worth Variable Length of Representations A review of semantic segmentation using deep neural networks.International journal of multimedia information retrieval, 7:87–93, 2018

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.758858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.310089Z digest=sha256:087c5614176ca6a1a6c16b543c8e0b64a307746a1eef25e9100f63f84ff1c178

Observation 019de7a3-6c95-41e8-9e30-3ea473633ec8 · outbound

This paper cites A brief survey on semantic segmentation with deep learning.

Images are Worth Variable Length of Representations A brief survey on semantic segmentation with deep learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.312920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.312920Z digest=sha256:44cb97ebe9677b7b931ae2bc4113364a42b62a7cb336618d7bd705853467ab41

Observation 48f972e5-dcc0-4816-9149-9f89f62bbbe1 · outbound

This paper cites Hierarchical cross-modal agent for robotics vision-and-language navigation.

Images are Worth Variable Length of Representations Hierarchical cross-modal agent for robotics vision-and-language navigation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.741335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.316327Z digest=sha256:419c6c8147a520e0e45d4f304618e05cda32a7141256e2e5ca64941ec9e1951a

Observation 6bece6e8-d26d-4f67-b9c6-82763fa8dad1 · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

Images are Worth Variable Length of Representations Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.319209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.319209Z digest=sha256:e8401bbba50a262807de4569c36dfa9aa8bfe16506dacc8e4c8f3298061332b3

Observation f3001747-ebb7-4a85-b3d1-db8d77ccf7f0 · outbound

This paper cites Perceiver: General perception with iterative attention.

Images are Worth Variable Length of Representations Perceiver: General perception with iterative attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.322257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.322257Z digest=sha256:f58dd8e72cb8919724b2b02df65caaf632e405bb609938f81ed273864dc63778

Observation 9b001137-dc8a-4ab0-a6e0-3bccaf5394aa · outbound

This paper cites Auto-encoding variational bayes, 2013.

Images are Worth Variable Length of Representations Auto-encoding variational bayes, 2013

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.324996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.324996Z digest=sha256:34570891b8adf545aac18875e7bc69fbac0815e7f3201931279771169892ae34

Observation 36a19078-aea1-4d25-ba63-b83344cfb3e3 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017.

Images are Worth Variable Length of Representations Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.327905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.327905Z digest=sha256:7f6ca99cff50134c04748db6fbe497920f9d777a0f25a3116bd0ab2d92360abe

Observation 41ee6020-27b0-4d1c-8767-75c13e86566b · outbound

This paper cites Learning multiple layers of features from tiny images.

Images are Worth Variable Length of Representations Learning multiple layers of features from tiny images

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.330945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.330945Z digest=sha256:81720a2d91ed48830379bdde56ccb24812888bdcc3ddbf65b4bd11be8181cbc3

Observation 77620b32-7059-4b66-a838-71e014c302c0 · outbound

This paper cites an unresolved cited work.

Images are Worth Variable Length of Representations Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.334047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.334047Z digest=sha256:166d1566e1a8672299ca4541f016d2ef40481d53a9c76796b8e0423f7b978720

Observation d9380372-0304-4957-9e22-d96bf19a583d · outbound

This paper cites The roles of vision and eye movements in the control of activities of daily living.Perception, 28(11):1311–1328, 1999.

Images are Worth Variable Length of Representations The roles of vision and eye movements in the control of activities of daily living.Perception, 28(11):1311–1328, 1999

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.698087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.336862Z digest=sha256:9168ffa684ccbf6f217f536c957e0bc2ad172c266255d029d662008f16a405ef

Observation 42c60635-8e1e-4e20-972e-3a8ba85d4ffe · outbound

This paper cites Microsoft coco: Common objects in context.

Images are Worth Variable Length of Representations Microsoft coco: Common objects in context

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.340117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.340117Z digest=sha256:e9b2e50cac023d949949ae89ef3bdcd6116ebc796635032d193c065e68831a2e

Observation f52a53eb-89ab-4977-b03b-1fc26cb8d4c5 · outbound

This paper cites Visual instruction tuning, 2023.

Images are Worth Variable Length of Representations Visual instruction tuning, 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.681102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.342828Z digest=sha256:01d2e8cb09f63d7eef31654b136702657e5af3f384c065e945d598e476509eef

Observation 96561ff4-5ad4-48de-8efe-babd7df27e49 · outbound

This paper cites A survey of image classification methods and techniques for improving classification performance.International journal of Remote sensing, 28(5):823–870, 2007.

Images are Worth Variable Length of Representations A survey of image classification methods and techniques for improving classification performance.International journal of Remote sensing, 28(5):823–870, 2007

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.670826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.345651Z digest=sha256:a0638746208e68896761a84883c24b758a0492b0972420441eaaa2b6cf8fb7d7

Observation 4e3f4811-8715-4b05-a7b0-9bbde605c2c1 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022.

Images are Worth Variable Length of Representations Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.348474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.348474Z digest=sha256:cb8bc895b4cc6439ea10501a932946befc893140b19c80844889fdc54688503b

Observation 5692391d-aa2f-4c6c-a73f-199b11dd5a1a · outbound

This paper cites Fine-grained visual classification of aircraft, 2013.

Images are Worth Variable Length of Representations Fine-grained visual classification of aircraft, 2013

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.351532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.351532Z digest=sha256:236db2256f80c7fefbd42f0da3c792442c1035532b7ce41d03486c5938a835e9

Observation 00df95de-f42a-4d2f-9812-6f9301f61c5f · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019.

Images are Worth Variable Length of Representations Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.354914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.354914Z digest=sha256:02a5afc82d21a7d4d91295291afc363fa523cc4c97dd35d04d1a8958255d26d3

Observation 7da9feef-d2de-41b9-a215-da20f4d4b611 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning, 2022.

Images are Worth Variable Length of Representations Chartqa: A benchmark for question answering about charts with visual and logical reasoning, 2022

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.357727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.357727Z digest=sha256:80acd1a847c35f882cbeec566a374af3bb4fd4a2daa4f918ac7e59c08b3c108a

Observation 7b8df15e-85fb-43e6-bb6a-f7c9738b71e3 · outbound

This paper cites V Jawahar.

Images are Worth Variable Length of Representations V Jawahar

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.360520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.360520Z digest=sha256:437663ad5532b3696a05f5297dbbaa50eaacbbd28bda48d716b3760fa15ee67c

Observation d92654c2-b6b8-4e24-9f1a-081cd8ab8de1 · outbound

This paper cites an unresolved cited work.

Images are Worth Variable Length of Representations Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.363555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.363555Z digest=sha256:05761965e3d886b981f074c81b8ab88ba2076bf167200674659aea455f695762

Observation aaeaaffa-f164-4f98-8d82-5f7d2d7c1063 · outbound

This paper cites Stl-10, nov 2024.

Images are Worth Variable Length of Representations Stl-10, nov 2024

Reference 45

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T11:05:22.619520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.366118Z digest=sha256:15560e77fd44fe500045c3f40feb4a3c32483a382b0ddcb317e22099e72219c7

Observation 2ae2af82-135c-4123-8a5c-2f48793d3f72 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Images are Worth Variable Length of Representations Learning transferable visual models from natural language supervision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.368954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.368954Z digest=sha256:77e588b9359879e3553f86a30e91c2e09d9fa4c35e3fab565a56599b8fcac516

Observation 82b92334-15b0-45d2-94c7-f836d7634716 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Images are Worth Variable Length of Representations Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.371771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.371771Z digest=sha256:a5ca7fe90c773e3eb8e755dccd32881a95053f06e10dfe7817295cb9d5fef866

Observation 10ca220c-212c-444f-8b32-e6b1ef1576ec · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.Advances in neural information processing systems, 34:13937–13949, 2021.

Images are Worth Variable Length of Representations Dynamicvit: Efficient vision transformers with dynamic token sparsification.Advances in neural information processing systems, 34:13937–13949, 2021

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.374870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.374870Z digest=sha256:b85baa75e20758e342e3ecf6229df632ff6f5d4bc15c38e1b97c731b9e22905b

Observation 3be09f3f-052c-41df-aa96-d8b96c5b0430 · outbound

This paper cites Generating diverse high-fidelity images with vq-vae-2.

Images are Worth Variable Length of Representations Generating diverse high-fidelity images with vq-vae-2

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.377480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.377480Z digest=sha256:55537b23a8c213b3c8723e3fba45166a7b7e3bcedc3a0431e4c9ce576e293dc8

Observation 88a6ba60-4198-4c95-bf35-0f7e44931193 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Images are Worth Variable Length of Representations High-resolution image synthesis with latent diffusion models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.381721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.381721Z digest=sha256:2316b24d5b0c37499fd1fa9a20fea0c54a1eb49bcc5e19d35bf7c5b80c1ae3b5

Observation fe8720a3-f6eb-49ed-81c7-91b2fcbd33f4 · outbound

This paper cites Towards vqa models that can read, 2019.

Images are Worth Variable Length of Representations Towards vqa models that can read, 2019

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.384499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.384499Z digest=sha256:bd82f307c3598c6db8002d22623fb072e1bd7a52f1a1f5b47709ef810c879b1a

Observation 05203ea4-08d8-4c28-9b26-bfc328771026 · outbound

This paper cites Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning.Artificial intelligence, 112(1-2):181–211, 1999.

Images are Worth Variable Length of Representations Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning.Artificial intelligence, 112(1-2):181–211, 1999

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.387502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.387502Z digest=sha256:c9a029c19856c0f502ea24fd3f061b1cfe72bd71cfad1d73a2840525394aa50b

Observation 8866789d-7efd-46b3-91d9-5bbeb1ed0fcb · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Images are Worth Variable Length of Representations Gemini: A Family of Highly Capable Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.390224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.390224Z digest=sha256:38fb1dbea1f77c37122af919adb785d606155d3e7376d276d673a120bb8b7123

Observation 6e42155a-daaf-4746-9f3f-11801312b986 · outbound

This paper cites Neural discrete representation learning.Advances in neural information processing systems, 30, 2017.

Images are Worth Variable Length of Representations Neural discrete representation learning.Advances in neural information processing systems, 30, 2017

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.393356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.393356Z digest=sha256:3117cd2e259626c91c638c3dd1561168eb7e31ecc4419d287ac688845459ffb8

Observation d13295d0-1535-4bae-b745-3635609c0cdd · outbound

This paper cites Attention is all you need.

Images are Worth Variable Length of Representations Attention is all you need

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.396119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.396119Z digest=sha256:0e0189e9a683bb59942774c5219da5fcc7d6aae2be8aa8127f394a383b305846

Observation c6c02abb-3fb2-480d-8926-60765e2584e4 · outbound

This paper cites Supervised hashing for image retrieval via image representation learning.

Images are Worth Variable Length of Representations Supervised hashing for image retrieval via image representation learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.558474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.398997Z digest=sha256:a846f295396b07711dd5fea05c96e96d5af4a9ddbf3f16cce1bfe7cbac1f4b08

Observation 584a92dd-b964-41ec-8b3b-20017a04943d · outbound

This paper cites Ehinger, Aude Oliva, and Antonio Torralba.

Images are Worth Variable Length of Representations Ehinger, Aude Oliva, and Antonio Torralba

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.548329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.401621Z digest=sha256:d45424fbba30653a5e3cf994b3a6e80ab32c3ad4e9f2f0fe87c048687955bfc0

Observation bc20ad02-831e-4c8b-b824-d8e030d24534 · outbound

This paper cites A-vit: Adaptive tokens for efficient vision transformer.

Images are Worth Variable Length of Representations A-vit: Adaptive tokens for efficient vision transformer

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.404732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.404732Z digest=sha256:7951c7f6645ff886d5da1d7f5434fa76c0bc8c4136acd4606e1eabaf54600048

Observation b7f27687-3010-43f3-bdbe-4f0cc2ef0e23 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

Images are Worth Variable Length of Representations Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.407403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.407403Z digest=sha256:669c27878ccde129b30a42acab38824d186745949b6fa2233a39a7a50aeb22b0

Observation 2f1ff9be-878a-4111-aedb-0634abbdcec8 · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.Advances in Neural Information Processing Systems, 37:128940–128966, 2024.

Images are Worth Variable Length of Representations An image is worth 32 tokens for reconstruction and generation.Advances in Neural Information Processing Systems, 37:128940–128966, 2024

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.410713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.410713Z digest=sha256:25009ad53f011b431d02f8899943d2d41c9066a51b4f7491fd2e3fdf4ae1a3e8

Observation 88729760-bbb3-4dce-b971-1b0cccb86c0b · outbound

This paper cites Object detection with deep learning: A review.IEEE transactions on neural networks and learning systems, 30(11):3212–3232, 2019.

Images are Worth Variable Length of Representations Object detection with deep learning: A review.IEEE transactions on neural networks and learning systems, 30(11):3212–3232, 2019

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.525814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.413512Z digest=sha256:17b6fe95f9bd4d07b0fd176d5e9c24d06775b185da9e0a02658a3c47cfc225e0

Observation 73c305e8-6a6d-4ba0-b741-f6e23dfe44ca · outbound

This paper cites STOP” on a sign as “SHOP.

Images are Worth Variable Length of Representations STOP” on a sign as “SHOP

Reference 62

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:05:22.514936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T11:05:22.416317Z digest=sha256:dce27061812b10e283ae58c5d5f49ab636356e71144d7fdef2b2b6f5f3d8ea95

Pith citing papers

Observation acf9eb43-b6b6-45a5-8c17-8e97430f7680 · inbound

DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking cites this paper.

DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking Images are Worth Variable Length of Representations

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:10:05.814629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T15:10:00.146694Z digest=sha256:912e88ef484bd8141dfadf47f0d7b69fcebc8c20d819bae007899adb8891fa55

Observation 2b64c07e-428e-4fbb-9902-0b96192c6b33 · inbound

ChannelTok: Efficient Flexible-Length Vision Tokenization cites this paper.

ChannelTok: Efficient Flexible-Length Vision Tokenization Images are Worth Variable Length of Representations

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:06:44.338154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-28T07:09:25.049534Z digest=sha256:5592736a16c0c43da03526b237b1f22071b62851fbf7a5e115690ad1be85f758