Pith. sign in

Paper Citation Record · LEDGER

Scalable Pre-training of Large Autoregressive Image Models

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2401.08541.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.08541 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T01:06:18.216606Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cc51c86e-8447-4528-87ac-473afe76094a · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Scalable Pre-training of Large Autoregressive Image Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-16T04:09:36.398124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:9a6e84809934c1a8bed1c5d6a08a92e831d26ef3227193a1e3c29c246649a6b7

Observation 25897824-a30e-4eba-b7bb-6b61635c7873 · inbound

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics cites this paper.

SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics Scalable Pre-training of Large Autoregressive Image Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:22:37.351022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T21:22:36.902119Z digest=sha256:0cc1dd24e85db029f5c0d3763a06eae5215bfd7d6c640a3813453e9c0c59b15a

Observation cce1f160-3ff1-43ee-b35f-ccb84a9f3b8b · inbound

Visual Pre-Training on Unlabeled Images using Reinforcement Learning cites this paper.

Visual Pre-Training on Unlabeled Images using Reinforcement Learning Scalable Pre-training of Large Autoregressive Image Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T01:06:18.216606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T01:06:18.216606Z digest=sha256:2d59118ed68720976e1155a3f5f8eb1112e80d4592916864d480911f22c5802d

Observation 4d8ec47f-687d-4fea-ab49-a7e1a0a11765 · inbound

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement cites this paper.

SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement Scalable Pre-training of Large Autoregressive Image Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:03.825159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:03.825159Z digest=sha256:55dac2e81fda7379fd07e82fb3644c40d698c3cc803ef9c49b805b78bb3cc6bc

Observation a87c79cf-73ab-4575-9598-d4760d41a977 · inbound

Hita: Holistic Tokenizer for Autoregressive Image Generation cites this paper.

Hita: Holistic Tokenizer for Autoregressive Image Generation Scalable Pre-training of Large Autoregressive Image Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T20:38:53.354253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:38:53.354253Z digest=sha256:049be0232710364eb169ca52de1ebca3d63e871e6f8945218cced61fbcd8a8ef

Observation 505638ea-f1f5-4a45-97f1-1463a38f024b · inbound

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning cites this paper.

OpenVision 2: A Family of Generative Pretrained Visual Encoders for Multimodal Learning Scalable Pre-training of Large Autoregressive Image Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:22:32.780054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:22:32.780054Z digest=sha256:ca3e2c186480478ab537bf5525727b689145b517ff5bf110b3e120586145238a

Observation 028815b8-227f-4d79-ac1f-cbd50fab44f5 · inbound

DifFoundMAD: Foundation Models meet Differential Morphing Attack Detection cites this paper.

DifFoundMAD: Foundation Models meet Differential Morphing Attack Detection Scalable Pre-training of Large Autoregressive Image Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:36:02.098890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:31:36.180645Z digest=sha256:5e1106eb8ed421a575b3ab4807f1fd70a03fc016853022f74ed51b14f7cb43cf

Observation 520346e1-1991-4a4b-a34e-0ef4c86f31bb · inbound

DifFoundMAD: Foundation Models meet Differential Morphing Attack Detection cites this paper.

DifFoundMAD: Foundation Models meet Differential Morphing Attack Detection Scalable Pre-training of Large Autoregressive Image Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T15:57:14.411668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:57:14.411668Z digest=sha256:a9d9bc514355b8c706255e83ece4fdca730d40b2ca7d67e85fffbbdec5ea80dc

Observation 1d1d70fe-05f9-4f77-8685-52af7b180094 · inbound

What Cohort INRs Encode and Where to Freeze Them cites this paper.

What Cohort INRs Encode and Where to Freeze Them Scalable Pre-training of Large Autoregressive Image Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:56:26.363547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-12T01:32:38.878151Z digest=sha256:5ee2b05a531fcbda162ff6c3a6c18f189ad99690ccd64a77d369fd672208ebda

Observation bbe2c6da-c266-4685-8d4d-3287e7ca10df · inbound

Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice cites this paper.

Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice Scalable Pre-training of Large Autoregressive Image Models

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:43:51.072841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-20T22:41:44.510546Z digest=sha256:f6af787e713e7a7a38afc36477061a547c2dd161421d8118e12ecfeb63422920

Observation 3f8d4011-4e1d-4d6f-bcc8-43b69459d3a4 · inbound

Weighted Reverse Convolution for Feature Upsampling cites this paper.

Weighted Reverse Convolution for Feature Upsampling Scalable Pre-training of Large Autoregressive Image Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:28:21.683562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T14:24:28.728963Z digest=sha256:d7301473e59066be8092f021682b74a3739baf4753b786be98a9d9ab214ffb5b

Observation b111038b-75b4-4db8-82f5-9f4b1e6b2a22 · inbound

Weighted Reverse Convolution for Feature Upsampling cites this paper.

Weighted Reverse Convolution for Feature Upsampling Scalable Pre-training of Large Autoregressive Image Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.817998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T08:15:52.610554Z digest=sha256:10a53feae2dab988b5ded1b6e76c98c9cb77dcee186033dfc6a59ab31f69aae6

Observation 58b74b70-2776-4b13-9694-c0a6fc467f3a · inbound

Uncovering the Latent Potential of Deep Intermediate Representations cites this paper.

Uncovering the Latent Potential of Deep Intermediate Representations Scalable Pre-training of Large Autoregressive Image Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:36:39.257486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-25T05:36:24.743558Z digest=sha256:d2a160774c04e0f41df857b6dc6559a1b5035c93eb0743c22ba9cd5580e896b7

Observation 2545e4e5-890a-4a15-862c-1485f8a63296 · inbound

A World Model of Radiologist Reading for Medical Image Representation Learning cites this paper.

A World Model of Radiologist Reading for Medical Image Representation Learning Scalable Pre-training of Large Autoregressive Image Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:15:01.230775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T18:55:50.523140Z digest=sha256:436c7473a8408c516aaaafa188f3e21079d5dfcff3796a36a83bf97471c1a6fb

Observation 2d749fa7-129a-4666-9008-4868700605f6 · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Scalable Pre-training of Large Autoregressive Image Models

Reference 244

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.519117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:4ab8772aa3a9fc438f817662644b90131fbd8f221e8e54c9d3c825352489d917

Observation a07b8580-0dfb-4a53-8ea3-122d735ad973 · inbound

The Galaxy's Guide to the Tokenizer: A Benchmark for Scientific Foundation Models cites this paper.

The Galaxy's Guide to the Tokenizer: A Benchmark for Scientific Foundation Models Scalable Pre-training of Large Autoregressive Image Models

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-25T20:38:19.798073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T20:30:45.814418Z digest=sha256:b1c787d768ea235a29023ae655111ce4f74a6683d77de4b9dc1ece3371d8113e

Observation d38ac3e9-eebc-4c52-b111-4f18ac8a2700 · inbound

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP cites this paper.

ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP Scalable Pre-training of Large Autoregressive Image Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:50.719757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T05:03:15.044146Z digest=sha256:a1f50d8b79151ff37724950d3da93a67c0f21ac82872f5abcfe12276420688ff

Observation 6bf5396e-4475-40ac-b767-1f148fe582af · inbound

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types cites this paper.

Adaptive Loss Balancing for Multi-Task Bioacoustic Classification of Bird Species and Call Types Scalable Pre-training of Large Autoregressive Image Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T03:24:13.814557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:24:13.814557Z digest=sha256:09eb16cf862b28b9b30ed5f44abba46308c4b4a61cec29971461cd8f5f2027ab

Observation 9585a693-900a-4e5b-b603-895a16a09f84 · inbound

Foundation Models for Astrophysics cites this paper.

Foundation Models for Astrophysics Scalable Pre-training of Large Autoregressive Image Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T04:31:48.707484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:31:48.707484Z digest=sha256:953b584d6911c701d0ffcbf750ce43264819a2642b8ab9d20b6e61eedcfffc5b