Pith. sign in

Paper Citation Record · LEDGER

Scaling Vision Transformers to 22 Billion Parameters

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2302.05442.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2302.05442 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:20:38.174869Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

118
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 23729b9b-4725-47c5-9564-2da4060c3202 · inbound

PaLM-E: An Embodied Multimodal Language Model cites this paper.

PaLM-E: An Embodied Multimodal Language Model Scaling Vision Transformers to 22 Billion Parameters

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:29:29.833939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T22:29:29.631351Z digest=sha256:45a8657f50440c473097b8a4d7e7e56b4bd0c5bdf190bed9a2d0c77d6200a2a4

Observation e95e3818-21d5-4680-8c7c-363a6e442c85 · inbound

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action cites this paper.

MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action Scaling Vision Transformers to 22 Billion Parameters

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T01:17:58.737145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T01:17:58.678036Z digest=sha256:3bfdafe513a7b308e7e158392df8c1754df0e31e238c5a20152a0bea12aae520

Observation 2688c09a-5d90-4442-b90c-d2cfa5d9aac2 · inbound

Scaling Data-Constrained Language Models cites this paper.

Scaling Data-Constrained Language Models Scaling Vision Transformers to 22 Billion Parameters

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:35:21.596114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T01:35:21.150772Z digest=sha256:16305ad1bd1e6bda4bcd46583df00e7a9461856c2cbf4b969e3a211192848ae1

Observation 9573a299-a857-4fa5-b700-90716e570d7b · inbound

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts cites this paper.

PiKE: Adaptive Data Mixing for Large-Scale Multi-Task Learning Under Low Gradient Conflicts Scaling Vision Transformers to 22 Billion Parameters

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T16:20:38.174869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:20:38.174869Z digest=sha256:62b9b2f1d923658ae0f4aeafcfc58dfe853507f9fd6903b6c361f17148e2bdf8

Observation 2133d3de-4b97-4476-9336-66851138504e · inbound

Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models cites this paper.

Direct Ascent Synthesis: Revealing Hidden Generative Capabilities in Discriminative Models Scaling Vision Transformers to 22 Billion Parameters

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T11:44:44.160662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:44:44.160662Z digest=sha256:66199d897cf5ab8e3ead6b127825baf034bbe0bd6d31275a8643be67ba426bdc

Observation 7d02eae8-43dc-4400-8256-542870d7463b · inbound

Adversarial Attacks on Robotic Vision Language Action Models cites this paper.

Adversarial Attacks on Robotic Vision Language Action Models Scaling Vision Transformers to 22 Billion Parameters

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:11:40.412772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:11:40.412772Z digest=sha256:f5f1cc673cb06e26da830e08abe8bd56dcbf12ac5676eb7878a96bb172c25719

Observation 1a4ac500-b51a-4f0b-8103-a72bd24ed971 · inbound

Distributed Cross-Channel Hierarchical Aggregation for Foundation Models cites this paper.

Distributed Cross-Channel Hierarchical Aggregation for Foundation Models Scaling Vision Transformers to 22 Billion Parameters

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:29:59.436591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:29:59.436591Z digest=sha256:6713a05908174b224319ee65f1eadcfc67cdac36372ad54862abd21830f8ba1a

Observation d38b0cd0-40aa-4a9d-a794-0a3d11f0e51d · inbound

Scalable Object Detection in the Car Interior With Vision Foundation Models cites this paper.

Scalable Object Detection in the Car Interior With Vision Foundation Models Scaling Vision Transformers to 22 Billion Parameters

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T21:11:50.831055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T21:07:31.589011Z digest=sha256:2de4da309f7daf5db786ca65a9087362fddeec1c2423c005f6ae934fa6f31bda

Observation adddb5ec-da02-467e-b030-836e6a73746c · inbound

LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems cites this paper.

LLMOrbit: A Circular Taxonomy of Large Language Models -From Scaling Walls to Agentic AI Systems Scaling Vision Transformers to 22 Billion Parameters

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T12:47:53.644339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T12:47:28.248540Z digest=sha256:c30b47e770f519a494d6a37cfeca00f39db2ff2064976cffbffc166f24d01a0d

Observation 91153d93-4a5c-4ad7-9ca3-b01c9c58f090 · inbound

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models cites this paper.

CG-MLLM: Captioning and Generating 3D content via Multi-modal Large Language Models Scaling Vision Transformers to 22 Billion Parameters

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-21T14:50:14.890183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T14:48:21.787919Z digest=sha256:55832990075a08842a1b804eee5b78a94c3db248db51e0f97ae31f1b384abcd1

Observation 5aa45d76-8595-4ff7-b135-eef1d3dfef7b · inbound

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use cites this paper.

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use Scaling Vision Transformers to 22 Billion Parameters

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:45:05.576201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T05:44:58.483183Z digest=sha256:b64eddd997a3834ae47af6c78b4b8422267b5cc1577c1afc9e64a50ebd106b7f

Observation 4da0c83c-c463-475b-8c3a-4ccd33e04f18 · inbound

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use cites this paper.

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use Scaling Vision Transformers to 22 Billion Parameters

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:03:46.390770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T21:02:51.889450Z digest=sha256:97353c353608329741f695bfda0e737d22b2593e3f33bc815a1e8c6a5faa3d58

Observation 331735e9-43e3-47b6-941d-60a359b6b60a · inbound

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use cites this paper.

VectraYX-Nano: A 42M-Parameter Spanish Cybersecurity Language Model with Curriculum Learning and Native Tool Use Scaling Vision Transformers to 22 Billion Parameters

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T09:31:22.762383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T09:29:30.950904Z digest=sha256:bd409d30781b117355ddb9961e76795b1dad5f5c62af82f7fd78740035ebb8ca

Observation 638b18c9-252e-4496-99a5-a2168512e572 · inbound

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability cites this paper.

A Readiness-Driven Runtime for Pipeline-Parallel Training under Runtime Variability Scaling Vision Transformers to 22 Billion Parameters

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:38:09.371957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T07:35:32.225708Z digest=sha256:483522082e469c9d3d927d83807d1550648255d8efb99f023534c6fa1e8e128b

Observation 927da818-1663-4882-bffc-704ad517ecdf · inbound

Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor cites this paper.

Most Transformer Modifications Still Do Not Transfer at 1-3B: A 2020-2026 Update to Narang et al. (2021) with Downstream Evaluation and a Noise Floor Scaling Vision Transformers to 22 Billion Parameters

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:19:42.020874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T06:15:47.451870Z digest=sha256:e3ca63ad024517413cf42d31c589ab674bd8a5cfd7a98afe257aa0878feec541

Observation 3e286a78-b870-4567-b118-a7e7fdb8cb3f · inbound

Multimodal Alignment and Preference Optimization for Zero-Shot Conditional RNA Generation cites this paper.

Multimodal Alignment and Preference Optimization for Zero-Shot Conditional RNA Generation Scaling Vision Transformers to 22 Billion Parameters

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:46.871113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T22:17:56.029187Z digest=sha256:0c673c63fcafa1bb01b8350d729a31970e8d5907b21118e5a676ba3e2e2899d8

Observation 5f9f321a-8d8f-4e16-8130-897384d7fcbb · inbound

Unsupervised Semantic Segmentation Facilitates Model Understanding cites this paper.

Unsupervised Semantic Segmentation Facilitates Model Understanding Scaling Vision Transformers to 22 Billion Parameters

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:43:15.473340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T08:37:02.350175Z digest=sha256:10740d2c75cb332c40691073a10bd55b19ddb34d1450421026512ad46750e25e

Observation d3d629c8-08dc-4f13-bf64-2fd1098ce959 · inbound

Unsupervised Semantic Segmentation Facilitates Model Understanding cites this paper.

Unsupervised Semantic Segmentation Facilitates Model Understanding Scaling Vision Transformers to 22 Billion Parameters

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:16.375380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-04T00:34:21.224797Z digest=sha256:d7759c9159a32ac922d8359769a94aad6af6052bf48715a413aa1d812718de8b

Observation 81530ff2-576f-4bfc-a440-8cc5e3d12828 · inbound

Scaling Laws for Neural-Network Quantum States cites this paper.

Scaling Laws for Neural-Network Quantum States Scaling Vision Transformers to 22 Billion Parameters

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:27.071592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T11:30:23.893472Z digest=sha256:fc63fc7bedf9eb7e85f5cbc21e6af1bd4471e0931728ca3caaf78756424cc57c

Observation 3a7d2d8c-ea92-4925-b702-774037e804a7 · inbound

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders cites this paper.

Size Doesn't Matter: Cosine-Scored Sparse Autoencoders Scaling Vision Transformers to 22 Billion Parameters

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:15:29.697712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T07:15:16.674714Z digest=sha256:bf5a097111ba4935069d457414c988e0c17f544c53fa686eb84882e67ed7afe0

Observation 694d21cb-4f0c-46e9-aa13-177c648bdd48 · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Scaling Vision Transformers to 22 Billion Parameters

Reference 160

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:30:07.807239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-25T20:05:09.179627Z digest=sha256:2fd227c62a0132e720a13d0d4b0ee6bbc60f57d0ad1aac000cfb67b952297ba6

Observation 3244b7a8-17a5-4b57-bf6b-2bc2a95d06ae · inbound

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors cites this paper.

Improving Neural Network Training by Decoupling the Magnitude and Direction of Weight Vectors Scaling Vision Transformers to 22 Billion Parameters

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T10:14:05.666991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:14:05.666991Z digest=sha256:98a99e91549bbfe4dd795a52124a66b635dd54ff7ab196f9fb8250111350579f

Observation fc1682b5-aa53-4aed-9792-b9fe167a72bb · inbound

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget cites this paper.

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Scaling Vision Transformers to 22 Billion Parameters

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T06:13:51.944737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:13:51.944737Z digest=sha256:86a8f78c77e76917c2684eef929da88ffebe8a673e4e602feaf7766a158f8e11

Observation 6d2c0ac0-df7d-4e71-92ba-38b944d8b4a7 · inbound

Predict before you train: Scaling Laws for particle physics foundation models cites this paper.

Predict before you train: Scaling Laws for particle physics foundation models Scaling Vision Transformers to 22 Billion Parameters

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-30T23:56:38.595543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T23:56:38.595543Z digest=sha256:4d8544871f0ba4a7088907d7a68a4cbb6431ae735231b15be11b688c14cfe268

Observation a09046c1-03fd-451d-91a8-f87589e196ff · inbound

Opt.Gear Technical Report cites this paper.

Opt.Gear Technical Report Scaling Vision Transformers to 22 Billion Parameters

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T00:40:38.653520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:40:38.653520Z digest=sha256:f6cc42415834e60d5cb99907690352bad9e5ea1bb398140db17f91d455c24444

Observation 698ecde2-46b1-494a-b9a0-6406481878e0 · inbound

One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse cites this paper.

One QK Channel, Many Sources: Guarding Low-Precision Attention Collapse Scaling Vision Transformers to 22 Billion Parameters

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T15:20:15.447616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T15:20:15.447616Z digest=sha256:3c8fa608378163a5ec2c92fe8169ced805bc3fb13bb73e6b448598f381cdbaff