Pith. sign in

Paper Citation Record · LEDGER

VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2403.00522.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.00522 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:16:46.284408Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T10:59:54.749546Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 13467226-435e-4bf0-a65d-6382134b43b6 · inbound

MpoxVLM: A Vision-Language Model for Diagnosing Skin Lesions from Mpox Virus Infection cites this paper.

MpoxVLM: A Vision-Language Model for Diagnosing Skin Lesions from Mpox Virus Infection VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:16:46.284408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:16:46.284408Z digest=sha256:58b0444cfdcd7319788d2e08162b997405e477302bb2519fda6baad6a8d642ae

Observation 60c69d0f-88ce-4201-a803-ab13abd4585b · inbound

Robust image classification with multi-modal large language models cites this paper.

Robust image classification with multi-modal large language models VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:00:13.052533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:00:13.052533Z digest=sha256:2cfe2aecbe8b3be778c4fe2195522801f70fd2fb0880906878a80eb90a6d5b11

Observation e3895d6a-575d-4893-977a-e15a9efbb9bf · inbound

DiC: Rethinking Conv3x3 Designs in Diffusion Models cites this paper.

DiC: Rethinking Conv3x3 Designs in Diffusion Models VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:52.510789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:52.510789Z digest=sha256:a9d4b2d20d9087eff6cf6838b6951840f99f2b5513838b2ed1f819e35b02a8dd

Observation ad34788c-ce41-4ede-95a2-2043812b615f · inbound

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding cites this paper.

Advancing General Multimodal Capability of Vision-language Models with Pyramid-descent Visual Position Encoding VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T18:53:21.155051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:53:21.155051Z digest=sha256:2fd39864ab7addd9a838cb17b38efbea67f92933d4b9ff4495d7965320bc5668

Observation 2ccfb1b7-20de-47af-933e-d0aeb3e43249 · inbound

UDiTQC: U-Net-Style Diffusion Transformer for Quantum Circuit Synthesis cites this paper.

UDiTQC: U-Net-Style Diffusion Transformer for Quantum Circuit Synthesis VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T15:06:40.512853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:06:40.512853Z digest=sha256:7d4a148c7557725b4fee46b6f354e84ab254c5df0305c147ca4e9e4e99211e43

Observation 7439bea8-daca-4d88-8499-61210ebfd298 · inbound

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures cites this paper.

GRR-CoCa: Leveraging LLM Mechanisms in Multimodal Model Architectures VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T14:42:59.996885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:42:59.996885Z digest=sha256:cac81b3d0a508a7ff38f4712679ad622926efbe0fe6822ee7892ca181f607336

Observation 1b636a14-23e1-4dd3-965e-2092e5b3d788 · inbound

PixNerd: Pixel Neural Field Diffusion cites this paper.

PixNerd: Pixel Neural Field Diffusion VisionLLaMA: A Unified LLaMA Backbone for Vision Tasks

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:59:54.761049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T10:59:53.695426Z digest=sha256:649b1fbdaed120d25508ad3fab3aa4fd9471a213dd41eb94da1b4194822adce2