Pith. sign in

Paper Citation Record · LEDGER

When Do We Not Need Larger Vision Models?

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2403.13043.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.13043 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:17:22.756008Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-05T11:41:02.805036Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cc85ff8f-9599-464e-8401-c799f953c6ac · inbound

Multimodal Autoregressive Pre-training of Large Vision Encoders cites this paper.

Multimodal Autoregressive Pre-training of Large Vision Encoders When Do We Not Need Larger Vision Models?

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-12T15:17:22.756008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:17:22.756008Z digest=sha256:234ca4b620d79861341cf05930735bc74c66afa98e771856b29dca4d0116273f

Observation 8bdecde8-5bd6-48cd-99c3-30c96ce04cf7 · inbound

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios cites this paper.

RefHCM: A Unified Model for Referring Perceptions in Human-Centric Scenarios When Do We Not Need Larger Vision Models?

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:52.763273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:05:52.763273Z digest=sha256:adbf059649ed32d046fe6774d79c8e16d9690134d17a80961472759d13d5aba6

Observation 46314a16-111e-4ea6-a0c2-69c8fa65bbba · inbound

FEA-SLT: A Gloss-Free End-to-End Framework for Facial-Expression-Aware Sign Language Translation cites this paper.

FEA-SLT: A Gloss-Free End-to-End Framework for Facial-Expression-Aware Sign Language Translation When Do We Not Need Larger Vision Models?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T12:19:05.755187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:19:05.755187Z digest=sha256:2ae5ebb80268a48772d7e09b8b291d9870dffedb5678ac495def7b2fc1544120

Observation 928a73ae-3339-4401-9510-5cdd5afe0bbf · inbound

ClawEnvKit: Automatic Environment Generation for Claw-Like Agents cites this paper.

ClawEnvKit: Automatic Environment Generation for Claw-Like Agents When Do We Not Need Larger Vision Models?

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-05T11:41:02.807457Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-05T11:32:36.356530Z digest=sha256:8a2f6fc8c0cb05c84c093a44c424786dff4b816c86c2736f0393462f82c495b9

Observation 72805b24-9ff7-427e-8f52-31466409cff8 · inbound

Advancing Vision Transformer with Enhanced Spatial Priors cites this paper.

Advancing Vision Transformer with Enhanced Spatial Priors When Do We Not Need Larger Vision Models?

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.635162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T05:22:21.264807Z digest=sha256:1854e81c675cc1c2dcb24d3b10457d40e8565425ec9dd119d90e1f32f9f196ca

Observation 02daeadd-fcdd-49e8-98ef-14e32e2938c4 · inbound

$A^2$: Smaller Self-Supervised ViTs Localize Better than Larger Ones cites this paper.

$A^2$: Smaller Self-Supervised ViTs Localize Better than Larger Ones When Do We Not Need Larger Vision Models?

Reference 31

Resolution
malformed identifier
arxiv_id, observed 2026-07-02T02:06:26.740154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T11:14:47.182429Z digest=sha256:d505d41add4399aea5357147635a14cce1f82020e54a6285ed01b12dfc917ea0

Observation 144f1609-5b75-41cc-9cd7-a031ea417612 · inbound

HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models cites this paper.

HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models When Do We Not Need Larger Vision Models?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T14:48:45.165044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:48:45.165044Z digest=sha256:96bc2c145a2abc104cfb9cc14223e6b35cbd0b576368262f11b91a9f0e3b221b