Pith. sign in

Paper Citation Record · LEDGER

Exploring Plain Vision Transformer Backbones for Object Detection

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2203.16527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2203.16527 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:16:12.363276Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T14:31:40.521481Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6fd6883d-85e4-451d-8bb0-c0bd008ac1b1 · inbound

Adding Conditional Control to Text-to-Image Diffusion Models cites this paper.

Adding Conditional Control to Text-to-Image Diffusion Models Exploring Plain Vision Transformer Backbones for Object Detection

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T22:43:10.956509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T22:43:10.880338Z digest=sha256:75662f97df94da6d86d3da2364bf54ea8f21cfaa6fd2e08eeb0918c2e795e291

Observation 6e598cf3-8252-4fa6-ba2f-fedc9c962aa8 · inbound

Revisiting Feature Prediction for Learning Visual Representations from Video cites this paper.

Revisiting Feature Prediction for Learning Visual Representations from Video Exploring Plain Vision Transformer Backbones for Object Detection

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T12:40:23.949043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T12:40:23.709098Z digest=sha256:6e75b994042727ecc997548fe69c11d45ea8f05a2dbeae608ffeaa6c0cdca26a

Observation def6d8b0-b2e9-49fd-9d5e-d0e419a94406 · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Exploring Plain Vision Transformer Backbones for Object Detection

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.112794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:7c4c7f1686e3d72feb1eedd4f4a21665498d848b81f256ec5659a674355aa803

Observation 597da9c4-4597-419d-8f5f-9c49c7be8822 · inbound

gen2seg: Generative Models Enable Generalizable Instance Segmentation cites this paper.

gen2seg: Generative Models Enable Generalizable Instance Segmentation Exploring Plain Vision Transformer Backbones for Object Detection

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-22T14:31:40.524150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T14:31:30.651144Z digest=sha256:c6164167302ee7fe35ae7d36753c9f6f15492ec48bf2fa9c8a6a8670d0451bd2

Observation b3524013-2008-4470-84d9-193b96830b34 · inbound

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces cites this paper.

Visual Embodied Brain: Let Multimodal Large Language Models See, Think, and Control in Spaces Exploring Plain Vision Transformer Backbones for Object Detection

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:16:12.363276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:16:12.363276Z digest=sha256:1f5777dc96cfac3a2fc47b1d701056c436f9c0aa3d66033fd828e8f8a2b2b045

Observation 8ab5abb8-4ca3-4a0d-ab6e-e39a49dfe3fa · inbound

Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings cites this paper.

Pan-Arctic Permafrost Landform and Human-built Infrastructure Feature Detection with Vision Transformers and Location Embeddings Exploring Plain Vision Transformer Backbones for Object Detection

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:17:54.608439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:17:54.608439Z digest=sha256:d3eec5f106a1bf1b2b4913ab67d971f99dad968021b4e7290009bf133986f599

Observation 9d5ec591-45c1-41cb-b6d8-244f741b30bb · inbound

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs cites this paper.

Manager: Aggregating Insights from Unimodal Experts in Two-Tower VLMs and MLLMs Exploring Plain Vision Transformer Backbones for Object Detection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:38.678581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:38.678581Z digest=sha256:39b99bb37acf34643e12a6fa6429cf59e1302e75cad34d9bc5c460dcdbe07689

Observation 22e4caec-79fe-4af0-9f69-fe861da24206 · inbound

SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images cites this paper.

SMARTIES: Spectrum-Aware Multi-Sensor Auto-Encoder for Remote Sensing Images Exploring Plain Vision Transformer Backbones for Object Detection

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:25.296451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:25.296451Z digest=sha256:48df622e81c09cccf692736abcbddd03ab511137f2f217cc49a0c57a7ca3d289

Observation d8ed63db-1e74-4e59-9763-5dfe625835d7 · inbound

ISRS-DETR: Detection-Guided Click Propagation for Remote Sensing Interactive Segmentation cites this paper.

ISRS-DETR: Detection-Guided Click Propagation for Remote Sensing Interactive Segmentation Exploring Plain Vision Transformer Backbones for Object Detection

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T06:51:04.256800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T06:51:04.256800Z digest=sha256:44d4a48f74adf2835bdd152efeed3fda7905993e7917c6d1560cf388864e9cf7

Observation 4d5017c6-564e-456b-ab00-699f500f4462 · inbound

FUSEP: A Multi-Center Benchmark for Diverse Tasks in Early Pregnancy Fetal Ultrasound Screening cites this paper.

FUSEP: A Multi-Center Benchmark for Diverse Tasks in Early Pregnancy Fetal Ultrasound Screening Exploring Plain Vision Transformer Backbones for Object Detection

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T17:14:45.185998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:14:45.185998Z digest=sha256:2a96c203f46f85785f7c2dc7e4bf4a530dc1091d23407a4e9eb230f1db010f20