Pith. sign in

Paper Citation Record · LEDGER

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers

As of 23 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 1 inbound Pith citation observation for arXiv:2501.09221.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09221 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:12:58.172261Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:35:03.349259Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T17:35:03.450625Z

Reference resolution

13 of 13 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1966fa80-2ed0-491e-949d-13af8a6fcf86 · outbound

This paper cites Describing objects by their attributes.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers Describing objects by their attributes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:12:58.355315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:12:58.140847Z digest=sha256:6619c7cb9ff26b2f162cf26e4865e3851dba6ac3990048195742da183e8c20c4

Observation 2fc6978c-abd0-4474-91ba-1071b70ea112 · outbound

This paper cites AST: Audio Spectrogram Transformer.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers AST: Audio Spectrogram Transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:58.144966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:58.144966Z digest=sha256:bbc7feaac4a390e5f1aa4ce9486cbd076ebe0f2f1189a1a87e507df126641112

Observation 2af4640b-0220-4526-a915-86b3be921b4e · outbound

This paper cites Are Convolutional Neural Networks or Transformers more like human vision?.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers Are Convolutional Neural Networks or Transformers more like human vision?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:58.153980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:58.153980Z digest=sha256:c3cf15b1f122cbea3b0532dbb35c1a6717ea18beb2d0370c38dc3bb178798fc8

Observation 1e30bcfb-b02e-4203-815c-e7bb59b0959e · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:58.163307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:58.163307Z digest=sha256:c39cbb2a9dedc83705ba6be1769afd828e03674a47620b945cb8984ad2eb718d

Observation 1e9718b5-4a7d-4876-a0d8-8e6e79d6de7c · outbound

This paper cites We set the patch size to correspond to 16x16 pixels.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers We set the patch size to correspond to 16x16 pixels

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:12:58.328782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:12:58.168030Z digest=sha256:c8bff8ce43b4e9a1d09e8f63d0791f6fd555dc64c6dd99b61e20669cf21a4ad8

Observation 5aedb210-1e7d-4602-9b5f-bfde6765f8ba · outbound

This paper cites The following convolution blocks consist of a single convolution layer of kernel size=3 and stride=2, as well as the batch norm and ReLU activations.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers The following convolution blocks consist of a single convolution layer of kernel size=3 and stride=2, as well as the batch norm and ReLU activations

Reference 1029

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:12:58.316300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:12:58.172261Z digest=sha256:5c983a553dbde887341e7472ad7be7678c408eaaca6ab5695b7387af4b233ad4

Observation 55785ad0-265a-4d90-a10f-99f7d97a28a6 · outbound

This paper cites The caltech-ucsd birds-200-2011 dataset.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers The caltech-ucsd birds-200-2011 dataset

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:12:58.341517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:12:58.158392Z digest=sha256:c57580f4bead98bc124990901266877660a80610548b3ed629136c84d2e91595

Observation febaf998-ae44-47df-9776-593b65b844ff · outbound

This paper cites Image transformer for explainable au- tonomous driving system.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers Image transformer for explainable au- tonomous driving system

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:12:58.370876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:12:58.136626Z digest=sha256:b7a148ca5f805990d6b943f6a9b312e8ddc3658052f904fc15d1131e97cf5d4e

Observation 654f5a22-82c6-4486-a61c-88f8aa928c89 · outbound

This paper cites Vision Transformer Adapter for Dense Predictions.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers Vision Transformer Adapter for Dense Predictions

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:58.127534Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:58.127534Z digest=sha256:19cf7e2fac4303a2b1ca64bca5f789699d8fb88309ac809bbf990e4f93143dae

Observation a742862d-fc28-4129-aaf0-71e73f46c555 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:58.131823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:58.131823Z digest=sha256:8a55118f94c3e80313c6170442f486cbc53971d972ff32c150d6dc88073fef41

Observation d54e9a35-8411-4432-82d3-9ea3956193b6 · outbound

This paper cites On the Opportunities and Risks of Foundation Models.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers On the Opportunities and Risks of Foundation Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:58.123392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:58.123392Z digest=sha256:8e5672ec0cf7ce1d54f4652b97acf49573bdaf5c74c53efb2c62112544281e5b

Observation ec637453-89cc-459b-88f8-9e10cb04b1c7 · outbound

This paper cites SiT: Self-supervised vIsion Transformer.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers SiT: Self-supervised vIsion Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T20:12:58.118649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:12:58.118649Z digest=sha256:2b36ae4b390290457e254907ad067363f088793ddb3b820233a5936c7072662a

Observation df43ab0e-c0d5-4cb3-aa0a-487c8a98c86d · outbound

This paper cites A Self-explaining Neural Architecture for Generalizable Concept Learning.

ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers A Self-explaining Neural Architecture for Generalizable Concept Learning

Reference 2024

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:12:58.236527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-10T20:12:58.149323Z digest=sha256:386c3a12f36832456fbe38fbbd1b3374de11b48266de2ee79ae3577dff74e1f3

Pith citing papers

Observation d069482b-01f8-4bcf-8cb7-08316db20626 · inbound

Integrating attention into explanation frameworks for language and vision transformers cites this paper.

Integrating attention into explanation frameworks for language and vision transformers ASCENT-ViT: Attention-based Scale-aware Concept Learning Framework for Enhanced Alignment in Vision Transformers

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-15T17:35:03.459622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T17:35:03.349259Z digest=sha256:9a56bd2c90098848ab6c399dc88203414037f524adb1989b6ec33727d057d1af