Pith. sign in

Paper Citation Record · LEDGER

Focal Self-attention for Local-Global Interactions in Vision Transformers

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2107.00641.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2107.00641 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T00:50:40.329762Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T18:23:07.249939Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c368a540-f6e6-4f13-b914-96d9a1d803bd · inbound

Florence: A New Foundation Model for Computer Vision cites this paper.

Florence: A New Foundation Model for Computer Vision Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:38:09.578306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T09:38:09.427509Z digest=sha256:dbff18b6cea8981b7a4b38e8e91eae25f9beef1775c83b7d050382475b8a3d74

Observation 71b7a824-4f2a-49f5-8c72-9721d3c7da9d · inbound

Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model cites this paper.

Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space Model Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:37:00.393190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T21:37:00.014069Z digest=sha256:0ccf3f79634e256a451bcff56d24187236463d71518e2f28c815f515f3ee2bb3

Observation 66fa8b1f-e2ff-4f41-894f-bb4f2f155834 · inbound

VMamba: Visual State Space Model cites this paper.

VMamba: Visual State Space Model Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-16T18:23:07.252179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T18:23:07.194824Z digest=sha256:fe8fcdc597feb077ab17104ce63e9ec6969b01380cc3f5af0cd663efc91945f6

Observation 3e793ade-1188-46ac-a119-e0fe0d08f68f · inbound

A Retrospective Systematic Study on Hierarchical Sparse Query Transformer-assisted Ultrasound Screening for Early Hepatocellular Carcinoma cites this paper.

A Retrospective Systematic Study on Hierarchical Sparse Query Transformer-assisted Ultrasound Screening for Early Hepatocellular Carcinoma Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T00:50:40.329762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T00:50:40.329762Z digest=sha256:4598711fe95eb8572fc43da883ea400533a812891702866e892181ff532f491e

Observation 640153bc-2fda-4b1d-8bf3-eca95f6ffafe · inbound

AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer cites this paper.

AnchorFormer: Differentiable Anchor Attention for Efficient Vision Transformer Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:05:20.916843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:05:20.916843Z digest=sha256:9e36c9f697b2ef060d63eaaf872e81b55c76e1cf403f212829cf3124c36d4fe5

Observation 8c4a6ed8-c6a8-4050-b989-c8a690177d4d · inbound

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads cites this paper.

ViT-Split: Unleashing the Power of Vision Foundation Models via Efficient Splitting Heads Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T11:09:31.821493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:09:31.821493Z digest=sha256:2da8a24a9bbe2465d41251c97cc0c8774cdacd673c5106a32ad307256611f675

Observation 11e92a37-aad1-48f9-a5f7-6c9e5cae0ffe · inbound

Resolving Token-Space Gradient Conflicts: Token Space Manipulation for Transformer-Based Multi-Task Learning cites this paper.

Resolving Token-Space Gradient Conflicts: Token Space Manipulation for Transformer-Based Multi-Task Learning Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T18:48:38.030618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:48:38.030618Z digest=sha256:8d2cebec18240a9299d5d1eb6c8a6fed63b10370fa67ec4389f2c2a64a8930d4

Observation 50ffbd01-c296-4531-bcc3-489afd507a6c · inbound

RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization cites this paper.

RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T21:04:17.197463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:04:17.197463Z digest=sha256:037ec69250d7006717f9e2dad724bdb82ee2021d5072bd75891e3320dba36faa

Observation 00199253-45fa-44de-aa34-916163508221 · inbound

Quantum-Enhanced Optimization by Warm Starts cites this paper.

Quantum-Enhanced Optimization by Warm Starts Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T17:26:02.327575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:26:02.327575Z digest=sha256:f2adad6c80e0e4665b62313b6d591e490378ae2c96d80f7adedfbacd50a1bc96

Observation afac8872-93a0-4c4c-8016-5053ab299a07 · inbound

Vision encoders should be image size agnostic and task driven cites this paper.

Vision encoders should be image size agnostic and task driven Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.454271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.454271Z digest=sha256:abc5ee58458681fd469f095e7744ef0087a99c874214739e3d112ad4fca5e576

Observation f2ce4b90-96a2-4e02-9506-611859152f2f · inbound

Partial Ring Scan: Revisiting Scan Order in Vision State Space Models cites this paper.

Partial Ring Scan: Revisiting Scan Order in Vision State Space Models Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T04:45:35.882372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:45:35.882372Z digest=sha256:21153de793411c8191674a35c22936a23fbc0b1f190f2541dea987cc3e1a02a2

Observation 54796c13-d5db-4434-8a3e-6ddbd04fde70 · inbound

MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model cites this paper.

MPDiT: Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:48:19.272275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T23:46:25.197344Z digest=sha256:1d27550959c16b22c47f2296265f746e06212afe46094ef8478b94fe93c8da19

Observation dbc77d95-fc45-4899-b5a0-db6a8f28e082 · inbound

Can Graphs Help Vision SSMs See Better? cites this paper.

Can Graphs Help Vision SSMs See Better? Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:17:22.837670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:17:19.534840Z digest=sha256:f847ba4798b3aa4b02a324fe8a98b2aa17e66ad7408d7f5a7f9c5f6fd5ecaec4

Observation 7739f1df-aa06-447c-8f3d-c850ad63b4d3 · inbound

MMA-Former: Multi-Window Mixture-of-Head Attention Transformer for Adaptive PNI Prediction in 3D MRI cites this paper.

MMA-Former: Multi-Window Mixture-of-Head Attention Transformer for Adaptive PNI Prediction in 3D MRI Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T07:50:26.363180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T07:50:26.363180Z digest=sha256:1c7e52749c515b7a23c1bcc80122780958492de7cc9144bc6ff0385bd73feac4