Pith. sign in

Paper Citation Record · LEDGER

Unified Local and Global Attention Interaction Modeling for Vision Transformers

As of 12 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2412.18778.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18778 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:34:02.333517Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved48
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dfca2a14-6f44-4723-92be-7475724edc22 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.622632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.075673Z digest=sha256:31358867258ccb296b90fc1ee29cad60ae186d66a0049f5f257c5760257bf1f3

Observation d5f60d4d-aebd-49e7-a3c1-2f604f23040e · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.080829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.080829Z digest=sha256:f78cbdf1e2883ee2c8bfc3ccfeeef298faafe02ce7953c08115dc27f66dc2467

Observation 3979f508-145e-4e9d-bb31-df8be4b80109 · outbound

This paper cites End-to-End Object Detection with Transformers.

Unified Local and Global Attention Interaction Modeling for Vision Transformers End-to-End Object Detection with Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.085201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.085201Z digest=sha256:39abf180c1072a804a48a4e2f84976af1a591eb31d73484f35978169b847d1f0

Observation 0a088262-0ecd-4887-b30e-553124069d4f · outbound

This paper cites SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More.

Unified Local and Global Attention Interaction Modeling for Vision Transformers SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.090460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.090460Z digest=sha256:1317eee573c1e8d9870ad91bc736d2e536806e4139304784e54b6bf2acc0ac6f

Observation 48f98527-4342-4056-9503-cfcd32e7f617 · outbound

This paper cites SAM Fails to Segment Anything? -- SAM-Adapter: Adapting SAM in Underperformed Scenes: Camouflage, Shadow, Medical Image Segmentation, and More.

Unified Local and Global Attention Interaction Modeling for Vision Transformers SAM Fails to Segment Anything? -- SAM-Adapter: Adapting SAM in Underperformed Scenes: Camouflage, Shadow, Medical Image Segmentation, and More

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.095509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.095509Z digest=sha256:6e8711f52dd70e5870a7663b97e6714402d2207317168b245bdd7a986894e8e8

Observation d920fe8d-ec60-44e5-a3db-d792211ba781 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-11T04:34:02.100667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.100667Z digest=sha256:44e5aa3b3ba868f20b99d96b9a3fc70fd9857007bcd408bbe46877b1304e4936

Observation 068a3ac1-b34c-4d50-8f9f-8bc81fd6bf04 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Unified Local and Global Attention Interaction Modeling for Vision Transformers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.105878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.105878Z digest=sha256:e2ab47da2d81e6e0b51aba9a461240afa5b407a1e58c7803cc7c30c60c070708

Observation 68fb8423-98e0-41b3-813a-0818bcd50c8d · outbound

This paper cites Concealed Object Detection.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Concealed Object Detection

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:03.274848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.110714Z digest=sha256:c63b14822ce1f111bf2649e1e5dc5dd4e20ac2eb75c868976f4bc03541478418

Observation 1a0bf6ae-f741-46b5-8b39-7b5e7eeff726 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.115665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.115665Z digest=sha256:739bf6a7ed9adf87746e6b3bd6c5e6f9ce6f8f02080bc91a888cbb3a626d382d

Observation 73b23bca-a75f-4879-b3ec-1d13cfe03512 · outbound

This paper cites Fast R-CNN.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Fast R-CNN

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.120229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.120229Z digest=sha256:098d0a8d33b68e65438c5efa8e53de19311d3949a8daa67dac2935b61472a572

Observation f47e55f7-bccf-4bcf-9040-b5bcb3d395d2 · outbound

This paper cites Rich feature hierarchies for accurate object detection and semantic segmentation.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Rich feature hierarchies for accurate object detection and semantic segmentation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.124878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.124878Z digest=sha256:98c06f0e0d47f75e59652d35e52ae5b44eeb720d9dfc45b12b221090bb8c6054

Observation 105fa00f-3743-44de-af76-3b0dce7e19c5 · outbound

This paper cites Goldberger, Luis A.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Goldberger, Luis A

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.129567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.129567Z digest=sha256:631e762d8a983d109d4edcc2ac508b154abb34d0ae0343cb5d185c9ae9a1551a

Observation 7b199c26-bcd3-467d-bd0b-566ff6fd9edb · outbound

This paper cites CMT: Convolutional Neural Networks Meet Vision Transformers.

Unified Local and Global Attention Interaction Modeling for Vision Transformers CMT: Convolutional Neural Networks Meet Vision Transformers

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:03.149792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.133793Z digest=sha256:d5580f3b6c6867321c158f58b9a817d65c19fd192f14c76402bcd57ef103f9cc

Observation 4d20afe7-52e0-44dd-9c4f-b818083f4852 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.608646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.138382Z digest=sha256:902fbe349950cc0527e5f58ef5ed8b1f1c90fb9d7998d7398b6d626bceaf27d5

Observation 3f033212-3b81-420b-b429-cefa7a0520a1 · outbound

This paper cites Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.146803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.146803Z digest=sha256:8764d2536382675a723cca19572e42056014f89128d71a7e9c70527670793fe7

Observation 4e8854c7-6b64-4e60-a66a-f5fe6cbfe725 · outbound

This paper cites Deep Residual Learning for Image Recognition.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Deep Residual Learning for Image Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.151323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.151323Z digest=sha256:c5262c9db460b0ac1b738adabc4e091fa3ad4077bd24c7c2ff366e8f588b638c

Observation 18bac6bf-cb98-427e-b1bd-b8855302fea2 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.595134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.156743Z digest=sha256:b655c1c4b6f78766fccff02eb9aa1236967426db37ffe95ec8e6009df2ea0349

Observation 5e0380c1-9d67-4cdf-8679-c1428be32ecf · outbound

This paper cites Berg, Wan- Yen Lo, Piotr DollÃąr, and Ross Girshick.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Berg, Wan- Yen Lo, Piotr DollÃąr, and Ross Girshick

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.166020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.166020Z digest=sha256:e64a9e2d0b6008bd13f2bf5bbcb0b1acaf1fd40c589e05e3dfb49ebf6d421a4a

Observation 41fab339-7f56-48a0-9a2f-c9fb48f9f650 · outbound

This paper cites Similarity of Neural Network Representations Revisited.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Similarity of Neural Network Representations Revisited

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.170332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.170332Z digest=sha256:2b0c7fe4c05f40f2f4b950fa47cff4069b34f4b1f5128352ca4e9d04e3a32c40

Observation 0fbd79ba-4e3a-41c8-86c9-8fcbffbf3bae · outbound

This paper cites CornerNet: Detecting Objects as Paired Keypoints.

Unified Local and Global Attention Interaction Modeling for Vision Transformers CornerNet: Detecting Objects as Paired Keypoints

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.174795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.174795Z digest=sha256:2e52cd0d766c7441a5c49bf5ab1188c73d41ac3c3afd77854856b76374a3da28

Observation 2a5c7759-8818-4780-8eff-913e32be5057 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 21

Resolution
verified exact
doi, observed 2026-08-11T04:34:02.413966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.178742Z digest=sha256:5828418d74b8f8b120794d75bc12e9d88857fac1e69d7acfdfa4593de6d0376b

Observation 1e348a27-9edb-4e0d-91db-1b652d3c0096 · outbound

This paper cites Feature Pyramid Networks for Object Detection.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Feature Pyramid Networks for Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.182700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.182700Z digest=sha256:c841a851263ef1e461fae5b078b900c97ec317ca9adc51b24f1d54cee019df40

Observation 25fce831-bc8a-4da2-abaf-37d25035ce0f · outbound

This paper cites Girshick, Kaiming He , and Piotr Dollár.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Girshick, Kaiming He , and Piotr Dollár

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:03.581831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.186774Z digest=sha256:8d32b8b0133d1d44f86e82df33feb92e441fb6ebf00b59ca1252529dca99ceec

Observation 49fa9b09-8bcd-46f6-851d-962f147d5bd9 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.567191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.195040Z digest=sha256:bcd05ecb19e1e21dfd61c73fc15d8173cba0942463d9cf2f18fb19e0f7a29d37

Observation c840ed8a-a38c-41c4-8cec-df60ba1f1fd0 · outbound

This paper cites SSD: Single Shot MultiBox Detector.

Unified Local and Global Attention Interaction Modeling for Vision Transformers SSD: Single Shot MultiBox Detector

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.204149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.204149Z digest=sha256:2619eb1ccb5a7a9bf6327b895fc79b18c49a253ec2ba71c24a4e4d88887d1f32

Observation e5939e69-769e-4628-bc2f-7753a423bb68 · outbound

This paper cites Swin Transformer V2: Scaling Up Capacity and Resolution.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Swin Transformer V2: Scaling Up Capacity and Resolution

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.208712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.208712Z digest=sha256:c93ce92836d18328c50fc07d19a41605858f47237fa4cbe12dd9b982e4cd86f1

Observation 0d9511f5-4150-4826-96ba-6f00a2c07783 · outbound

This paper cites RTMDet: An Empirical Study of Designing Real-Time Object Detectors.

Unified Local and Global Attention Interaction Modeling for Vision Transformers RTMDet: An Empirical Study of Designing Real-Time Object Detectors

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.217206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.217206Z digest=sha256:01ac57c6d6e2cd4d9b64fb8c0a9376bfaa1585f3025b54b67d429793e7d7084f

Observation c1c495be-6c38-42eb-b4fb-2fd689ef2ef1 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.540028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.221711Z digest=sha256:1a5a8e093afea972177f31dcac79538b7c240eca14c30db4ff75fd19bbbd1cfa

Observation d6eb5e06-7cb7-4b5f-ac19-fbf15ba2712c · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.526104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.230494Z digest=sha256:aedf4d6205024d33bab00fce36ec674c6bcb1186ccfccf05c590df0a4f186e26

Observation 87c790a6-80dd-4393-9a0f-9a4c6f5a99d6 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.234743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.234743Z digest=sha256:1d07dbbb875575606e91716c5fe397674c12238f100408b96551137abcd2c4b2

Observation c6c93072-9be3-4538-a6fb-5a78a6b4ac23 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.511260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.239108Z digest=sha256:4c2caa91862b9792258dc5e831edd87e1db88026779878edaf01a96acb72067d

Observation 834b2e6d-3e01-47f3-9800-b6ba83bd9411 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.495960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.243315Z digest=sha256:20e3cf1507648b71bacabee05e1e88767c83f9b157bb4071d701cfbd64ed3ea7

Observation aa493100-75dd-4e97-98e5-8a631db8c28d · outbound

This paper cites In Computer Vision âĂŞ ECCV 2022 Workshops: Tel A viv, Israel, October 23âĂŞ27, 202 2, Proceedings, Part VII (Tel Aviv, Israel).

Unified Local and Global Attention Interaction Modeling for Vision Transformers In Computer Vision âĂŞ ECCV 2022 Workshops: Tel A viv, Israel, October 23âĂŞ27, 202 2, Proceedings, Part VII (Tel Aviv, Israel)

Reference 34

Resolution
verified exact
doi, observed 2026-08-11T04:34:02.399237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.226254Z digest=sha256:d7438c36814a55643d1c8e73f225a9bb91cee8f89be133bf579f62e975f82514

Observation 3608339d-eaf3-479d-b709-927492599dfb · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.252092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.252092Z digest=sha256:7b99f2d6c89d9010ef8a2f2406bae717b084655e43c2c7c07665d91d606c37cb

Observation 6349ec22-1f98-4cb8-852a-556dc607d144 · outbound

This paper cites You Only Look Once: Unified, Real-Time Object Detection.

Unified Local and Global Attention Interaction Modeling for Vision Transformers You Only Look Once: Unified, Real-Time Object Detection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.256193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.256193Z digest=sha256:1dc51736bb5fefffefa485a1e79d7f117398b9d8748fbd194044ef7f98a762c9

Observation 92ad948b-d174-40e3-a31e-ad24d503f29b · outbound

This paper cites Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.260971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.260971Z digest=sha256:11c61ed003f0e8939dc5e9ede5d716bbd29408eb266f95769a77fc1d293ed596

Observation 6e0e3ab7-ea8d-4460-8a77-87b671302cff · outbound

This paper cites Wu, Safwan S.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Wu, Safwan S

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-11T04:34:02.265565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.265565Z digest=sha256:a23de2977861b21ab32d95c4b253e79393b1be79729f2295a687cbb14b7d25bd

Observation 3c005129-e5c7-4535-b446-c23dfc39da68 · outbound

This paper cites Designing Network Design Spaces.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Designing Network Design Spaces

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.247765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.247765Z digest=sha256:5b115840c7994362aa70f4137bf81f85b7ddc98bcfeecfef21b7d134f394a4ce

Observation ba13177f-9bf0-471e-9195-d98182a76eef · outbound

This paper cites EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks.

Unified Local and Global Attention Interaction Modeling for Vision Transformers EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.274332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.274332Z digest=sha256:3acf3aa874954ecf224ae488b2e788ec9360cb7e31b701f532252a7285cdb432

Observation 16d5f83c-c569-413c-9e76-288417e3fba6 · outbound

This paper cites Attention Is All You Need.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Attention Is All You Need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.278900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.278900Z digest=sha256:238c024969a39f5ed794108f27910fbaced86a12cd02762d31adde23e20dfaa3

Observation 777a7d41-9015-43dd-b854-7f2413e41bec · outbound

This paper cites Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.283296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.283296Z digest=sha256:ea0b72fb2205a1f9385eefa3cf3b220480fd84be544819be49c72ca6a3562fda

Observation 3e7a63b4-e830-4d54-8eb5-e6acd6f21135 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.287811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.287811Z digest=sha256:ff5bf88330b9abdbbbb05f9afe7bb769b0011570a06de1ba99f25718fff9aed1

Observation 7a4333fc-7eb7-43bd-b467-216f5ccae013 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.269929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.269929Z digest=sha256:e72ea376d3adbfebda131a50907201af0277ece9105b4f9563e3960c40e41464

Observation a5c2843f-2825-4012-a6e8-b5efc510faab · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.482056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.296177Z digest=sha256:14591d25277a671469af197e73d8df298879646b9b528d8a88fa5a75f8dd8c8f

Observation c18b2d64-e096-4035-98aa-d7ef4360257b · outbound

This paper cites CvT: Introducing Convolutions to Vision Transformers.

Unified Local and Global Attention Interaction Modeling for Vision Transformers CvT: Introducing Convolutions to Vision Transformers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.303884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.303884Z digest=sha256:f9d71a4d09e9f30fb0cd892ef3e55cf49c1b1e8b1ddcae193f8b3981467695a5

Observation 7e4fd844-f2b2-405a-bf3a-98dd106fe840 · outbound

This paper cites Vision Transformer with Deformable Attention.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Vision Transformer with Deformable Attention

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.307775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.307775Z digest=sha256:d91c235d0dacf3318e5a01e937636e81709c1d1bf328e84280d0f5f62e6ee405

Observation 0cd8eb6a-d2a0-466b-a740-54b8b849facf · outbound

This paper cites DAT++: Spatially Dynamic Vision Transformer with Deformable Attention.

Unified Local and Global Attention Interaction Modeling for Vision Transformers DAT++: Spatially Dynamic Vision Transformer with Deformable Attention

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.312181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.312181Z digest=sha256:0b91f1177de81334b1bb49ba1288282760b2719bbf3b1e968b3dcd058f19cbe8

Observation 111a8269-85dd-4803-87e2-86f69389b322 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.292082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.292082Z digest=sha256:728b296fa0f17b5052ace944618b1a4a0deaf626c1b9552f3f3c0fbf967b49ed

Observation 02108eb6-d90c-429b-9cdb-33eb7c7c405f · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.439246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.320053Z digest=sha256:f9b5e1a30e8683b22623b15a28e3fba6da54bd37e4ada9dfb4edce84b42d8e01

Observation b92fc2b6-7f21-40f3-bdb9-c60dbe31c84b · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 51

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T04:34:03.468138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.300079Z digest=sha256:dad3915120368a27ab424b1a4d7eedfb22da659ecbc3883c51de9b81af975805

Observation e8ab321e-0a2c-4bc5-9805-592a55e3392e · outbound

This paper cites Detecting Twenty-thousand Classes using Image-level Supervision.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Detecting Twenty-thousand Classes using Image-level Supervision

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.329007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.329007Z digest=sha256:0785303a24af9d0fee5fca750fa018e43496adc77923f73f1fdfd579b32f5624

Observation a20076af-da48-43d9-98a2-2ffe6485f391 · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.333517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.333517Z digest=sha256:307b796d9652b238753b8f734518614cd25f47e5d94ccb0c3f8befac4e79ee07

Observation 425a0832-06ff-48f9-aa15-a3124d76c745 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.453889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.316240Z digest=sha256:79ed56589180198f39c29912868add1e965f86046501291f1fdb0f8f92202091

Observation 2d71f19f-a4f0-471d-915b-09ef6c0ecb1f · outbound

This paper cites Refiner: Refining Self-attention for Vision Transformers.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Refiner: Refining Self-attention for Vision Transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.324295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.324295Z digest=sha256:1750b4dbb648e23ef3cf7be72a86ea8954d6b8ef2edaac80d446347823657899

Observation 8a833448-8f08-4bad-9398-4f50009ee08c · outbound

This paper cites Focal Loss for Dense Object Detection.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Focal Loss for Dense Object Detection

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.190793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.190793Z digest=sha256:cf8d1bd4bbea948142878f1dab31f185c664cb29edb3ecadc6ba885e6dbc31a4

Observation 37315f81-30d0-4d19-b9a3-87a0d21739dc · outbound

This paper cites In European Conference on Computer Vision.

Unified Local and Global Attention Interaction Modeling for Vision Transformers In European Conference on Computer Vision

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:03.553975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.199377Z digest=sha256:353ca84ba1f9a479140ee72a02d6d5db6237f50b5309b0fff9ef4106afcf2342

Observation 2abe69fa-245a-45ba-b72f-dc6f285dc221 · outbound

This paper cites In 2023 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR).

Unified Local and Global Attention Interaction Modeling for Vision Transformers In 2023 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR)

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.142594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.142594Z digest=sha256:ce372fa7c24d8bae713a7804a03669ef3698d5f5f40874dc9989a09eea0ecc16

Observation c0795a57-27f5-48de-a036-56cbc66feb09 · outbound

This paper cites SurANet: Surrounding-Aware Network for Concealed Object Detection via Highly-Efficient Interactive Contrastive Learning Strategy.

Unified Local and Global Attention Interaction Modeling for Vision Transformers SurANet: Surrounding-Aware Network for Concealed Object Detection via Highly-Efficient Interactive Contrastive Learning Strategy

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T04:34:03.027113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:02.161232Z digest=sha256:aebe5f2d7b757adb23d36de641204e92757da056bda40770aaf285a35eb2c33a

Pith citing papers

No inbound Pith citation observations are available.