Pith. sign in

Paper Citation Record · LEDGER

Revisiting the Integration of Convolution and Attention for Vision Backbone

As of 21 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 0 inbound Pith citation observations for arXiv:2411.14429.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14429 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:15:22.887919Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

85 of 85 outbound references displayed

  • verified exact0
  • verified fuzzy48
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85ec6c65-1333-490e-ad08-3faccc2cd0f9 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Revisiting the Integration of Convolution and Attention for Vision Backbone Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.473367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.473367Z digest=sha256:106e0274ff88902b8ff56324becfc73febffe0d2845076ab90ba6ed2a8af8472

Observation 6e13fda8-0217-41df-b529-f87e235e8181 · outbound

This paper cites High-performance large-scale image recognition without normalization.

Revisiting the Integration of Convolution and Attention for Vision Backbone High-performance large-scale image recognition without normalization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.479201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.479201Z digest=sha256:2dfd20ad51ee43744d8ff7c0ed3be6990411d450ff8db707505b791f42769401

Observation 575d3e38-d4ce-4536-8244-5200b0306881 · outbound

This paper cites Regionvit: Regional-to-local attention for vision transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Regionvit: Regional-to-local attention for vision transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.485564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.485564Z digest=sha256:a460e4a8afcb6f00390f57584d7f394751ed07b385930b65a17f4fdce7f80f00

Observation 92577866-f032-4765-a7f6-826adf46ba12 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

Revisiting the Integration of Convolution and Attention for Vision Backbone MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.490646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.490646Z digest=sha256:5011af83a0304a5e1d2421609eeea214a5eedf14d9be2214b0c79b2dc717d787

Observation 1b1964f5-579b-45c3-ab6c-d88843bae7cf · outbound

This paper cites Mixformer: Mixing features across windows and dimensions.

Revisiting the Integration of Convolution and Attention for Vision Backbone Mixformer: Mixing features across windows and dimensions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.497258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.497258Z digest=sha256:1dd10a68e6b87bf85701405bce268c6232256088cd059dc883586fdb821d8b34

Observation e0a296c7-cb84-4a71-bf8e-59af214b155d · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Generating Long Sequences with Sparse Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.502531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.502531Z digest=sha256:990a2b55f0890adbc1a1347171aa3fd92bb3a9f007df64bda82ded0054aa4d1c

Observation 8c995fce-d942-4a08-b3c4-1581c1abf564 · outbound

This paper cites Rethinking Attention with Performers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Rethinking Attention with Performers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.508849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.508849Z digest=sha256:61abf54cb3a6b46cbd8a152709684e0fa9a1279deef71262fa2d6b9ea482bcdc

Observation aa796205-2e8a-4e63-bbc0-620eee8759ff · outbound

This paper cites Twins: Revisiting the design of spatial attention in vision transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Twins: Revisiting the design of spatial attention in vision transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.515367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.515367Z digest=sha256:ee87966afe41e64a536379f61503bdad74a1837908f65dd6476d7dba4c52f7f8

Observation 601dfdc7-aa3e-4e1e-afa1-e0cdc0d3945e · outbound

This paper cites Conditional positional encodings for vision transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Conditional positional encodings for vision transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:24.070229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.520564Z digest=sha256:1cb4b6fea277a00086acaa6fc4104b66756347d5db952e89902743fdff397594

Observation 10146473-e53e-44d2-8817-76f176627026 · outbound

This paper cites Openmmlab semantic segmentation toolbox and benchmark.

Revisiting the Integration of Convolution and Attention for Vision Backbone Openmmlab semantic segmentation toolbox and benchmark

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:24.054908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.525310Z digest=sha256:d3fdc750d14e1368648e0471384752394539f317dd508eaa02f248c875579443

Observation 39f87492-986b-4906-b614-23c2a040a523 · outbound

This paper cites Randaugment: Practical automated data augmentation with a reduced search space.

Revisiting the Integration of Convolution and Attention for Vision Backbone Randaugment: Practical automated data augmentation with a reduced search space

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:24.039588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.530096Z digest=sha256:df10d7bfa53d7d0952638e65e1a27a130cb2c2508faaf80e75c8b4eda3695272

Observation b631f5e7-ed14-4224-87e3-bea6341558bd · outbound

This paper cites Coatnet: Marrying convolution and attention for all data sizes.

Revisiting the Integration of Convolution and Attention for Vision Backbone Coatnet: Marrying convolution and attention for all data sizes

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.534952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.534952Z digest=sha256:e639ff4f10e6dd5a24ab90a953cc3b59f5ef90b16a6cfd9640db9587c1f68def

Observation 69998699-e92a-4ad8-8e35-04f0500bb553 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Revisiting the Integration of Convolution and Attention for Vision Backbone Imagenet: A large-scale hierarchical image database

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.539649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.539649Z digest=sha256:fc6e965d0898f7fd66eb15473023f83a4dd5eb8dc52c3d5a32ff58457e58c4bd

Observation f5561a5d-fb6b-4a22-903b-7574b64cc779 · outbound

This paper cites Cswin transformer: A general vision transformer backbone with cross-shaped windows.

Revisiting the Integration of Convolution and Attention for Vision Backbone Cswin transformer: A general vision transformer backbone with cross-shaped windows

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:24.004354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.544343Z digest=sha256:489cb04c91482a220fa3be251e7eb82f436a48acaf90efa1d3657eec66a5cfe0

Observation 5de1cfa8-2c6e-4e69-ad9b-f9bc722a1bd7 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Revisiting the Integration of Convolution and Attention for Vision Backbone An image is worth 16x16 words: Transformers for image recognition at scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.548993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.548993Z digest=sha256:06f402df5a0b7a3740eb7ddfacad634577dc38fc0ea37c20ffd7729b6080f7b2

Observation 277458b5-8800-4377-939c-1bb54742e2b2 · outbound

This paper cites Study on density peaks clustering based on k-nearest neighbors and principal component analysis.

Revisiting the Integration of Convolution and Attention for Vision Backbone Study on density peaks clustering based on k-nearest neighbors and principal component analysis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.977514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.553760Z digest=sha256:c2ab86068a6d0faeb8bbf1b636079007b07de1517860621a7c836a1401d18a38

Observation f7b292ce-d2c2-4461-b342-f63166df136f · outbound

This paper cites Paca-vit: learning patch-to-cluster attention in vision transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Paca-vit: learning patch-to-cluster attention in vision transformers

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.961116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.558335Z digest=sha256:28c676ef9f5ccf1580ad8e30dc6501f690ce9bf4f156925761022b1bfa7a4118

Observation 07734779-3b05-4a23-a3e5-dd1419365dbb · outbound

This paper cites Cmt: Convolutional neural networks meet vision transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Cmt: Convolutional neural networks meet vision transformers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.945582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.562881Z digest=sha256:62fa6cf243d4b00d6e2aefff492cd45829d13d892bbeee5127080b0c11bb50e1

Observation 0c3ed845-77d1-4541-8d97-172f286a8e12 · outbound

This paper cites Flatten transformer: Vision transformer using focused linear attention.

Revisiting the Integration of Convolution and Attention for Vision Backbone Flatten transformer: Vision transformer using focused linear attention

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.930569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.567905Z digest=sha256:b251bcd0219e008318a646f99d41daddb821ae701efd77e6ae3810e67bc41e9e

Observation 060daf9b-4efd-49d2-84d0-09e3485cdb35 · outbound

This paper cites Neighborhood attention transformer.

Revisiting the Integration of Convolution and Attention for Vision Backbone Neighborhood attention transformer

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.915125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.572522Z digest=sha256:437573bd2bef6f57a6a9418b31277e1ae2f69141b6c209fa6354861fdfd14476

Observation 8d34bee4-5db0-4c50-8e03-f995303685a7 · outbound

This paper cites Mask r-cnn.

Revisiting the Integration of Convolution and Attention for Vision Backbone Mask r-cnn

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.898705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.577426Z digest=sha256:a7b31b25eba5482462d4d9f2b435502e2c0bfd11c7abd19abb206f11ec90d8fb

Observation 4e9946a9-0230-4ee2-90dc-e53d280b4777 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Revisiting the Integration of Convolution and Attention for Vision Backbone Masked autoencoders are scalable vision learners

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.582165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.582165Z digest=sha256:34cce84f985e587890a54ed66e4458af4f3f555f504c340b2176bdf4918caa3d

Observation 2456e701-6b9d-4a64-bc77-d77180d10247 · outbound

This paper cites Squeeze-and-excitation networks.

Revisiting the Integration of Convolution and Attention for Vision Backbone Squeeze-and-excitation networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.586747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.586747Z digest=sha256:0eb4d3cbbde218d0c4f3a95123ecf9e7c26cd0211a8c247c585df876d0ee74c9

Observation 29559a36-52fa-4884-86a9-4c6356a7e59e · outbound

This paper cites Deep networks with stochastic depth.

Revisiting the Integration of Convolution and Attention for Vision Backbone Deep networks with stochastic depth

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.863077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.591293Z digest=sha256:adc44b1f003c81213347b64ff6f29a822ae29c6354f552851c58c6c20792f7a9

Observation 349aa6a6-6c53-4e1f-b8f5-6caea1440096 · outbound

This paper cites How Much Position Information Do Convolutional Neural Networks Encode?.

Revisiting the Integration of Convolution and Attention for Vision Backbone How Much Position Information Do Convolutional Neural Networks Encode?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.595672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.595672Z digest=sha256:2ed3211ce8b3e8fb5263ba755025ff520031d4271ff3809430d20d1ad79319d2

Observation 11ea65da-416b-48b3-a128-3407c1d3b1fd · outbound

This paper cites All tokens matter: Token labeling for training better vision transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone All tokens matter: Token labeling for training better vision transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.601161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.601161Z digest=sha256:04368c16bab34910ac8071062f20143117f2373e0bf75236a99e25d1f41f4b4f

Observation 3afcd94a-b170-45e8-ae78-976303ae3437 · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

Revisiting the Integration of Convolution and Attention for Vision Backbone Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.605844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.605844Z digest=sha256:e49c0266aaf8ced402280094409fbc286b9b6703ba0c0c9be09eaa7ad8dd7a02

Observation 56b2cf3a-529c-43b0-84e4-3dd26e28a4e8 · outbound

This paper cites Panoptic feature pyramid networks.

Revisiting the Integration of Convolution and Attention for Vision Backbone Panoptic feature pyramid networks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.827731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.610377Z digest=sha256:f4534479d7f5fda65d89c6f842d48c130650334372e5f3f1b57c278ff4f547ee

Observation 3f2e5312-18fa-4a6e-9acd-2b1a84ea9e56 · outbound

This paper cites Uniformer: Unifying convolution and self-attention for visual recognition.

Revisiting the Integration of Convolution and Attention for Vision Backbone Uniformer: Unifying convolution and self-attention for visual recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.812703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.615020Z digest=sha256:fe27faf085b9fd7f0470bef365e22d2606444ba59065f6e366dfcf3b527686d8

Observation 8478aeed-fa30-49c4-8dc7-8bece1ab76df · outbound

This paper cites Clusterfomer: clustering as a universal visual learner.

Revisiting the Integration of Convolution and Attention for Vision Backbone Clusterfomer: clustering as a universal visual learner

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.797270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.619573Z digest=sha256:fedbea670c58fe2fa3791ba326401d0b0b22ed8bca73ffc35c3caf1ca990e2b5

Observation b677d9d5-eff8-4b45-a739-db88e90842ed · outbound

This paper cites Microsoft coco: Common objects in context.

Revisiting the Integration of Convolution and Attention for Vision Backbone Microsoft coco: Common objects in context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.781713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.624233Z digest=sha256:441f3cf046cbe4bbe1f0d54ff4c9b72d9b41ae36bf097a87b7c859943a53db17

Observation 28f92658-c9cc-4cad-9021-b179b3f898bd · outbound

This paper cites Feature pyramid networks for object detection.

Revisiting the Integration of Convolution and Attention for Vision Backbone Feature pyramid networks for object detection

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.628846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.628846Z digest=sha256:775f751f5761bbb7fdd6d4653f47f2de8cf2168eb4385becb763932d5531f3f0

Observation 0814d827-6fec-46ea-bb0d-ed6ede6cc3f1 · outbound

This paper cites Focal loss for dense object detection.

Revisiting the Integration of Convolution and Attention for Vision Backbone Focal loss for dense object detection

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.633545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.633545Z digest=sha256:f96ccad9253949bbd0093df164b91af8b580495a48c1861ffea70d666db63dd5

Observation b2ed5c2b-24ed-4b35-8291-8bd5a6cc7efe · outbound

This paper cites Scale-aware modulation meet transformer.

Revisiting the Integration of Convolution and Attention for Vision Backbone Scale-aware modulation meet transformer

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.746302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.638218Z digest=sha256:a9e8898e1e4ef14698a5d454d8cb21b03235504ba71e40ebb6777b54f1618a2d

Observation d78dbcfa-7982-42f5-80ef-046cbea6162e · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Revisiting the Integration of Convolution and Attention for Vision Backbone Swin transformer: Hierarchical vision transformer using shifted windows

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.642872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.642872Z digest=sha256:8752f8c6fb78d8777922505819036d4d64bbb2cb470b9c1bd85bbb8691345d02

Observation 20fd367f-32ad-450f-ae16-7586f1e1f546 · outbound

This paper cites A convnet for the 2020s.

Revisiting the Integration of Convolution and Attention for Vision Backbone A convnet for the 2020s

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.648958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.648958Z digest=sha256:ef0307e44efabd90619cf7a8e5870d1971d6a5637df93e5f7d665c88d3a80141

Observation 08e274b4-2eb3-434d-a2a2-6cec33af6aa0 · outbound

This paper cites Stochastic gradient descent with warm restarts.

Revisiting the Integration of Convolution and Attention for Vision Backbone Stochastic gradient descent with warm restarts

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.710531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.653860Z digest=sha256:81cc59bb3c080df915041ed0d2cc9777698d02afb5a347df166ca3cf5770101c

Observation 67f1a967-156b-4f1d-9352-ed25ba37e34e · outbound

This paper cites Decoupled weight decay regularization.

Revisiting the Integration of Convolution and Attention for Vision Backbone Decoupled weight decay regularization

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.694264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.658430Z digest=sha256:b601e087e586259d8f30033765aa9b433e0fa0c9561879fb32583f6e2b145ef2

Observation 811dda20-5f93-4a53-bf67-b916f19d64f4 · outbound

This paper cites Understanding the effective receptive field in deep convolutional neural networks.

Revisiting the Integration of Convolution and Attention for Vision Backbone Understanding the effective receptive field in deep convolutional neural networks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.663125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.663125Z digest=sha256:249e797e818d49597ba1cda9d8e5d8a0b5fbc51d45a7260ea548feedc1c3beec

Observation d26f9cf3-eedb-4ff9-94bf-bedf9a98ed08 · outbound

This paper cites On the integration of self-attention and convolution.

Revisiting the Integration of Convolution and Attention for Vision Backbone On the integration of self-attention and convolution

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.669184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.667715Z digest=sha256:51676c768ed95f9c8f8489884ecc56c29074647dd428a6807bc10b1a3f762070

Observation 86f18108-76e6-4d03-89f4-aad15f100b5e · outbound

This paper cites How do vision transformers work? In International Conference on Learning Representations.

Revisiting the Integration of Convolution and Attention for Vision Backbone How do vision transformers work? In International Conference on Learning Representations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.653251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.673152Z digest=sha256:df35587f3871af03be51faaedb59212f1745d230564d3c63944bde08c552e253

Observation 8c52ff4f-e782-485d-8c4a-7917b806d9f7 · outbound

This paper cites From Sparse to Soft Mixtures of Experts.

Revisiting the Integration of Convolution and Attention for Vision Backbone From Sparse to Soft Mixtures of Experts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.677701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.677701Z digest=sha256:df7f934e1382452b002da270f93367ce4d81e54f78298c14375a008c01e62e2c

Observation 7d2d9425-0051-4a0c-9647-f8003afb9af8 · outbound

This paper cites Sg-former: Self-guided transformer with evolving token reallocation.

Revisiting the Integration of Convolution and Attention for Vision Backbone Sg-former: Self-guided transformer with evolving token reallocation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.637641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.683934Z digest=sha256:b3fa963b1947c015d650d75f36b56a574aeaeba478876a8df85516778f151475

Observation 8f447da8-b00f-4745-a3d5-37c81f920de4 · outbound

This paper cites GLU Variants Improve Transformer.

Revisiting the Integration of Convolution and Attention for Vision Backbone GLU Variants Improve Transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.688586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.688586Z digest=sha256:06f792e358934cf37b5c21fb3f4f1e10fbcc87fe33a00636aabe6da857e7247f

Observation c957a076-e7b5-462b-bc6a-a893ef048b2e · outbound

This paper cites Rethinking the inception architecture for computer vision.

Revisiting the Integration of Convolution and Attention for Vision Backbone Rethinking the inception architecture for computer vision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.620814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.693643Z digest=sha256:e5ea49b9b9c53095dc0d247340cbd54ec49b2846727b8f754413a2e7dc393755

Observation b8d1b302-4d1a-4165-9f47-6bae7e6a7358 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

Revisiting the Integration of Convolution and Attention for Vision Backbone Training data-efficient image transformers & distillation through attention

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.699611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.699611Z digest=sha256:0ccd816d7639cdaf5f1e73db3bab84a0263df39dbb043ba5d8c8e5631130082f

Observation b3f5029a-4064-4d07-bc5e-a06e52549b7c · outbound

This paper cites Maxvit: Multi-axis vision transformer.

Revisiting the Integration of Convolution and Attention for Vision Backbone Maxvit: Multi-axis vision transformer

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.595037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.704544Z digest=sha256:819438fd7e09ef33621700526483adaf563f5474adb26528166ccbd405fcbf52

Observation 097efe03-527f-4a90-80af-8915d55197d6 · outbound

This paper cites Attention is all you need.

Revisiting the Integration of Convolution and Attention for Vision Backbone Attention is all you need

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.709136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.709136Z digest=sha256:e4558ae2654e74fb7948edfe0a9ea4e4c51eb81ccd28940163c6017557712da5

Observation 25f076be-721c-4358-a2c3-630252303d85 · outbound

This paper cites Fast transformers with clustered attention.

Revisiting the Integration of Convolution and Attention for Vision Backbone Fast transformers with clustered attention

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.569689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.713791Z digest=sha256:83c8666b862690a547b4776ce648095f3973c608872e58ca21e172ddb6c4d1e0

Observation 671dc6b3-bceb-4adb-ba08-247306c7051a · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Revisiting the Integration of Convolution and Attention for Vision Backbone Linformer: Self-Attention with Linear Complexity

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.718558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.718558Z digest=sha256:35376f56a756c883561958bce49be47ff88a227d77b32bfa36c12d7e65d4db8b

Observation 3f160672-9376-4a07-8645-75a305ce7584 · outbound

This paper cites Pyramid vision transformer: A versatile backbone for dense prediction without convolutions.

Revisiting the Integration of Convolution and Attention for Vision Backbone Pyramid vision transformer: A versatile backbone for dense prediction without convolutions

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.554280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.723620Z digest=sha256:d9cbd8d5a52d639da5f0c2d6671241b6f08ed510105b2f989f1b54747289734c

Observation 46981908-6003-4abf-aba5-5d9a28dbf5c7 · outbound

This paper cites Internimage: Exploring large-scale vision foundation models with deformable convolutions.

Revisiting the Integration of Convolution and Attention for Vision Backbone Internimage: Exploring large-scale vision foundation models with deformable convolutions

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.728316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.728316Z digest=sha256:ea07ed7c4ef1fc9daafd07249e91b1db80b9288e3f179b88f9b492dded0437c8

Observation 37546beb-fec2-4028-8a5a-67d65e1bb292 · outbound

This paper cites Crossformer: A versatile vision transformer hinging on cross-scale attention.

Revisiting the Integration of Convolution and Attention for Vision Backbone Crossformer: A versatile vision transformer hinging on cross-scale attention

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.526889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.732770Z digest=sha256:b9989028c3ae5a31d49cb69b2d7699ad627b7b0ffa2cf419ee5042c08b79a0cd

Observation e25f9852-2ff8-4577-b39d-c1a3fd08fe2a · outbound

This paper cites Pytorch image models.

Revisiting the Integration of Convolution and Attention for Vision Backbone Pytorch image models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.737393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.737393Z digest=sha256:7d8b6f157e49da8f29187936912f61e13ca96ee0b292e95dc4c399aa5c4c09f9

Observation 2363af27-ee8c-462c-b387-5b23004cf27b · outbound

This paper cites Cvt: Introducing convolutions to vision transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Cvt: Introducing convolutions to vision transformers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.742139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.742139Z digest=sha256:3518ab2b0fcbfff29ae380095b42fd1f177fd2c06420e9b770e061cf2e936a7f

Observation 72b7e3a2-3828-4fb0-8c1f-9b6fac3b54cd · outbound

This paper cites Vision transformer with deformable attention.

Revisiting the Integration of Convolution and Attention for Vision Backbone Vision transformer with deformable attention

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.490534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.746602Z digest=sha256:509c8a060c6bdb730b22f635b3d8905fc80464ab169b7f98931e9caec3d59f75

Observation 432a43ca-7941-4232-8daa-c5a5f9d37b39 · outbound

This paper cites Unified perceptual parsing for scene understanding.

Revisiting the Integration of Convolution and Attention for Vision Backbone Unified perceptual parsing for scene understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.471885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.751192Z digest=sha256:9efbaf3259082045e1a061b33f5a2fd778ab7a2bb4dd651b107a3ad4f9902214

Observation c57ac0cd-f838-4904-b964-df1166874530 · outbound

This paper cites ClusTR: Exploring Efficient Self-attention via Clustering for Vision Transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone ClusTR: Exploring Efficient Self-attention via Clustering for Vision Transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.755707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.755707Z digest=sha256:353078bbeb36b33326d7ee1f359a238b11e29ce130047571343524180da0a4ab

Observation 6273dae9-623d-47b8-a650-862cb13558df · outbound

This paper cites Moat: Alternating mobile convolution and attention brings strong vision models.

Revisiting the Integration of Convolution and Attention for Vision Backbone Moat: Alternating mobile convolution and attention brings strong vision models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.456508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.760844Z digest=sha256:4d3b82bf0b7fef4a8c9ef95e91092d631ffbaf33155572aecac5287e5923f47f

Observation 127dccab-3294-49da-a40b-af70cca364a3 · outbound

This paper cites Scalablevit: Rethinking the context-oriented generalization of vision transformer.

Revisiting the Integration of Convolution and Attention for Vision Backbone Scalablevit: Rethinking the context-oriented generalization of vision transformer

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.440771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.765408Z digest=sha256:ed4b410465be35ce98eed870e97eb200c5e19ae56acf93c9e8e922e5ec135d40

Observation 781f0422-c2e3-43c2-9731-2aadaa43b699 · outbound

This paper cites Wave-vit: Unifying wavelet and transformers for visual representation learning.

Revisiting the Integration of Convolution and Attention for Vision Backbone Wave-vit: Unifying wavelet and transformers for visual representation learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.423595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.769944Z digest=sha256:d0b5b3f4605d12bc60c95fe725322099750475597d11e8da85754707bf504fcf

Observation c6cc6a64-fbd2-430b-b27f-7529164bee92 · outbound

This paper cites Dual vision transformer.

Revisiting the Integration of Convolution and Attention for Vision Backbone Dual vision transformer

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.408500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.774538Z digest=sha256:3962822d4f214d91f5dd2feb6227fd98b8ff327a7be26f42dab75858d457c4ad

Observation 3104cb3c-de13-4103-9518-77fe4e6b2634 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Revisiting the Integration of Convolution and Attention for Vision Backbone mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.778999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.778999Z digest=sha256:51b94062c1f2d5d0d45654c60e359e0a2dc2214b61b6570782618cdb0186fa98

Observation f37b6ea9-5c22-4d9f-9f7c-e9eb59b32ebd · outbound

This paper cites Metaformer is actually what you need for vision.

Revisiting the Integration of Convolution and Attention for Vision Backbone Metaformer is actually what you need for vision

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.783904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.783904Z digest=sha256:0ed306d5eb17b1426690f2fbb831b4f9e0e6389bd173cb53c498e55165939277

Observation 96127753-1662-4284-bb34-aa11a01db413 · outbound

This paper cites Cutmix: Regularization strategy to train strong classifiers with localizable features.

Revisiting the Integration of Convolution and Attention for Vision Backbone Cutmix: Regularization strategy to train strong classifiers with localizable features

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.789468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.789468Z digest=sha256:6f6f39ebd578d9f851cb92baaee40418d5358c25310998bd5e0601175a8b5c54

Observation d8f1911c-4a38-40ff-b5e0-44c2b85ef318 · outbound

This paper cites Not all tokens are equal: Human-centric visual analysis via token clustering transformer.

Revisiting the Integration of Convolution and Attention for Vision Backbone Not all tokens are equal: Human-centric visual analysis via token clustering transformer

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.373093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.794116Z digest=sha256:d63f4f5046b5d3473bfbf463fe1b05345e79cfc23ea48c1f151a8ad7d68977c6

Observation f871d40c-118c-4d34-a125-1b3cb92efa2a · outbound

This paper cites Dauphin, and David Lopez-Paz.

Revisiting the Integration of Convolution and Attention for Vision Backbone Dauphin, and David Lopez-Paz

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.357107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.799921Z digest=sha256:693f355c62a1de0350ff3ecfe673a5b97fc638d448d708217cdd60ac9dac110d

Observation 99390d75-5c34-4505-9944-84d99160ca84 · outbound

This paper cites Dense distinct query for end-to-end object detection.

Revisiting the Integration of Convolution and Attention for Vision Backbone Dense distinct query for end-to-end object detection

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.342338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.804553Z digest=sha256:f8e4a7f0928260e1ac7cea18fefc9a7809e89004d82b10ef7da6bec9a207bb2e

Observation 101fe6df-b352-47df-a1b5-995a62413be6 · outbound

This paper cites Semantic understanding of scenes through the ade20k dataset.

Revisiting the Integration of Convolution and Attention for Vision Backbone Semantic understanding of scenes through the ade20k dataset

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.809225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.809225Z digest=sha256:4fe4c1442b3a237d34158df4ca4abff64f59c22ebb4f18b1fbf1565151f63e9d

Observation 177a1496-b5a7-426c-a01f-b677ad510484 · outbound

This paper cites Biformer: Vision transformer with bi-level routing attention.

Revisiting the Integration of Convolution and Attention for Vision Backbone Biformer: Vision transformer with bi-level routing attention

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.316422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.813803Z digest=sha256:c7c0346773522c98df185fa0b63e9cce1c51e0bb8c03aa510e04ea2cf6524195

Observation 47e9e79a-7146-40d9-968c-9b10fe497bb2 · outbound

This paper cites The contributions are summarized at the end of the Introduction (Sec.

Revisiting the Integration of Convolution and Attention for Vision Backbone The contributions are summarized at the end of the Introduction (Sec

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.301750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.818678Z digest=sha256:b444a5a1ca04392a5bd00ba1f816dfaf6d3face3094fc339402bb149205a6a60

Observation 61e3f61e-ef81-41b8-bad8-74c9db0db45c · outbound

This paper cites Limitations.

Revisiting the Integration of Convolution and Attention for Vision Backbone Limitations

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.286061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.823780Z digest=sha256:8ffaddf9eab4925f1a8765a3256e3d75274fd590cc98d0fbd7003d4cbe9f862f

Observation 64ea8b7e-7c8b-4a03-ba5f-f73450c340ae · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Revisiting the Integration of Convolution and Attention for Vision Backbone Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.271088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.828772Z digest=sha256:649a79b8adc7cdf81d3351969f1110a86e3807d61358656358a392ec7c57e085

Observation 54e4fca1-90b6-4f1d-9cd1-bb07ad6ab8e5 · outbound

This paper cites an unresolved cited work.

Revisiting the Integration of Convolution and Attention for Vision Backbone Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:15:23.255235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.833652Z digest=sha256:68e8195bd8faccfddec3836d0deb2d4cd25102bf760ef4e71fa5564a07d6dd46

Observation 6093176e-68bb-4a1c-9bb9-65cdeda95d99 · outbound

This paper cites We will release the code upon the acceptance of this paper.

Revisiting the Integration of Convolution and Attention for Vision Backbone We will release the code upon the acceptance of this paper

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.239351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.838807Z digest=sha256:33f017e22169d1196433867200d3dab91deec1ce9ffc0bf8d28b60d8d4406070

Observation c1aad93e-ca16-4315-a9a3-41f42c0312e3 · outbound

This paper cites an unresolved cited work.

Revisiting the Integration of Convolution and Attention for Vision Backbone Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:15:23.223830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.843504Z digest=sha256:66e83ce58d5dbdda33d6bc9d4a6dbb6d6717983c6490fc908456627b3023b14b

Observation 39acc319-7392-46c0-b341-9a35519da408 · outbound

This paper cites The experimental results are not sensitive to random initialization.

Revisiting the Integration of Convolution and Attention for Vision Backbone The experimental results are not sensitive to random initialization

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.208439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.849169Z digest=sha256:3cee2784508197701ffb5ee482060d44f26b02980ac7e0f7d7b6ef85b69f443b

Observation ba19d142-b8d1-42a2-b77a-88c08d78cd56 · outbound

This paper cites an unresolved cited work.

Revisiting the Integration of Convolution and Attention for Vision Backbone Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:15:23.193193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.853815Z digest=sha256:397559d26f9db41721871d7e865ebd7fc27963f7ab4cef3e036419d2ec260b5f

Observation 5b5c2858-117e-415c-ba42-37ba7f8a4875 · outbound

This paper cites We only use existing and publicly available datasets for evaluations.

Revisiting the Integration of Convolution and Attention for Vision Backbone We only use existing and publicly available datasets for evaluations

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.177914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.858221Z digest=sha256:66b164865bac5727373f166100f87b62b1d3fa995b8798aa3ca99cd728e1dad0

Observation 5174d3b6-a2ba-4438-bae1-c4cd00b5589f · outbound

This paper cites It is too broad to discuss the societal impacts of such a general topic.

Revisiting the Integration of Convolution and Attention for Vision Backbone It is too broad to discuss the societal impacts of such a general topic

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.162268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.863391Z digest=sha256:f9391977b19c28578ef27cec8727bf17e02eb5b98f9f6d511025b3a06363d262

Observation accd94e5-e3eb-474d-9757-15c9fb2dfc09 · outbound

This paper cites It poses no such risks.

Revisiting the Integration of Convolution and Attention for Vision Backbone It poses no such risks

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.145084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.868619Z digest=sha256:93780eef6527280303be603c1752fb3f7e638d21582859bc8bc6de9dfd4d28f5

Observation 5df8bdaf-1c5c-416d-98ae-f241e660f4b4 · outbound

This paper cites These works are properly cited.

Revisiting the Integration of Convolution and Attention for Vision Backbone These works are properly cited

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.129743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.873273Z digest=sha256:a12f835685b03c10047a0f4d74381920d68d73bdd848774ae03d0e1eb482fa75

Observation fc9cc2ca-ae0b-41e7-bfe3-79412e961492 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Revisiting the Integration of Convolution and Attention for Vision Backbone Guidelines: • The answer NA means that the paper does not release new assets

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.114445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.878078Z digest=sha256:5eed744daeed3501735c95a340ec5fc0cb70dd88cda94acaea90445e97a7331e

Observation af19c8cc-0c12-4688-bcdc-09c8d8769274 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Revisiting the Integration of Convolution and Attention for Vision Backbone Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.098801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.883149Z digest=sha256:5e427353608ea8ebdd4b9cb070bb5b117a18cf1855b79cd79f2c868a17275f5c

Observation 55fae290-e9b2-42a3-a17b-34eb33397eda · outbound

This paper cites Guidelines: 23 • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Revisiting the Integration of Convolution and Attention for Vision Backbone Guidelines: 23 • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.082865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-12T15:15:22.887919Z digest=sha256:6ebe20ad4b0813d945a718b43a98465a4d4531b9fa742dff71e5615145f7261a

Pith citing papers

No inbound Pith citation observations are available.