Pith. sign in

Paper Citation Record · LEDGER

Revisiting the Integration of Convolution and Attention for Vision Backbone

As of 13 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 0 inbound Pith citation observations for arXiv:2411.14429.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.14429 v1

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:15:22.887919Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

85 of 85 outbound references displayed

  • verified exact0
  • verified fuzzy48
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 85ec6c65-1333-490e-ad08-3faccc2cd0f9 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Revisiting the Integration of Convolution and Attention for Vision Backbone Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.473367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.473367Z digest=sha256:bf05cb0a42af291b8cb14c1cd0d89e96d26683751af611fbf1d4144ed42a9873

Observation 6e13fda8-0217-41df-b529-f87e235e8181 · outbound

This paper cites High-performance large-scale image recognition without normalization.

Revisiting the Integration of Convolution and Attention for Vision Backbone High-performance large-scale image recognition without normalization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.479201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.479201Z digest=sha256:0ac6075fc267674cbc23d3694cb220a82dc246e69b159e405969853170b8ddf4

Observation 575d3e38-d4ce-4536-8244-5200b0306881 · outbound

This paper cites Regionvit: Regional-to-local attention for vision transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Regionvit: Regional-to-local attention for vision transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.485564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.485564Z digest=sha256:f79692fbfb34d776863060c7fb702fb75bc0896c3a6ebc53f470dfe598da10e8

Observation 92577866-f032-4765-a7f6-826adf46ba12 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

Revisiting the Integration of Convolution and Attention for Vision Backbone MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.490646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.490646Z digest=sha256:ab6b51e73f09b025b3d2f1a57c6f2bb44f2c89e8cc5477048c95a1d8ca4017fc

Observation 1b1964f5-579b-45c3-ab6c-d88843bae7cf · outbound

This paper cites Mixformer: Mixing features across windows and dimensions.

Revisiting the Integration of Convolution and Attention for Vision Backbone Mixformer: Mixing features across windows and dimensions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.497258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.497258Z digest=sha256:f03bfff05f2752c3cbcc66cd24867cb78e28ea4ff4e83beb8e7c8cb7af2704bd

Observation e0a296c7-cb84-4a71-bf8e-59af214b155d · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Generating Long Sequences with Sparse Transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.502531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.502531Z digest=sha256:cd3533553be07444a64b1a42c91b0d5b02c2bed3cf6b29240ccf849e04c5cef0

Observation 8c995fce-d942-4a08-b3c4-1581c1abf564 · outbound

This paper cites Rethinking Attention with Performers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Rethinking Attention with Performers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.508849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.508849Z digest=sha256:29ab45d69eaf1f5b41a2e5cc638ca414d302e6a4c49a819a9a9a03a266b9b233

Observation aa796205-2e8a-4e63-bbc0-620eee8759ff · outbound

This paper cites Twins: Revisiting the design of spatial attention in vision transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Twins: Revisiting the design of spatial attention in vision transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.515367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.515367Z digest=sha256:797f85964670e8be4894948419068adf0b6bd14770e15b302858fae745ea8ee9

Observation 601dfdc7-aa3e-4e1e-afa1-e0cdc0d3945e · outbound

This paper cites Conditional positional encodings for vision transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Conditional positional encodings for vision transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:24.070229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.520564Z digest=sha256:42b2ea23ebe7681a27e98931715ce5e0c0505116ea4b56c882773dc77c2f4d38

Observation 10146473-e53e-44d2-8817-76f176627026 · outbound

This paper cites Openmmlab semantic segmentation toolbox and benchmark.

Revisiting the Integration of Convolution and Attention for Vision Backbone Openmmlab semantic segmentation toolbox and benchmark

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:24.054908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.525310Z digest=sha256:c1ddf3b821c4fd1a1354d56b586b54588bdbecaec878a9c1ec7a025a50a48a6d

Observation 39f87492-986b-4906-b614-23c2a040a523 · outbound

This paper cites Randaugment: Practical automated data augmentation with a reduced search space.

Revisiting the Integration of Convolution and Attention for Vision Backbone Randaugment: Practical automated data augmentation with a reduced search space

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:24.039588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.530096Z digest=sha256:004a4ee4bf9f351a9eccf58c97b45046687c6981ad64379e2d6eefefa468d661

Observation b631f5e7-ed14-4224-87e3-bea6341558bd · outbound

This paper cites Coatnet: Marrying convolution and attention for all data sizes.

Revisiting the Integration of Convolution and Attention for Vision Backbone Coatnet: Marrying convolution and attention for all data sizes

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.534952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.534952Z digest=sha256:8cfaee5417255df7abb15f3e03d1a861a68e3d52d049db328839060e9f199e7e

Observation 69998699-e92a-4ad8-8e35-04f0500bb553 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Revisiting the Integration of Convolution and Attention for Vision Backbone Imagenet: A large-scale hierarchical image database

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.539649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.539649Z digest=sha256:1da43985c9065d356062340be659d1ce987f300647705f7256b0a2cc547ffa64

Observation f5561a5d-fb6b-4a22-903b-7574b64cc779 · outbound

This paper cites Cswin transformer: A general vision transformer backbone with cross-shaped windows.

Revisiting the Integration of Convolution and Attention for Vision Backbone Cswin transformer: A general vision transformer backbone with cross-shaped windows

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:24.004354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.544343Z digest=sha256:b9324fcab87f872fc6021026b0fcc64adb8899d1ea301012d50ad07cfe35f264

Observation 5de1cfa8-2c6e-4e69-ad9b-f9bc722a1bd7 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Revisiting the Integration of Convolution and Attention for Vision Backbone An image is worth 16x16 words: Transformers for image recognition at scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.548993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.548993Z digest=sha256:796f03294a01a760122f83ce681b5bb94cc85df046b9e7d283d75182ad4d0559

Observation 277458b5-8800-4377-939c-1bb54742e2b2 · outbound

This paper cites Study on density peaks clustering based on k-nearest neighbors and principal component analysis.

Revisiting the Integration of Convolution and Attention for Vision Backbone Study on density peaks clustering based on k-nearest neighbors and principal component analysis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.977514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.553760Z digest=sha256:694d8d0c9fc7af42f85f189abb97fc01cdafbd6b3af7f6a3850e84dd206dbfce

Observation f7b292ce-d2c2-4461-b342-f63166df136f · outbound

This paper cites Paca-vit: learning patch-to-cluster attention in vision transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Paca-vit: learning patch-to-cluster attention in vision transformers

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.961116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.558335Z digest=sha256:c76ac8553ac34d464252dce28c50be403841df4a24280b6403c336767ce18a0f

Observation 07734779-3b05-4a23-a3e5-dd1419365dbb · outbound

This paper cites Cmt: Convolutional neural networks meet vision transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Cmt: Convolutional neural networks meet vision transformers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.945582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.562881Z digest=sha256:177e9b2006a45a3181113b16dc3ec75080a115849d6c43327228bd7111e16b0e

Observation 0c3ed845-77d1-4541-8d97-172f286a8e12 · outbound

This paper cites Flatten transformer: Vision transformer using focused linear attention.

Revisiting the Integration of Convolution and Attention for Vision Backbone Flatten transformer: Vision transformer using focused linear attention

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.930569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.567905Z digest=sha256:ae0cfd17234a6325781be0e8e1f155176ac00e71ba5339e684cf5eef955625f8

Observation 060daf9b-4efd-49d2-84d0-09e3485cdb35 · outbound

This paper cites Neighborhood attention transformer.

Revisiting the Integration of Convolution and Attention for Vision Backbone Neighborhood attention transformer

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.915125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.572522Z digest=sha256:316dccf95dca4e12399c663c58b8b22e2ae2a3c619d697ff1e5d87c39ab02cea

Observation 8d34bee4-5db0-4c50-8e03-f995303685a7 · outbound

This paper cites Mask r-cnn.

Revisiting the Integration of Convolution and Attention for Vision Backbone Mask r-cnn

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.898705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.577426Z digest=sha256:c6cb8897c44aea394fbb5147af240f86febef8087436e3ac1314da91c0514c62

Observation 4e9946a9-0230-4ee2-90dc-e53d280b4777 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Revisiting the Integration of Convolution and Attention for Vision Backbone Masked autoencoders are scalable vision learners

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.582165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.582165Z digest=sha256:c2777553d10f83fca7877b3f6c8af8debd7b6987655ca8e7947627b897d71771

Observation 2456e701-6b9d-4a64-bc77-d77180d10247 · outbound

This paper cites Squeeze-and-excitation networks.

Revisiting the Integration of Convolution and Attention for Vision Backbone Squeeze-and-excitation networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.586747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.586747Z digest=sha256:e63fbb7e61eb2b9184877e24f429aded30ab873191189ef183c49b2d5cf81cd2

Observation 29559a36-52fa-4884-86a9-4c6356a7e59e · outbound

This paper cites Deep networks with stochastic depth.

Revisiting the Integration of Convolution and Attention for Vision Backbone Deep networks with stochastic depth

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.863077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.591293Z digest=sha256:e5cefe51002550ff06122fc45c7a5d3ba812bba9ff7a25400710b5bb05691476

Observation 349aa6a6-6c53-4e1f-b8f5-6caea1440096 · outbound

This paper cites How Much Position Information Do Convolutional Neural Networks Encode?.

Revisiting the Integration of Convolution and Attention for Vision Backbone How Much Position Information Do Convolutional Neural Networks Encode?

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.595672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.595672Z digest=sha256:f3f6caa6ac4d8a52315886cb02d3ad06beb563ace3c73779eef3fa5365f25d40

Observation 11ea65da-416b-48b3-a128-3407c1d3b1fd · outbound

This paper cites All tokens matter: Token labeling for training better vision transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone All tokens matter: Token labeling for training better vision transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.601161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.601161Z digest=sha256:d1be0e5430677e413e656420d2a271660ff572d30115c8fb90de987f5da410ab

Observation 3afcd94a-b170-45e8-ae78-976303ae3437 · outbound

This paper cites Transformers are rnns: Fast autoregressive transformers with linear attention.

Revisiting the Integration of Convolution and Attention for Vision Backbone Transformers are rnns: Fast autoregressive transformers with linear attention

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.605844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.605844Z digest=sha256:fca905c2dac7b8625ce416d9b79f134382272c68ed96e84f812e3d634c07232c

Observation 56b2cf3a-529c-43b0-84e4-3dd26e28a4e8 · outbound

This paper cites Panoptic feature pyramid networks.

Revisiting the Integration of Convolution and Attention for Vision Backbone Panoptic feature pyramid networks

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.827731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.610377Z digest=sha256:30ebfc9f16646a7f70f72cb3cf3ce1cafc33e8eaae438adaf2c8c0985164c07c

Observation 3f2e5312-18fa-4a6e-9acd-2b1a84ea9e56 · outbound

This paper cites Uniformer: Unifying convolution and self-attention for visual recognition.

Revisiting the Integration of Convolution and Attention for Vision Backbone Uniformer: Unifying convolution and self-attention for visual recognition

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.812703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.615020Z digest=sha256:4c1794aaf7ebcec32bbeaa6b34b380a9003e6fb2ba93e7336b112bd4b9f6106a

Observation 8478aeed-fa30-49c4-8dc7-8bece1ab76df · outbound

This paper cites Clusterfomer: clustering as a universal visual learner.

Revisiting the Integration of Convolution and Attention for Vision Backbone Clusterfomer: clustering as a universal visual learner

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.797270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.619573Z digest=sha256:d105e56df714e0564a512f014ed23225b83bfbec6b23bb26f536c83fdcb4ee08

Observation b677d9d5-eff8-4b45-a739-db88e90842ed · outbound

This paper cites Microsoft coco: Common objects in context.

Revisiting the Integration of Convolution and Attention for Vision Backbone Microsoft coco: Common objects in context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.781713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.624233Z digest=sha256:156d07becc8e6f235f2796a8fed8f04f3d918b3be834db897fc62ba52cf8fd4a

Observation 28f92658-c9cc-4cad-9021-b179b3f898bd · outbound

This paper cites Feature pyramid networks for object detection.

Revisiting the Integration of Convolution and Attention for Vision Backbone Feature pyramid networks for object detection

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.628846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.628846Z digest=sha256:1b01a2d463fd1a5bc0fc96c798ea650cd63429840a60c308ee65aa388d245fb1

Observation 0814d827-6fec-46ea-bb0d-ed6ede6cc3f1 · outbound

This paper cites Focal loss for dense object detection.

Revisiting the Integration of Convolution and Attention for Vision Backbone Focal loss for dense object detection

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.633545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.633545Z digest=sha256:295d956010a53280edaa7527e4a6f6d87ccdb3189ea4b042d2373e5c651229e7

Observation b2ed5c2b-24ed-4b35-8291-8bd5a6cc7efe · outbound

This paper cites Scale-aware modulation meet transformer.

Revisiting the Integration of Convolution and Attention for Vision Backbone Scale-aware modulation meet transformer

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.746302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.638218Z digest=sha256:c96e93066c15fcaa526967967c1541ba13383015d8bed40b9dca61f1dfb37fb3

Observation d78dbcfa-7982-42f5-80ef-046cbea6162e · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Revisiting the Integration of Convolution and Attention for Vision Backbone Swin transformer: Hierarchical vision transformer using shifted windows

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.642872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.642872Z digest=sha256:89970661e183733502c4a6eb931a7c4bce8785c255222a68fc3f4a56fd8c57d6

Observation 20fd367f-32ad-450f-ae16-7586f1e1f546 · outbound

This paper cites A convnet for the 2020s.

Revisiting the Integration of Convolution and Attention for Vision Backbone A convnet for the 2020s

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.648958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.648958Z digest=sha256:b2d8c21b19d6646003e25ea69dc7c49ec75e27096d5c4b8f6c4b82a25f9d3d1b

Observation 08e274b4-2eb3-434d-a2a2-6cec33af6aa0 · outbound

This paper cites Stochastic gradient descent with warm restarts.

Revisiting the Integration of Convolution and Attention for Vision Backbone Stochastic gradient descent with warm restarts

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.710531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.653860Z digest=sha256:6b7f512eede6c3cc60412daf27860f711808538e3307d9a9a08d96e8f68a31bd

Observation 67f1a967-156b-4f1d-9352-ed25ba37e34e · outbound

This paper cites Decoupled weight decay regularization.

Revisiting the Integration of Convolution and Attention for Vision Backbone Decoupled weight decay regularization

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.694264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.658430Z digest=sha256:3ac5136b6df0a22acc4d2bba18588f0acd70961311778a336a403d74b0e9c892

Observation 811dda20-5f93-4a53-bf67-b916f19d64f4 · outbound

This paper cites Understanding the effective receptive field in deep convolutional neural networks.

Revisiting the Integration of Convolution and Attention for Vision Backbone Understanding the effective receptive field in deep convolutional neural networks

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.663125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.663125Z digest=sha256:bc89996189561ee61cb1a49fa503aaa9c5a8e56227028e0bcdef7a899b45e4ba

Observation d26f9cf3-eedb-4ff9-94bf-bedf9a98ed08 · outbound

This paper cites On the integration of self-attention and convolution.

Revisiting the Integration of Convolution and Attention for Vision Backbone On the integration of self-attention and convolution

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.669184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.667715Z digest=sha256:b36fe5e53ebc9f75b9234514108d19277cc131d688cb572b8d84aaedc0795c82

Observation 86f18108-76e6-4d03-89f4-aad15f100b5e · outbound

This paper cites How do vision transformers work? In International Conference on Learning Representations.

Revisiting the Integration of Convolution and Attention for Vision Backbone How do vision transformers work? In International Conference on Learning Representations

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.653251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.673152Z digest=sha256:f029730b44ab082de6ddf8a8f64ee4d9e3e95a17faee82b70b7a76ca73d9ba3b

Observation 8c52ff4f-e782-485d-8c4a-7917b806d9f7 · outbound

This paper cites From Sparse to Soft Mixtures of Experts.

Revisiting the Integration of Convolution and Attention for Vision Backbone From Sparse to Soft Mixtures of Experts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.677701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.677701Z digest=sha256:7265119ceb85f2178db98a739905b5ed75cdc20422e543d8c1e3f37c4327a5a5

Observation 7d2d9425-0051-4a0c-9647-f8003afb9af8 · outbound

This paper cites Sg-former: Self-guided transformer with evolving token reallocation.

Revisiting the Integration of Convolution and Attention for Vision Backbone Sg-former: Self-guided transformer with evolving token reallocation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.637641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.683934Z digest=sha256:aff94c2153e8a049a0213b2e895dbb7d9f828f839099dbd53966be774f298d47

Observation 8f447da8-b00f-4745-a3d5-37c81f920de4 · outbound

This paper cites GLU Variants Improve Transformer.

Revisiting the Integration of Convolution and Attention for Vision Backbone GLU Variants Improve Transformer

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.688586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.688586Z digest=sha256:c4f50fa28963e684be8932a710cf740cf82a025d08c244c6cbbbbf8c8d7eed69

Observation c957a076-e7b5-462b-bc6a-a893ef048b2e · outbound

This paper cites Rethinking the inception architecture for computer vision.

Revisiting the Integration of Convolution and Attention for Vision Backbone Rethinking the inception architecture for computer vision

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.620814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.693643Z digest=sha256:c51e6d798da1a5137d2e1f098529251586182f1f680530abda6f8f8398c12baa

Observation b8d1b302-4d1a-4165-9f47-6bae7e6a7358 · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

Revisiting the Integration of Convolution and Attention for Vision Backbone Training data-efficient image transformers & distillation through attention

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.699611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.699611Z digest=sha256:e64a082513d01b4b16343815d13d7f00113f2eaafe68eb385c8276ee69a41dae

Observation b3f5029a-4064-4d07-bc5e-a06e52549b7c · outbound

This paper cites Maxvit: Multi-axis vision transformer.

Revisiting the Integration of Convolution and Attention for Vision Backbone Maxvit: Multi-axis vision transformer

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.595037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.704544Z digest=sha256:f77fccf55bc4dfde9f82c4440a4f2bdee134174d5895bc8d8e66673de0fd6ec3

Observation 097efe03-527f-4a90-80af-8915d55197d6 · outbound

This paper cites Attention is all you need.

Revisiting the Integration of Convolution and Attention for Vision Backbone Attention is all you need

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.709136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.709136Z digest=sha256:68ba3cd8eaa4396db4c843f496f427e51a73e31e82eee4039661a634ef6aa231

Observation 25f076be-721c-4358-a2c3-630252303d85 · outbound

This paper cites Fast transformers with clustered attention.

Revisiting the Integration of Convolution and Attention for Vision Backbone Fast transformers with clustered attention

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.569689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.713791Z digest=sha256:20d99d7ef4d53a3dfea10a7a8d8ba386fd008501b55676c31c41e46061630b0a

Observation 671dc6b3-bceb-4adb-ba08-247306c7051a · outbound

This paper cites Linformer: Self-Attention with Linear Complexity.

Revisiting the Integration of Convolution and Attention for Vision Backbone Linformer: Self-Attention with Linear Complexity

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.718558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.718558Z digest=sha256:dd18f556823a4241ea06284a29f59d5271ca0c1d3721f6f1f867a560f9548898

Observation 3f160672-9376-4a07-8645-75a305ce7584 · outbound

This paper cites Pyramid vision transformer: A versatile backbone for dense prediction without convolutions.

Revisiting the Integration of Convolution and Attention for Vision Backbone Pyramid vision transformer: A versatile backbone for dense prediction without convolutions

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.554280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.723620Z digest=sha256:571dda7dfc61c4f5b6f1876dc4111493a1871a36fa816e527da2523de8a47058

Observation 46981908-6003-4abf-aba5-5d9a28dbf5c7 · outbound

This paper cites Internimage: Exploring large-scale vision foundation models with deformable convolutions.

Revisiting the Integration of Convolution and Attention for Vision Backbone Internimage: Exploring large-scale vision foundation models with deformable convolutions

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.728316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.728316Z digest=sha256:7eb1a8b54ccd971496f1dde4612dcee7e9e0bb89e37165a710af2eec8353a7c7

Observation 37546beb-fec2-4028-8a5a-67d65e1bb292 · outbound

This paper cites Crossformer: A versatile vision transformer hinging on cross-scale attention.

Revisiting the Integration of Convolution and Attention for Vision Backbone Crossformer: A versatile vision transformer hinging on cross-scale attention

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.526889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.732770Z digest=sha256:725c7ef63bc7345c42f9613ba0082415429c779dd9a5b1330dde1b576a33713e

Observation e25f9852-2ff8-4577-b39d-c1a3fd08fe2a · outbound

This paper cites Pytorch image models.

Revisiting the Integration of Convolution and Attention for Vision Backbone Pytorch image models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.737393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.737393Z digest=sha256:4c4118bc8fcb88f9cb693133757d296470f30d12b45311d4297d06370fee08ff

Observation 2363af27-ee8c-462c-b387-5b23004cf27b · outbound

This paper cites Cvt: Introducing convolutions to vision transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone Cvt: Introducing convolutions to vision transformers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.742139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.742139Z digest=sha256:7791b972226fe91e4d050df8553d067c6f4b2048fff56c9cc36a2a4624c713d2

Observation 72b7e3a2-3828-4fb0-8c1f-9b6fac3b54cd · outbound

This paper cites Vision transformer with deformable attention.

Revisiting the Integration of Convolution and Attention for Vision Backbone Vision transformer with deformable attention

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.490534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.746602Z digest=sha256:9a63a7720034d216001cd8cb1b62b1a8474b1e8ac3ad6c00b3ed136fe5acfcec

Observation 432a43ca-7941-4232-8daa-c5a5f9d37b39 · outbound

This paper cites Unified perceptual parsing for scene understanding.

Revisiting the Integration of Convolution and Attention for Vision Backbone Unified perceptual parsing for scene understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.471885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.751192Z digest=sha256:cb441617e8edc2117e1eed810cc0a08bc4835985316f4d55568173ffcbdd9d84

Observation c57ac0cd-f838-4904-b964-df1166874530 · outbound

This paper cites ClusTR: Exploring Efficient Self-attention via Clustering for Vision Transformers.

Revisiting the Integration of Convolution and Attention for Vision Backbone ClusTR: Exploring Efficient Self-attention via Clustering for Vision Transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.755707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.755707Z digest=sha256:0f2302a559c837db855de901b595a9342b58b34b4f73b16a36afd30662ada430

Observation 6273dae9-623d-47b8-a650-862cb13558df · outbound

This paper cites Moat: Alternating mobile convolution and attention brings strong vision models.

Revisiting the Integration of Convolution and Attention for Vision Backbone Moat: Alternating mobile convolution and attention brings strong vision models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.456508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.760844Z digest=sha256:1d8d56cd3e53a59bb34e3d5c228bedb769476768c5f919b62975001b48c7d7a0

Observation 127dccab-3294-49da-a40b-af70cca364a3 · outbound

This paper cites Scalablevit: Rethinking the context-oriented generalization of vision transformer.

Revisiting the Integration of Convolution and Attention for Vision Backbone Scalablevit: Rethinking the context-oriented generalization of vision transformer

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.440771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.765408Z digest=sha256:c8ebad297b5c0712d5afd67b0bdc353966efcd4f0cf5129db6a7625cdafc51d6

Observation 781f0422-c2e3-43c2-9731-2aadaa43b699 · outbound

This paper cites Wave-vit: Unifying wavelet and transformers for visual representation learning.

Revisiting the Integration of Convolution and Attention for Vision Backbone Wave-vit: Unifying wavelet and transformers for visual representation learning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.423595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.769944Z digest=sha256:66dfec4117abae47fa4f57d9057fe723f1cc5fef900ffb0951c954b2a0552e34

Observation c6cc6a64-fbd2-430b-b27f-7529164bee92 · outbound

This paper cites Dual vision transformer.

Revisiting the Integration of Convolution and Attention for Vision Backbone Dual vision transformer

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.408500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.774538Z digest=sha256:0ba5aef77c0daee9018ca4f23afe14a1cc2b5e7304eaf19fae9706cc151a163f

Observation 3104cb3c-de13-4103-9518-77fe4e6b2634 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Revisiting the Integration of Convolution and Attention for Vision Backbone mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.778999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.778999Z digest=sha256:aa2e38fe673d957a8313d0088d35abd65a7a3e4c072d5ae6a2b8a1fa2e1c2124

Observation f37b6ea9-5c22-4d9f-9f7c-e9eb59b32ebd · outbound

This paper cites Metaformer is actually what you need for vision.

Revisiting the Integration of Convolution and Attention for Vision Backbone Metaformer is actually what you need for vision

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.783904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.783904Z digest=sha256:93e5808c1675b319a45f102dade970412b0bc95af9fe23b12acf1dd2158205ad

Observation 96127753-1662-4284-bb34-aa11a01db413 · outbound

This paper cites Cutmix: Regularization strategy to train strong classifiers with localizable features.

Revisiting the Integration of Convolution and Attention for Vision Backbone Cutmix: Regularization strategy to train strong classifiers with localizable features

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.789468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.789468Z digest=sha256:6350db24d6b984204d6eb34d1839ec35cc4516eec498be8c2d17f01fcd48aba2

Observation d8f1911c-4a38-40ff-b5e0-44c2b85ef318 · outbound

This paper cites Not all tokens are equal: Human-centric visual analysis via token clustering transformer.

Revisiting the Integration of Convolution and Attention for Vision Backbone Not all tokens are equal: Human-centric visual analysis via token clustering transformer

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.373093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.794116Z digest=sha256:4c5724f165333133e934748b986b8545c3e5b13c565b4541d4c7caeb2835c9e0

Observation f871d40c-118c-4d34-a125-1b3cb92efa2a · outbound

This paper cites Dauphin, and David Lopez-Paz.

Revisiting the Integration of Convolution and Attention for Vision Backbone Dauphin, and David Lopez-Paz

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.357107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.799921Z digest=sha256:3486736e62aa338ef31804e5bebcf132d02db7f27d2d682b5e0acaa8b913117c

Observation 99390d75-5c34-4505-9944-84d99160ca84 · outbound

This paper cites Dense distinct query for end-to-end object detection.

Revisiting the Integration of Convolution and Attention for Vision Backbone Dense distinct query for end-to-end object detection

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.342338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.804553Z digest=sha256:7f898651266f74273f58c3371f4067c535a56e252729081f6e2d294d93095e83

Observation 101fe6df-b352-47df-a1b5-995a62413be6 · outbound

This paper cites Semantic understanding of scenes through the ade20k dataset.

Revisiting the Integration of Convolution and Attention for Vision Backbone Semantic understanding of scenes through the ade20k dataset

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T15:15:22.809225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:15:22.809225Z digest=sha256:d11e75dc42d62e71f40b9aefd7276b6592d8480645107fee5a80d0205db1e264

Observation 177a1496-b5a7-426c-a01f-b677ad510484 · outbound

This paper cites Biformer: Vision transformer with bi-level routing attention.

Revisiting the Integration of Convolution and Attention for Vision Backbone Biformer: Vision transformer with bi-level routing attention

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.316422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.813803Z digest=sha256:3e2bb8c1d1abdaf6e1809450091713407e341202a3c48ac4dc981c8176fb0298

Observation 47e9e79a-7146-40d9-968c-9b10fe497bb2 · outbound

This paper cites The contributions are summarized at the end of the Introduction (Sec.

Revisiting the Integration of Convolution and Attention for Vision Backbone The contributions are summarized at the end of the Introduction (Sec

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.301750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.818678Z digest=sha256:bf6791a7f767d167511710d08e826d39433aeac01ea565f61e9bfab6e836b4f7

Observation 61e3f61e-ef81-41b8-bad8-74c9db0db45c · outbound

This paper cites Limitations.

Revisiting the Integration of Convolution and Attention for Vision Backbone Limitations

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.286061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.823780Z digest=sha256:8cfd95071ce7268a9e9f2dc9de55d2f7000295a905ca036635a54795259e22c9

Observation 64ea8b7e-7c8b-4a03-ba5f-f73450c340ae · outbound

This paper cites Guidelines: • The answer NA means that the paper does not include theoretical results.

Revisiting the Integration of Convolution and Attention for Vision Backbone Guidelines: • The answer NA means that the paper does not include theoretical results

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.271088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.828772Z digest=sha256:0d9ca57fab249c26942e4b1da6781ebc0f26e2c41931b43dde514b1fbdc4f92d

Observation 54e4fca1-90b6-4f1d-9cd1-bb07ad6ab8e5 · outbound

This paper cites an unresolved cited work.

Revisiting the Integration of Convolution and Attention for Vision Backbone Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:15:23.255235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.833652Z digest=sha256:76cdb71e0f4768e7e28460ae2fda86cc62ae17ae368915922f6831ddf6b5459c

Observation 6093176e-68bb-4a1c-9bb9-65cdeda95d99 · outbound

This paper cites We will release the code upon the acceptance of this paper.

Revisiting the Integration of Convolution and Attention for Vision Backbone We will release the code upon the acceptance of this paper

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.239351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.838807Z digest=sha256:90d7c8681bc069274d46cd9c2187a91f468603e975562d0b57590c3381bee396

Observation c1aad93e-ca16-4315-a9a3-41f42c0312e3 · outbound

This paper cites an unresolved cited work.

Revisiting the Integration of Convolution and Attention for Vision Backbone Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:15:23.223830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.843504Z digest=sha256:47c310ee0c323f48701660dcccbac9858642161c3538276c8093c14195cd1de1

Observation 39acc319-7392-46c0-b341-9a35519da408 · outbound

This paper cites The experimental results are not sensitive to random initialization.

Revisiting the Integration of Convolution and Attention for Vision Backbone The experimental results are not sensitive to random initialization

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.208439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.849169Z digest=sha256:e5c2e4f6370b9cf27575f192341482b6e8a6f646417e4ed1138ccd6965eed3dc

Observation ba19d142-b8d1-42a2-b77a-88c08d78cd56 · outbound

This paper cites an unresolved cited work.

Revisiting the Integration of Convolution and Attention for Vision Backbone Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-12T15:15:23.193193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.853815Z digest=sha256:59edb5cb8235ae917008cf41834337b29403967f7113267ad1bcfdde9b5f0551

Observation 5b5c2858-117e-415c-ba42-37ba7f8a4875 · outbound

This paper cites We only use existing and publicly available datasets for evaluations.

Revisiting the Integration of Convolution and Attention for Vision Backbone We only use existing and publicly available datasets for evaluations

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.177914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.858221Z digest=sha256:dd57630904d0a8be5ad55a8ca91c02c9b066071ca7bf54c3aa79ef499faa29b6

Observation 5174d3b6-a2ba-4438-bae1-c4cd00b5589f · outbound

This paper cites It is too broad to discuss the societal impacts of such a general topic.

Revisiting the Integration of Convolution and Attention for Vision Backbone It is too broad to discuss the societal impacts of such a general topic

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.162268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.863391Z digest=sha256:9d5e6538368d085ad2d35864bb0ad84e47a06cf63b5ef09dfb63f20f8639fbc0

Observation accd94e5-e3eb-474d-9757-15c9fb2dfc09 · outbound

This paper cites It poses no such risks.

Revisiting the Integration of Convolution and Attention for Vision Backbone It poses no such risks

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.145084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.868619Z digest=sha256:f09e976abd8aceb33049e981c26527e12e7757928a33e7a8be8d434eb54036c9

Observation 5df8bdaf-1c5c-416d-98ae-f241e660f4b4 · outbound

This paper cites These works are properly cited.

Revisiting the Integration of Convolution and Attention for Vision Backbone These works are properly cited

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.129743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.873273Z digest=sha256:bde866337027069889463e099de2448a0d187fe156d26ffec0495ae79fae82e0

Observation fc9cc2ca-ae0b-41e7-bfe3-79412e961492 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not release new assets.

Revisiting the Integration of Convolution and Attention for Vision Backbone Guidelines: • The answer NA means that the paper does not release new assets

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.114445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.878078Z digest=sha256:a43500058f10f1d10c93aa67a7ba21d50df95744bca9595b2f08af978ee52191

Observation af19c8cc-0c12-4688-bcdc-09c8d8769274 · outbound

This paper cites Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Revisiting the Integration of Convolution and Attention for Vision Backbone Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.098801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.883149Z digest=sha256:035cb19ce3b8891c97ddfddfa39330c83b10fc2177d3fb1d74fea9ad89209197

Observation 55fae290-e9b2-42a3-a17b-34eb33397eda · outbound

This paper cites Guidelines: 23 • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects.

Revisiting the Integration of Convolution and Attention for Vision Backbone Guidelines: 23 • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:15:23.082865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-12T15:15:22.887919Z digest=sha256:3694c359c8d8df1a09f24ccbeb05d7f66186c84d2de8c93edae874d940d707c8

Pith citing papers

No inbound Pith citation observations are available.