On 4xH100 nodes, FP16, pin_memory, and NHWC/DALI speed up image recognition, while LoRA is faster than DPO and QLoRA for LLM tuning, and PyTorch DataLoader loses scaling beyond 2 GPUs.
Attention-based Image Upsampling
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Convolutional layers are an integral part of many deep neural network solutions in computer vision. Recent work shows that replacing the standard convolution operation with mechanisms based on self-attention leads to improved performance on image classification and object detection tasks. In this work, we show how attention mechanisms can be used to replace another canonical operation: strided transposed convolution. We term our novel attention-based operation attention-based upsampling since it increases/upsamples the spatial dimensions of the feature maps. Through experiments on single image super-resolution and joint-image upsampling tasks, we show that attention-based upsampling consistently outperforms traditional upsampling methods based on strided transposed convolution or based on adaptive filters while using fewer parameters. We show that the inherent flexibility of the attention mechanism, which allows it to use separate sources for calculating the attention coefficients and the attention targets, makes attention-based upsampling a natural choice when fusing information from multiple image modalities.
citation-role summary
citation-polarity summary
fields
cs.DC 1years
2025 1verdicts
CONDITIONAL 1roles
background 1polarities
unclear 1representative citing papers
citing papers explorer
-
Profiling and optimization of multi-card GPU machine learning jobs
On 4xH100 nodes, FP16, pin_memory, and NHWC/DALI speed up image recognition, while LoRA is faster than DPO and QLoRA for LLM tuning, and PyTorch DataLoader loses scaling beyond 2 GPUs.