REVIEW 3 cited by
A Systematic Performance Analysis of Deep Perceptual Loss Networks: Breaking Transfer Learning Conventions
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
A Systematic Performance Analysis of Deep Perceptual Loss Networks: Breaking Transfer Learning Conventions
read the original abstract
In recent years, deep perceptual loss has been widely and successfully used to train machine learning models for many computer vision tasks, including image synthesis, segmentation, and autoencoding. Deep perceptual loss is a type of loss function for images that computes the error between two images as the distance between deep features extracted from a neural network. Most applications of the loss use pretrained networks called loss networks for deep feature extraction. However, despite increasingly widespread use, the effects of loss network implementation on the trained models have not been studied. This work rectifies this through a systematic evaluation of the effect of different pretrained loss networks on four different application areas. Specifically, the work evaluates 14 different pretrained architectures with four different feature extraction layers. The evaluation reveals that VGG networks without batch normalization have the best performance and that the choice of feature extraction layer is at least as important as the choice of architecture. The analysis also reveals that deep perceptual loss does not adhere to the transfer learning conventions that better ImageNet accuracy implies better downstream performance and that feature extraction from the later layers provides better performance.
Forward citations
Cited by 3 Pith papers
-
DRIFT: Deep Restoration, ISP Fusion, and Tone-mapping
DRIFT uses a multi-frame network trained with adversarial perceptual loss for alignment, denoising, demosaicing and super-resolution, followed by an efficient deep tone-mapping module that supports tunability and refe...
-
Can LLMs Reason About Attention? Towards Zero-Shot Analysis of Multimodal Classroom Behavior
A pipeline uses OpenPose and Gaze-LLE to extract pose and gaze data from classroom videos, deletes the raw footage, and applies an LLM for zero-shot behavioral analysis of student attention.
-
Can LLMs Reason About Attention? Towards Zero-Shot Analysis of Multimodal Classroom Behavior
A single-GPU pipeline extracts OpenPose skeletons and Gaze-LLE attention, discards frames, and uses QwQ-32B zero-shot to produce classroom attention summaries and heatmaps, with preliminary evidence of limited spatial...
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.