Pith. sign in

REVIEW 3 cited by

A Systematic Performance Analysis of Deep Perceptual Loss Networks: Breaking Transfer Learning Conventions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.04032 v3 pith:77FWAE5B submitted 2023-02-08 cs.CV cs.LG

A Systematic Performance Analysis of Deep Perceptual Loss Networks: Breaking Transfer Learning Conventions

classification cs.CV cs.LG
keywords lossdeepnetworksdifferentextractionfeatureperceptualperformance
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
Share X Bluesky LinkedIn Reddit HN
read the original abstract

In recent years, deep perceptual loss has been widely and successfully used to train machine learning models for many computer vision tasks, including image synthesis, segmentation, and autoencoding. Deep perceptual loss is a type of loss function for images that computes the error between two images as the distance between deep features extracted from a neural network. Most applications of the loss use pretrained networks called loss networks for deep feature extraction. However, despite increasingly widespread use, the effects of loss network implementation on the trained models have not been studied. This work rectifies this through a systematic evaluation of the effect of different pretrained loss networks on four different application areas. Specifically, the work evaluates 14 different pretrained architectures with four different feature extraction layers. The evaluation reveals that VGG networks without batch normalization have the best performance and that the choice of feature extraction layer is at least as important as the choice of architecture. The analysis also reveals that deep perceptual loss does not adhere to the transfer learning conventions that better ImageNet accuracy implies better downstream performance and that feature extraction from the later layers provides better performance.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. DRIFT: Deep Restoration, ISP Fusion, and Tone-mapping

    eess.IV 2026-04 unverdicted novelty 5.0

    DRIFT uses a multi-frame network trained with adversarial perceptual loss for alignment, denoising, demosaicing and super-resolution, followed by an efficient deep tone-mapping module that supports tunability and refe...

  2. Can LLMs Reason About Attention? Towards Zero-Shot Analysis of Multimodal Classroom Behavior

    cs.HC 2026-04 unverdicted novelty 4.0

    A pipeline uses OpenPose and Gaze-LLE to extract pose and gaze data from classroom videos, deletes the raw footage, and applies an LLM for zero-shot behavioral analysis of student attention.

  3. Can LLMs Reason About Attention? Towards Zero-Shot Analysis of Multimodal Classroom Behavior

    cs.HC 2026-04 unverdicted novelty 4.0

    A single-GPU pipeline extracts OpenPose skeletons and Gaze-LLE attention, discards frames, and uses QwQ-32B zero-shot to produce classroom attention summaries and heatmaps, with preliminary evidence of limited spatial...