Pith. sign in

REVIEW 2 cited by

PIM: Video Coding using Perceptual Importance Maps

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.10674 v2 pith:SYNF2OTD submitted 2022-12-20 eess.IV

classification eess.IV
keywords importancevideosmapsperceptualqualityspatio-temporalvideodataset
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
abstract

Human perception is at the core of lossy video compression, with numerous approaches developed for perceptual quality assessment and improvement over the past two decades. In the determination of perceptual quality, different spatio-temporal regions of the video differ in their relative importance to the human viewer. However, since it is challenging to infer or even collect such fine-grained information, it is often not used during compression beyond low-level heuristics. We present a framework which facilitates research into fine-grained subjective importance in compressed videos, which we then utilize to improve the rate-distortion performance of an existing video codec (x264). The contributions of this work are threefold: (1) we introduce a web-tool which allows scalable collection of fine-grained perceptual importance, by having users interactively paint spatio-temporal maps over encoded videos; (2) we use this tool to collect a dataset with 178 videos with a total of 14443 frames of human annotated spatio-temporal importance maps over the videos; and (3) we use our curated dataset to train a lightweight machine learning model which can predict these spatio-temporal importance regions. We demonstrate via a subjective study that encoding the videos in our dataset while taking into account the importance maps leads to higher perceptual quality at the same bitrate, with the videos encoded with importance maps preferred $1.8 \times$ over the baseline videos. Similarly, we show that for the 18 videos in test set, the importance maps predicted by our model lead to higher perceptual quality videos, $2 \times$ preferred over the baseline at the same bitrate.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Rethink Before You Execute: Adaptive Execution for World Action Models

    cs.RO 2026-08 reject novelty 5.0 of 10

    TempoWAM adapts the replanning frequency of world action models based on an online estimate of task progress, reducing inference calls on easy tasks and improving success on hard tasks.

  2. Emerging Advances in Learned Video Compression: Models, Systems and Beyond

    eess.IV 2025-04 conditional novelty 3.0 of 10

    A survey of end-to-end learned video compression that categorizes P-frame and B-frame neural codecs, reviews optimization and system implementation, and benchmarks several learned codecs against standard codecs.

Pith tools