Pith. sign in

REVIEW 1 cited by

Colorization Transformer

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.04432 v2 pith:33SN762G submitted 2021-02-08 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords colorizationtransformerimagegrayscaleresolutioncoarsecoloredcolorings
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present the Colorization Transformer, a novel approach for diverse high fidelity image colorization based on self-attention. Given a grayscale image, the colorization proceeds in three steps. We first use a conditional autoregressive transformer to produce a low resolution coarse coloring of the grayscale image. Our architecture adopts conditional transformer layers to effectively condition grayscale input. Two subsequent fully parallel networks upsample the coarse colored low resolution image into a finely colored high resolution image. Sampling from the Colorization Transformer produces diverse colorings whose fidelity outperforms the previous state-of-the-art on colorising ImageNet based on FID results and based on a human evaluation in a Mechanical Turk test. Remarkably, in more than 60% of cases human evaluators prefer the highest rated among three generated colorings over the ground truth. The code and pre-trained checkpoints for Colorization Transformer are publicly available at https://github.com/google-research/google-research/tree/master/coltran

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Enhancing Video Understanding: Deep Neural Networks for Spatiotemporal Analysis

    cs.CV 2025-02 unverdicted

    A narrative review of spatiotemporal deep neural networks for video understanding, with tables of benchmark datasets and reported model results.

Pith tools