Pith. sign in

REVIEW 2 cited by

Efficient Self-Ensemble for Semantic Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2111.13280 v2 pith:7AZKCUHE submitted 2021-11-26 cs.CV

classification cs.CV
keywords ensemblesegmentationsemanticself-ensembleapproachcreatingheavymethods
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Ensemble of predictions is known to perform better than individual predictions taken separately. However, for tasks that require heavy computational resources, e.g. semantic segmentation, creating an ensemble of learners that needs to be trained separately is hardly tractable. In this work, we propose to leverage the performance boost offered by ensemble methods to enhance the semantic segmentation, while avoiding the traditional heavy training cost of the ensemble. Our self-ensemble approach takes advantage of the multi-scale features set produced by feature pyramid network methods to feed independent decoders, thus creating an ensemble within a single model. Similar to the ensemble, the final prediction is the aggregation of the prediction made by each learner. In contrast to previous works, our model can be trained end-to-end, alleviating the traditional cumbersome multi-stage training of ensembles. Our self-ensemble approach outperforms the current state-of-the-art on the benchmark datasets Pascal Context and COCO-Stuff-10K for semantic segmentation and is competitive on ADE20K and Cityscapes. Code is publicly available at github.com/WalBouss/SenFormer.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. EDTformer: An Efficient Decoder Transformer for Visual Place Recognition

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A decoder transformer with learnable queries plus a low-rank parallel adapter for frozen DINOv2 achieves state-of-the-art visual place recognition on multiple benchmarks with reduced training memory.

  2. Chanel-Orderer: A Channel-Ordering Predictor for Tri-Channel Natural Images

    cs.CV 2024-11 reject novelty 5.0 of 10

    Chanel-Orderer uses semantic-mask-weighted channel scores to predict the correct R/G/B ordering of a permuted tri-channel image, reporting up to 98.5% accuracy on SiftFlow.

Pith tools