Pith. sign in

REVIEW 9 cited by

Query2Label: A Simple Transformer Way to Multi-Label Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.10834 v1 pith:HYJIGEIC submitted 2021-07-22 cs.CV

Query2Label: A Simple Transformer Way to Multi-Label Classification

classification cs.CV
keywords classificationmulti-labelsimpletransformereffectiveapproachexistencefeatures
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

This paper presents a simple and effective approach to solving the multi-label classification problem. The proposed approach leverages Transformer decoders to query the existence of a class label. The use of Transformer is rooted in the need of extracting local discriminative features adaptively for different labels, which is a strongly desired property due to the existence of multiple objects in one image. The built-in cross-attention module in the Transformer decoder offers an effective way to use label embeddings as queries to probe and pool class-related features from a feature map computed by a vision backbone for subsequent binary classifications. Compared with prior works, the new framework is simple, using standard Transformers and vision backbones, and effective, consistently outperforming all previous works on five multi-label classification data sets, including MS-COCO, PASCAL VOC, NUS-WIDE, and Visual Genome. Particularly, we establish $91.3\%$ mAP on MS-COCO. We hope its compact structure, simple implementation, and superior performance serve as a strong baseline for multi-label classification tasks and future studies. The code will be available soon at https://github.com/SlongLiu/query2labels.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. PHOEBI: An Open-World Benchmark for Bacterial Identification in Phase-Contrast Microscopy

    cs.CV 2026-06 unverdicted novelty 7.0

    PHOEBI is a benchmark dataset and LCO evaluation protocol for open-world multi-label bacterial species identification from phase-contrast microscopy of polymicrobial samples.

  2. ELDOR: A Dataset and Benchmark for Illegal Gold Mining in the Amazon Rainforest

    cs.CV 2026-05 unverdicted novelty 7.0

    Introduces the ELDOR UAV dataset and four benchmark tasks for semantic segmentation and classification of mining disturbances and ecological recovery in rainforest imagery.

  3. SenBen: Sensitive Scene Graphs for Explainable Content Moderation

    cs.CV 2026-04 unverdicted novelty 7.0

    SenBen is the first large-scale scene graph benchmark for sensitive content, paired with a 241M distilled model that outperforms most VLMs and safety APIs on grounded detection while running much faster.

  4. Disentangled Fine-Grained Prototype Learning for Incomplete Image-Tabular Classification

    cs.CV 2026-06 unverdicted novelty 6.0

    DFPL introduces prototype-based disentanglement and alignment modules to preserve fine-grained consistency across heterogeneous modalities for better performance under missing data conditions.

  5. CXR-LT 2026 Challenge: Multi-Center Long-Tailed and Zero Shot Chest X-ray Classification

    cs.CV 2026-04 accept novelty 6.0

    CXR-LT 2026 introduces a radiologist-annotated multi-center dataset of 145k+ CXRs to benchmark robust multi-label classification on known classes and open-world generalization to unseen rare diseases.

  6. SenBen: Sensitive Scene Graphs for Explainable Content Moderation

    cs.CV 2026-04 conditional novelty 6.0

    A 241M multi-task student trained with suffix identity, VAR loss, and a decoupled Q2L head matches or beats most VLMs and safety APIs on grounded sensitive scene graphs at 7.6× lower latency.

  7. FISHER: Gradient-Decoupled Hierarchical Multi-Task Learning for Fine-Grained Aquatic Species Recognition

    q-bio.QM 2026-07 conditional novelty 5.0

    FISHER improves fine-grained fish recognition by detaching gradients between hierarchical tasks, raising ultra-rare species accuracy from 50.4% to 63.8% on Fish-Vista.

  8. Multimodal Group Emotion Recognition In-the-Wild Towards a Privacy-Safe Non-Individual Approach

    cs.CV 2026-05 unverdicted novelty 4.0

    Proposes cross-attention audio-video fusion and VE-MD latent-space models for group emotion recognition that avoid individual cues and report competitive performance via ablation studies on synthetic and real data.

  9. Intuitive Surgical SurgToolLoc and SurgVU Challenges Results: 2022-2025

    cs.CV 2023-05 unverdicted novelty 2.0

    The paper summarizes results from the SurgToolLoc and SurgVU challenges held at MICCAI conferences from 2022 to 2025.