Pith. sign in

REVIEW 4 cited by

Multimodal Conditional Information Bottleneck for Generalizable AI-Generated Image Detection

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.15217 v1 pith:DCHVDZZ6 submitted 2025-05-21 cs.CV

classification cs.CV
keywords featuresbottleneckimageinformationtextai-generatedconditionalfeature
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Although existing CLIP-based methods for detecting AI-generated images have achieved promising results, they are still limited by severe feature redundancy, which hinders their generalization ability. To address this issue, incorporating an information bottleneck network into the task presents a straightforward solution. However, relying solely on image-corresponding prompts results in suboptimal performance due to the inherent diversity of prompts. In this paper, we propose a multimodal conditional bottleneck network to reduce feature redundancy while enhancing the discriminative power of features extracted by CLIP, thereby improving the model's generalization ability. We begin with a semantic analysis experiment, where we observe that arbitrary text features exhibit lower cosine similarity with real image features than with fake image features in the CLIP feature space, a phenomenon we refer to as "bias". Therefore, we introduce InfoFD, a text-guided AI-generated image detection framework. InfoFD consists of two key components: the Text-Guided Conditional Information Bottleneck (TGCIB) and Dynamic Text Orthogonalization (DTO). TGCIB improves the generalizability of learned representations by conditioning on both text and class modalities. DTO dynamically updates weighted text features, preserving semantic information while leveraging the global "bias". Our model achieves exceptional generalization performance on the GenImage dataset and latest generative models. Our code is available at https://github.com/Ant0ny44/InfoFD.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. VGIF-Score: Interpretable and Diagnostic Evaluation of Spatio-Temporal Instruction Following in Video Generation

    cs.CV 2026-07 conditional novelty 7.0 of 10

    VGIF-Score decomposes video prompts into dependency graphs and uses a VLM to diagnose which instruction constraints models satisfy, revealing strong failures on causal and late-prompt constraints.

  2. CoDA: Color Distribution Probing for Efficient and Generalizable AI-Generated Image Detection

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    CoDA is a lightweight detector using a Noise-Quantization Probe on color non-uniformity that reports strong cross-domain results on the new FakeForm benchmark and competitive cross-model performance on standard tests.

  3. IncreFA: Breaking the Static Wall of Generative Model Attribution

    cs.CV 2026-04 unverdicted novelty 6.0 of 10

    IncreFA uses hierarchical constraints with learnable orthogonal priors and a latent memory bank to enable continual adaptation for attributing images to new generative models, reporting SOTA accuracy and 98.93% unseen...

  4. FlexiGrad: Adaptive Gradient Modulation for Hierarchical Fine-Grained Classification

    cs.CV 2026-07 conditional novelty 4.0 of 10

    FlexiGrad selectively removes conflicting and reinforces agreeing gradient components between hierarchy levels, improving multi-granularity accuracy on CUB, FGVC-Aircraft and Stanford Cars.

Pith tools