Pith. sign in

REVIEW 9 cited by

GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.04607 v8 pith:JNDCAL3K submitted 2023-06-07 cs.CV cs.AI

classification cs.CVcs.AI
keywords geometricconditionsdatadiffusiongenerationgeodiffusionmodelsdetection
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models have attracted significant attention due to the remarkable ability to create content and generate data for tasks like image classification. However, the usage of diffusion models to generate the high-quality object detection data remains an underexplored area, where not only image-level perceptual quality but also geometric conditions such as bounding boxes and camera views are essential. Previous studies have utilized either copy-paste synthesis or layout-to-image (L2I) generation with specifically designed modules to encode the semantic layouts. In this paper, we propose the GeoDiffusion, a simple framework that can flexibly translate various geometric conditions into text prompts and empower pre-trained text-to-image (T2I) diffusion models for high-quality detection data generation. Unlike previous L2I methods, our GeoDiffusion is able to encode not only the bounding boxes but also extra geometric conditions such as camera views in self-driving scenes. Extensive experiments demonstrate GeoDiffusion outperforms previous L2I methods while maintaining 4x training time faster. To the best of our knowledge, this is the first work to adopt diffusion models for layout-to-image generation with geometric conditions and demonstrate that L2I-generated images can be beneficial for improving the performance of object detectors.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. To Blend In, First Decouple: Rethinking Camouflage Image Generation via Context-Decoupled Representations

    cs.CV 2026-07 conditional novelty 6.0 of 10

    CamoDreamer generates camouflage images by decoupling foreground and background control in a diffusion model, reporting a 15.5-point FID gain over prior state of the art on LAKE-RED.

  2. Towards Continual Expansion of Data Coverage: Automatic Text-guided Edge-case Synthesis

    cs.CV 2025-09 unverdicted novelty 6.0 of 10

    Automated LLM-based prompt engineering for text-to-image edge-case synthesis improves object detection robustness on the FishEye8K benchmark over naive augmentation and manual prompts.

  3. FICGen: Frequency-Inspired Contextual Disentanglement for Layout-driven Degraded Image Generation

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A frequency-guided layout-to-image generation framework, FICGen, improves fidelity, layout alignment, and detector trainability on degraded scenes across five benchmarks.

  4. Detail++: Training-Free Detail Enhancer for T2I Diffusion Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Detail++ uses progressive multi-branch prompt injection and test-time attention optimization to improve attribute binding in text-to-image generation.

  5. Controllable Video Object Insertion via Multiview Priors

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    A multi-view prior-based framework for video object insertion that uses dual-path conditioning and an integration-aware consistency module to improve appearance stability and occlusion handling.

  6. Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method

    cs.CV 2025-10 conditional novelty 5.0 of 10

    UniScenev2 scales occupancy-centric driving-scene generation to NuPlan scale, releasing a 3.6M-frame semantic-occupancy dataset and jointly generating occupancy, video, and LiDAR that beats published baselines on its ...

  7. FLORA: Efficient Synthetic Data Generation for Object Detection in Low-Data Regimes via finetuning Flux LoRA

    cs.CV 2025-08 conditional novelty 5.0 of 10

    A LoRA fine-tuned Flux inpainting pipeline generates synthetic object detection images that outperform ODGEN's 10x larger synthetic set in downstream mAP.

  8. Edge-case Synthesis for Fisheye Object Detection: A Data-centric Perspective

    cs.CV 2025-07 reject novelty 5.0 of 10

    Edge-case synthesis with a fine-tuned text-to-image model improves fisheye object detection, but the gain is not isolated from simply adding more data.

  9. ECCV 2024 W-CODA: 1st Workshop on Multimodal Perception and Comprehension of Corner Cases in Autonomous Driving

    cs.CV 2025-07 unverdicted novelty 1.0 of 10

    A workshop report documenting the ECCV 2024 W-CODA event, its accepted papers, speakers, and the dual-track corner case understanding and generation challenge.

Pith tools