Pith. sign in

REVIEW 4 major objections 4 minor 113 references

Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding

T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read POBF repaints only the background around a target object, then filters synthetic samples with teacher scores, yielding an average 5.83% gain over real-data-only visual grounding training in data-scarce settings.

desk verdict A useful synthetic-data recipe for data-scarce visual grounding, with a plausible label-alignment story that needs verification. read the letter →

arxiv 2412.00684 v2 pith:I3QVJPTU submitted 2024-12-01 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords visualgroundingdata-scarcelearningsynthetictrainingdataimageinpaintingfilteringpruningreferringexpressioncomprehensiondiffusionmodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper claims that visual grounding can be learned from very little real data if you generate extra training images by repainting everything outside the target bounding box and then keep only the synthetic samples a teacher model judges useful. The proposed framework, POBF, keeps the object inside the box untouched, so the original text-box annotation stays aligned, and selects one of several generated variants per real sample using a score that combines how easy the sample is, whether the background could cause overfitting, and a penalty for overly easy or uninformative backgrounds. Across four benchmark datasets at 1% of real training data, POBF reports an average gain of 5.83% over real-data-only training and margins of 2.29% to 3.85% over prior synthetic-data methods. The authors also show the gains persist across different generators, captioners, data sizes, and grounding architectures.

What carries the argument

The load-bearing object is the paint-outside-the-box inpainting step: an off-the-shelf diffusion inpainting model regenerates everything outside the bounding box at high strength while the pixels inside the box are copied unchanged, so the synthetic image is guaranteed to align with the box by construction. On top of that sits a three-term selection score computed by a teacher network trained only on the scarce real data. The hardness score $S_1$ is the teacher's IoU with the true box given the text; the overfitting score $S_2 = 1 - \mathrm{IoU}(\mathrm{teacher}(\mathrm{masked\ image}, \mathrm{text}), \mathrm{box})$ flags backgrounds that leak the answer; and the penalty term $P$ is the IoU the teacher gets from the image alone with an empty text prompt. Together they rank the generated variants and one winner per real sample enters the student's training set.

What would settle it

Compare POBF's selected synthetic set against a version in which each generated image's box content is validated by an independent object detector or human annotator; if removing samples with corrupted box contents does not lower accuracy or the gain disappears, the assumed label alignment is not what drives the result.

Watch

Extended reading notes

Core claim

The central claim is that the label misalignment that plagues synthetic visual grounding data comes from editing the object rather than the context, and that painting outside the box fixes it. Starting from a real image, a real text query, and its bounding box, the method captions the image, inpaints the region outside the box with high strength so the background changes substantially, and pairs the new image with the original box and text. Because the box pixels are never regenerated, the synthetic sample inherits the exact ground-truth localization. A teacher trained on the scarce real data then scores each generated image: $S_1$ measures how confidently the teacher localizes the target with the text, $S_2$ measures whether the teacher can still localize it when the box is masked (a sign the background carries unintended shortcut features), and $P$ measures localization from the image alone with an empty text prompt. The weighted sum of the three normalized scores, tuned by grid search, picks one synthetic image per real sample, and the student trains on real plus selected synthetic data, sometimes with regenerated captions.

Load-bearing premise

The approach assumes that inpainting the background leaves the object inside the bounding box visually unchanged and semantically aligned with the original text; if the generator alters the object or introduces overlapping artifacts, the synthetic labels are wrong even after filtering.

Editorial extensions

If this is right

  • If the claim holds, dense region-text annotations are not a hard requirement for visual grounding: a small real set plus repainted backgrounds and teacher-filtered selection can lift accuracy by 5.83% on average.
  • Because the box pixels are never regenerated, label misalignment is avoided by construction, which is the paper's diagnosis of why prior object-editing and text-to-image baselines underperform.
  • The filter's three scores beat the common CLIP-similarity selection rule by an average of 1.32%, suggesting that hardness and overfitting capture complementary quality signals.
  • The gains are not tied to one generator or architecture: the paper reports consistent improvements across two alternative image generators, two alternative captioners, three grounding models, and three data-scarce budgets at 0.5%, 1%, and 2% of real data.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same paint-outside-the-box rule should transfer to other region-labeled tasks such as instance segmentation, detection, and keypoint localization: preserve the annotated region and regenerate its context, then filter by a teacher; this extension is not tested in the paper.
  • The per-sample 'keep one winner' rule limits synthetic expansion to 2x; a diversity-aware selection that keeps multiple high-scoring variants could use the same scores to get more benefit from a fixed generator budget.
  • Because the teacher is trained on the same tiny real set, the filter inherits the teacher's blind spots; averaging scores across multiple teacher seeds or architectures could make the selection more stable, but the paper's teacher set is single.
  • The paper's own hint that small objects gain less suggests a testable extension: constrain the inpainted background area or use an object-aware mask so the generator does not have to synthesize huge surrounding regions.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes POBF, a framework for data-scarce visual grounding. POBF generates synthetic training images by inpainting the region outside the ground-truth bounding box while keeping the boxed object intact, and it augments text by captioning the cropped box. A teacher model trained only on the limited real data then scores each synthetic image with a hardness score (IoU of the teacher's prediction with the box), an overfitting score (1 minus IoU when the box is masked), and a penalty term (IoU with an empty text query); the highest-scoring synthetic image per real sample is kept and added to the training set. Experiments on RefCOCO, RefCOCO+, RefCOCOg, and ReferItGame report an average gain of 5.83% over real-data-only training and consistent gains across different generators, captioners, data sizes, and student architectures.

Significance. If the empirical claims hold, POBF is a practical and flexible data-augmentation method for visual grounding under severe data scarcity. Its strengths are that the generation pipeline uses only off-the-shelf image-caption pretrained models, the filtering scheme is compared against several data-selection baselines, and the robustness experiments span four datasets, three architectures, three generators/captioners, and three training-data sizes. The paper is also transparent about the teacher being trained on real data only and about the held-out evaluation protocol. The main unresolved risk is that the central label-alignment assumption is asserted rather than verified, and the ablation table as printed does not support all of the stated component-wise gains.

major comments (4)
  1. [§4.1 and §5 Implementation Details] Label alignment is asserted rather than verified. The claim that painting outside the box leaves the object 'unchanged' and therefore 'strictly align[s] with the ground truth bounding box' is not demonstrated for the chosen strength of 0.9, which adds substantial noise to the base image and can alter pixels inside the box. Even if the box pixels were exactly preserved, replacing the background can invalidate relational referring expressions such as 'right cow' or 'bus in front of a building,' so the original text T may no longer uniquely refer to B in the new image. The hardness score in Eq. (1) checks only whether the teacher predicts B from T; it does not verify that T is a valid description of the boxed object in the synthetic image. I request a label-validity check (human evaluation on a sample, or an automated VQA/CLIP-based check) and a report of how often the box content changes at strength=0.9, because the 5.83% gain is attributed to a generation strategy whose labels are assumed correct.
  2. [Table 2] The headline comparison against X-Paste, Gen2Det, and GeoDiffusion is based on a single randomly sampled subset with no error bars, while the three-subset average is reported only for the Real and POBF rows. Thus the claim of outperforming leading baselines by 2.29%–3.85% rests on one draw. Please report the three-subset average (or at least standard deviations) for the baselines as well, or provide a significance test. In addition, the header 'All methods employ the proposed filtering scheme' needs a clear statement of which hyperparameters (K, q, lambda) were shared by the baselines, since the table otherwise reads as a comparison of full pipelines rather than generation strategies under a common filter.
  3. [§5.2, Table 3] The ablation table does not support the stated per-component gains. The text claims that S1, S2, and P give gains of 1.23%, 0.71%, and 1.14% 'compared to the variant without any filtering,' but Table 3 contains no visible row for S2 alone or P alone, and the differences between the no-filter Real+Synimg row (31.31) and the filtering rows are not 1.23, 0.71, and 1.14. The reported 'improvement of 1.96% over the baseline without filtering' for the full scheme also does not match any visible no-filter baseline; the only visible pairwise difference equal to 1.96 is with the S1-only row, which is itself a filtering variant. Please clarify the row encodings or add the missing rows so that each component's individual contribution can be verified.
  4. [§4.2, Eq. (3)] The penalty term P is computed as IoU(T(I', ∅), B), but TransVG and most grounding models require a non-empty text query. The manuscript does not specify how the empty string is encoded or why a box prediction with no text measures 'prior knowledge.' Since the penalty term is one of the three components of the final score in Eq. (4), this operational detail directly affects reproducibility of the filtering scheme.
minor comments (4)
  1. [Abstract vs. Conclusion] The reported margins over baselines are inconsistent: the Abstract and §5.1 state 2.29%–3.85%, while the Conclusion states 2.74%–4.35%; please correct one of these.
  2. [§4.2, Eq. (4)] Please specify the population over which the three scores are normalized; if the normalization is per real image (over its K synthetic variants) rather than global, the selection behavior of Eq. (4) is different and should be described.
  3. [§5.1] The explanation that the modest improvement on ReferIt stems from small objects and large inpainted background regions is a hypothesis; either provide supporting evidence (for example, an object-size analysis) or clearly mark it as speculative.
  4. [Table 5] The rows labeled 'Replace the Image Captioner with ...' do not show the default BLIP captioner row in the table; adding it would make the comparison easier to read.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the central gains are empirical comparisons against held-out benchmarks with independently defined filtering scores.

full rationale

POBF's derivation chain is not circular. The generation strategy (paint outside the box) is a data-augmentation recipe, not a theorem: the claim that the object inside the box remains aligned is an empirical assumption about the inpainting model, and any failure of that assumption is a correctness risk rather than a logical circularity. The filtering scores S1, S2, and P are defined through a teacher model trained only on the limited real data, and the final student is evaluated on held-out test splits, so the reported 5.83% gain over the real-only baseline is an external empirical result rather than a quantity forced by the definitions. The only fitted parameters are the three lambda weights, which are tuned on the validation set, not on the test set, and the scores themselves are not constructed from the student's test performance. No equation in the paper reduces to its own input, no self-citation carries the central load, and no known result is merely renamed. The paper's own limitation regarding label alignment at high inpainting strength is a concern about validity of synthetic labels, not circularity.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

No invented physical or conceptual entities are introduced. The method relies on off-the-shelf generators and standard transformer grounding models. The main load-bearing assumptions are that outside-box inpainting preserves label validity, that teacher scores rank synthetic sample usefulness, and that easier or low-overfit samples transfer to the student. The three lambda weights are free parameters tuned on validation, while K and q are hand-chosen hyperparameters.

free parameters (5)
  • lambda_1 (hardness score weight) = 1.0 or 0.5 via grid search on validation; exact value per dataset not reported
    Weight for the hardness score S1 in Eq. (4), tuned on the validation set.
  • lambda_2 (overfitting score weight) = 1.0 or 0.5 via grid search on validation; exact value per dataset not reported
    Weight for the overfitting score S2 in Eq. (4), tuned on the validation set.
  • lambda_P (penalty term weight) = 1.0 or 0.5 via grid search on validation; exact value per dataset not reported
    Weight for the penalty term P in Eq. (4), tuned on the validation set.
  • K (synthetic images per real sample) = 4
    Number of generated images per real sample, chosen by hand; it determines the candidate pool from which one sample is selected.
  • q (caption replacement probability) = 0.3
    Probability of replacing the real text with a generated caption during student training, chosen by hand.
assumptions (4)
  • domain assumption The inpainting model preserves the content inside the bounding box exactly when painting outside, so the original box remains a valid label for the generated image.
    Invoked in Sec. 4.1. If inpainting alters the object or adds overlapping content, label misalignment persists. The paper provides qualitative examples but no quantitative check of box validity.
  • domain assumption A teacher model trained on 1% real data produces hardness and overfitting scores that rank synthetic samples by their usefulness for student training.
    Sec. 4.2. The filter's efficacy depends on teacher predictions being informative in a low-data regime.
  • domain assumption Selecting easier synthetic samples and samples that do not allow box prediction from background improves student generalization.
    Sec. 4.2. This is an empirical hypothesis supported by prior pruning literature [57], not a derived guarantee.
  • domain assumption IoU-based top-1 accuracy with a 0.5 threshold is an appropriate proxy for visual grounding quality.
    Evaluation metric in Sec. 5; common in the field but a simplifying convention.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding." pith.science (2026). https://pith.science/paper/I3QVJPTU

@misc{pith2026241200684,
  author       = {Pith},
  title        = {Pith review of: Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/I3QVJPTU}},
  note         = {Machine review of arXiv:2412.00684}
}
read the original abstract

Visual grounding aims to localize the image regions based on a textual query. Given the difficulty of large-scale data curation, we investigate how to effectively learn visual grounding under data-scarce settings in this paper. To address the data scarcity, we propose a novel framework, POBF (Paint Outside the Box and Filter). POBF synthesizes images by inpainting outside the box, tackling a label misalignment issue encountered in previous works. Furthermore, POBF leverages an innovative filtering scheme to select the most effective training data. This scheme combines a hardness score and an overfitting score, balanced by a penalty term. Extensive experiments across four benchmark datasets demonstrate that POBF consistently improves performance, achieving an average gain of 5.83\% over the real-data-only method and outperforming leading baselines by 2.29\%-3.85\% in accuracy. Additionally, we validate the robustness and generalizability of POBF across various generative models, training data sizes, and model architectures.

Figures

Figures reproduced from arXiv: 2412.00684 by the authors.

Figure 1
Figure 1. Examples of generated images from different methods. X-Paste [ [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of our proposed framework POBF, which consists of four steps: data generation, teacher training, data [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Illustration of the inputs used to compute each [PITH_FULL_IMAGE:figures/full_fig_p004_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Scatter plot illustrating the relationship between un [PITH_FULL_IMAGE:figures/full_fig_p007_4.png]
Figure 5
Figure 5. Figure 5: Qualitative examples illustrating the effectiveness [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

113 extracted references · 55 canonical work pages

  1. [1]

    Amro Abbas, Kushal Tirumala, Dániel Simig, Surya Ganguli, and Ari S Mor- cos. 2023. Semdedup: Data-efficient learning at web-scale through semantic deduplication. arXiv preprint arXiv:2303.09540 (2023)

  2. [2]

    Sara Beery, Grant Van Horn, and Pietro Perona. 2018. Recognition in terra incognita. In Proceedings of the ECCV . 456–473

  3. [3]

    Fangyi Chen, Han Zhang, Zhantao Yang, Hao Chen, Kai Hu, and Marios Sav- vides. 2024. RTGen: Generating Region-Text Pairs for Open-Vocabulary Object Detection. arXiv preprint arXiv:2405.19854 (2024)

  4. [4]

    Gongwei Chen, Leyang Shen, Rui Shao, Xiang Deng, and Liqiang Nie. 2024. Lion: Empowering multimodal large language model with dual-level visual knowledge. In Proceedings of the IEEE/CVF CVPR . 26540–26550

  5. [5]

    Kai Chen, Enze Xie, Zhe Chen, Yibo Wang, HONG Lanqing, Zhenguo Li, and Dit-Yan Yeung. 2024. GeoDiffusion: Text-Prompted Geometric Control for Object Detection Data Generation. In The Twelfth International Conference on Learning Representations

  6. [6]

    Keqin Chen, Zhao Zhang, Weili Zeng, Richong Zhang, Feng Zhu, and Rui Zhao

  7. [7]

    Pengfei Chen, Ben Ben Liao, Guangyong Chen, and Shengyu Zhang. 2019. Un- derstanding and utilizing deep neural networks trained with noisy labels. In International Conference on Machine Learning . PMLR, 1062–1070

  8. [8]

    Sijia Chen and Baochun Li. 2022. Multi-modal dynamic graph transformer for visual grounding. In Proceedings of the IEEE/CVF CVPR . 15534–15543

Show all 113 references
  1. [9]

    Sijia Chen and Baochun Li. 2023. Language-guided diffusion model for visual grounding. arXiv preprint arXiv:2308.09599 (2023)

  2. [10]

    Shengxin Chen, Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Guannan Jiang, and Ron- grong Ji. 2024. QueryMatch: A Query-based Contrastive Learning Framework for Weakly Supervised Visual Grounding. In Proceedings of the 32nd ACM Inter- national Conference on Multimedia . 4177–4186

  3. [11]

    Yicheng Chen, Xiangtai Li, Yining Li, Yanhong Zeng, Jianzong Wu, Xiangyu Zhao, and Kai Chen. 2024. Auto Cherry-Picker: Learning from High-quality Generative Data Driven by Language. arXiv preprint arXiv:2406.20085 (2024)

  4. [12]

    Jiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou, and Houqiang Li

  5. [13]

    Jiajun Deng, Zhengyuan Yang, Daqing Liu, Tianlang Chen, Wengang Zhou, Yanyong Zhang, Houqiang Li, and Wanli Ouyang. 2023. Transvg++: End-to-end visual grounding with language conditioned vision transformer.IEEE transactions on pattern analysis and machine intelligence (2023)

  6. [14]

    Zilin Du, Yunxin Li, Xu Guo, Yidan Sun, and Boyang Li. 2023. Training Multi- media Event Extraction With Generated Images and Captions. arXiv preprint arXiv:2306.08966 (2023)

  7. [15]

    Chengxiang Fan, Muzhi Zhu, Hao Chen, Yang Liu, Weijia Wu, Huaqi Zhang, and Chunhua Shen. 2024. DiverGen: Improving Instance Segmentation by Learning Wider Data Distribution with More Diverse Generative Data. In Proceedings of the IEEE/CVF CVPR. 3986–3995

  8. [16]

    Haoyang Fang, Boran Han, Shuai Zhang, Su Zhou, Cuixiong Hu, and Wen-Ming Ye. 2024. Data augmentation for object detection via controllable diffusion models. In Proceedings of the IEEE/CVF W ACV. 1257–1266

  9. [17]

    Chengjian Feng, Yujie Zhong, Zequn Jie, Weidi Xie, and Lin Ma. 2024. Instagen: Enhancing object detection by training on synthetic dataset. In Proceedings of the IEEE/CVF CVPR. 14121–14130

  10. [18]

    Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. 2020. Shortcut learning in deep neural networks. Nature Machine Intelligence 2, 11 (2020)

  11. [19]

    Zeyu Han, Fangrui Zhu, Qianru Lao, and Huaizu Jiang. 2024. Zero-shot referring expression comprehension via structural similarity between images and captions. In Proceedings of the IEEE/CVF CVPR . 14364–14374

  12. [20]

    Ruozhen He, Paola Cascante-Bonilla, Ziyan Yang, Alexander C Berg, and Vicente Ordonez. 2024. Improved Visual Grounding through Self-Consistent Explanations. In Proceedings of the IEEE/CVF CVPR . 13095–13105

  13. [21]

    Ruozhen He, Paola Cascante-Bonilla, Ziyan Yang, Alexander C Berg, and Vicente Ordonez. 2024. Learning from Models and Data for Visual Grounding. arXiv preprint arXiv:2403.13804 (2024)

  14. [22]

    Haojun Jiang, Yuanze Lin, Dongchen Han, Shiji Song, and Gao Huang. 2022. Pseudo-q: Generating pseudo language queries for visual grounding. In Proceed- ings of the IEEE/CVF CVPR

  15. [23]

    Lu Jiang, Zhengyuan Zhou, Thomas Leung, Li-Jia Li, and Li Fei-Fei. 2018. Mentor- net: Learning data-driven curriculum for very deep neural networks on corrupted labels. In International conference on machine learning . PMLR

  16. [24]

    Yiding Jiang, Allan Zhou, Zhili Feng, Sadhika Malladi, and J Zico Kolter. 2024. Adaptive Data Optimization: Dynamic Sample Selection with Scaling Laws.arXiv preprint arXiv:2410.11820 (2024)

  17. [25]

    Yang Jiao, Zequn Jie, Jingjing Chen, Lin Ma, and Yu-Gang Jiang. 2023. Suspected Objects Matter: Rethinking Model’s Prediction for One-stage Visual Grounding. In Proceedings of the 31st ACM International Conference on Multimedia . 17–26

  18. [26]

    Lei Jin, Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Guannan Jiang, Annan Shu, and Rongrong Ji. 2023. Refclip: A universal teacher for weakly supervised referring expression comprehension. In Proceedings of the IEEE/CVF CVPR . 2681–2690

  19. [27]

    Aishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve, Ishan Misra, and Nicolas Carion. 2021. Mdetr-modulated detection for end-to-end multi-modal understanding. In Proceedings of the IEEE/CVF ICCV . 1780–1790

  20. [28]

    Seil Kang, Jinyeong Kim, Junhyeok Kim, and Seong Jae Hwang. 2025. Your Large Vision-Language Model Only Needs A Few Attention Heads For Visual Grounding. arXiv preprint arXiv:2503.06287 (2025)

  21. [29]

    Sahar Kazemzadeh, Vicente Ordonez, Mark Matten, and Tamara Berg. 2014. Referitgame: Referring to objects in photographs of natural scenes. InProceedings of the 2014 conference on EMNLP . 787–798

  22. [30]

    Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al

  23. [31]

    Katherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang, Douglas Eck, Chris Callison-Burch, and Nicholas Carlini. 2021. Deduplicating training data makes language models better. arXiv preprint arXiv:2107.06499 (2021)

  24. [32]

    Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. 2022. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. In International conference on machine learning . PMLR, 12888–12900

  25. [33]

    In Proceedings of the IEEE/CVF ICCV

    Segment anything. In Proceedings of the IEEE/CVF ICCV . 4015–4026

  26. [34]

    Mingchen Li, Mahdi Soltanolkotabi, and Samet Oymak. 2020. Gradient descent with early stopping is provably robust to label noise for overparameterized neural networks. In International conference on artificial intelligence and statistics . 4313

  27. [35]

    Ziyi Li, Qinye Zhou, Xiaoyun Zhang, Ya Zhang, Yanfeng Wang, and Weidi Xie

  28. [36]

    Junnan Li, Richard Socher, and Steven CH Hoi. 2020. Dividemix: Learning with noisy labels as semi-supervised learning. arXiv preprint arXiv:2002.07394 (2020)

  29. [37]

    Pengyue Lin, Ruifan Li, Yuzhe Ji, Zhihan Yu, Fangxiang Feng, Zhanyu Ma, and Xiaojie Wang. 2024. Triple Alignment Strategies for Zero-shot Phrase Grounding under Weak Supervision. InProceedings of the 32nd ACM International Conference on Multimedia. 4312–4321

  30. [38]

    Daqing Liu, Hanwang Zhang, Feng Wu, and Zheng-Jun Zha. 2019. Learning to assemble neural module tree networks for visual grounding. In Proceedings of the IEEE/CVF ICCV. 4673–4682

  31. [39]

    InProceedings of the IEEE/CVF ICCV

    Open-vocabulary object segmentation with diffusion models. InProceedings of the IEEE/CVF ICCV . 7667

  32. [40]

    Yaoyuan Liang, Zhao Yang, Yansong Tang, Jiashuo Fan, Ziran Li, Jingang Wang, Philip HS Torr, and Shao-Lun Huang. 2023. Luna: Language as continuing anchors for referring expression comprehension. In Proceedings of the 31st ACM International Conference on Multimedia . 5174–5184

  33. [41]

    Yongfei Liu, Bo Wan, Lin Ma, and Xuming He. 2021. Relation-aware instance refinement for weakly supervised visual grounding. InProceedings of the IEEE/CVF CVPR. 5612–5621

  34. [42]

    Yang Liu, Jiahua Zhang, Qingchao Chen, and Yuxin Peng. 2023. Confidence-aware pseudo-label learning for weakly supervised visual grounding. In Proceedings of the IEEE/CVF ICCV. 2828–2838

  35. [43]

    Haotian Liu, Chunyuan Li, Qingyang Wu, and Yong Jae Lee. 2024. Visual instruc- tion tuning. Advances in neural information processing systems 36 (2024)

  36. [44]

    Xihui Liu, Zihao Wang, Jing Shao, Xiaogang Wang, and Hongsheng Li. 2019. Improving referring expression grounding with cross-modal attention-guided erasing. In Proceedings of the IEEE/CVF CVPR . 1950–1959

  37. [45]

    when to update

    Eran Malach and Shai Shalev-Shwartz. 2017. Decoupling “when to update” from “how to update”. Advances in neural information processing systems 30 (2017)

  38. [46]

    Junhua Mao, Jonathan Huang, Alexander Toshev, Oana Camburu, Alan L Yuille, and Kevin Murphy. 2016. Generation and comprehension of unambiguous object descriptions. In Proceedings of the IEEE CVPR . 11–20. Conference, Zilin Du, Haoxin Li, Jianfei Yu, and Boyang Li

  39. [47]

    Tao Ma, Bing Bai, Haozhe Lin, Heyuan Wang, Yu Wang, Lin Luo, and Lu Fang

  40. [48]

    Mansheej Paul, Surya Ganguli, and Gintare Karolina Dziugaite. 2021. Deep learning on a data diet: Finding important examples early in training. Advances in neural information processing systems 34 (2021), 20596–20607

  41. [49]

    Anas Mahmoud, Mostafa Elhoushi, Amro Abbas, Yu Yang, Newsha Ardalani, Hugh Leather, and Ari S Morcos. 2024. Sieve: Multimodal dataset pruning using image captioning models. In Proceedings of the IEEE/CVF CVPR . 22423–22432

  42. [50]

    Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach. 2023. Sdxl: Improving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952 (2023)

  43. [51]

    Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF CVPR . 10684–10695

  44. [52]

    Varun K Nagaraja, Vlad I Morariu, and Larry S Davis. 2016. Modeling context between objects for referring expression understanding. In Computer Vision– ECCV 2016: 14th European Conference . Springer, 792–807

  45. [53]

    Noam Rotstein, David Bensaïd, Shaked Brody, Roy Ganz, and Ron Kimmel. 2024. Fusecap: Leveraging large language models for enriched fused image captions. In Proceedings of the IEEE/CVF W ACV. 5689–5700

  46. [54]

    Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Alessandro Cappelli, Hamza Alobeidli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The RefinedWeb dataset for Falcon LLM: outperforming curated corpora with web data, and web data only.arX...

  47. [55]

    Fengyuan Shi, Ruopeng Gao, Weilin Huang, and Limin Wang. 2023. Dynamic mdetr: A dynamic multimodal transformer decoder for visual grounding. IEEE Transactions on Pattern Analysis and Machine Intelligence (2023)

  48. [56]

    Hwanjun Song, Minseok Kim, Dongmin Park, and Jae-Gil Lee. 2019. How does early stopping help generalization against label noise? arXiv preprint arXiv:1911.08059 (2019)

  49. [57]

    Daniel Rose, Vaishnavi Himakunthala, Andy Ouyang, Ryan He, Alex Mei, Yujie Lu, Michael Saxon, Chinmay Sonar, Diba Mirza, and William Yang Wang. 2023. Visual Chain of Thought: Bridging Logical Gaps with Multimodal Infillings.arXiv preprint arXiv:2305.02317 (2023)

  50. [58]

    Wei Su, Peihan Miao, Huanzhang Dou, and Xi Li. 2024. Scanformer: Referring expression comprehension by iteratively scanning. InProceedings of the IEEE/CVF CVPR. 13449–13458

  51. [59]

    Ramprasaath R Selvaraju, Michael Cogswell, Abhishek Das, Ramakrishna Vedan- tam, Devi Parikh, and Dhruv Batra. 2017. Grad-cam: Visual explanations from deep networks via gradient-based localization. In Proceedings of the IEEE ICCV . 618–626

  52. [60]

    Jiamu Sun, Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Guannan Jiang, Zhiyu Wang, and Rongrong Ji. 2023. Refteacher: A strong baseline for semi-supervised referring expression comprehension. In Proceedings of the IEEE/CVF CVPR . 19144–19154

  53. [61]

    Mingjie Sun, Jimin Xiao, Eng Gee Lim, Si Liu, and John Y Goulermas. 2021. Discriminative triad matching and reconstruction for weakly referring expression grounding. IEEE transactions on pattern analysis and machine intelligence 43, 11 (2021), 4189–4195

  54. [62]

    Ben Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli, and Ari Morcos

  55. [63]

    Saksham Suri, Fanyi Xiao, Animesh Sinha, Sean Culatana, Raghuraman Krish- namoorthi, Chenchen Zhu, and Abhinav Shrivastava. 2024. Gen2Det: Generate to Detect. In Synthetic Data for Computer Vision Workshop@ CVPR 2024

  56. [64]

    Tianyi Tang, Yushuo Chen, Yifan Du, Junyi Li, Wayne Xin Zhao, and Ji-Rong Wen

  57. [65]

    Wei Su, Peihan Miao, Huanzhang Dou, Gaoang Wang, Liang Qiao, Zheyang Li, and Xi Li. 2023. Language adaptive weight generation for multi-task visual grounding. In Proceedings of the IEEE/CVF CVPR . 10857–10866

  58. [66]

    Mariya Toneva, Alessandro Sordoni, Remi Tachet des Combes, Adam Trischler, Yoshua Bengio, and Geoffrey J Gordon. 2018. An empirical study of example forgetting during deep neural network learning. arXiv preprint arXiv:1812.05159 (2018)

  59. [67]

    Shijie Wang, Dahun Kim, Ali Taalimi, Chen Sun, and Weicheng Kuo. 2024. Learn- ing Visual Grounding from Generative Vision and Language Model.arXiv preprint arXiv:2407.14563 (2024)

  60. [68]

    Yuxi Sun, Shanshan Feng, Xutao Li, Yunming Ye, Jian Kang, and Xu Huang

  61. [69]

    In Proceedings of the 30th ACM International conference on Multimedia

    Visual grounding in remote sensing images. In Proceedings of the 30th ACM International conference on Multimedia . 404–412

  62. [70]

    Dulanga Weerakoon, Vigneshwaran Subbaraju, Tuan Tran, and Archan Misra

  63. [71]

    Xiaobo Xia, Jiale Liu, Jun Yu, Xu Shen, Bo Han, and Tongliang Liu. 2022. Moderate coreset: A universal method of data selection for real-world data-efficient deep learning. In The Eleventh International Conference on Learning Representations

  64. [72]

    arXiv preprint arXiv:2305.16944 (2023)

    Learning to Imagine: Visually-Augmented Natural Language Generation. arXiv preprint arXiv:2305.16944 (2023)

  65. [73]

    Yonglong Tian, Lijie Fan, Kaifeng Chen, Dina Katabi, Dilip Krishnan, and Phillip Isola. 2023. Learning vision from models rivals learning vision from data. arXiv preprint arXiv:2312.17742 (2023)

  66. [74]

    Linhui Xiao, Xiaoshan Yang, Fang Peng, Yaowei Wang, and Changsheng Xu. 2024. HiVG: Hierarchical Multimodal Fine-grained Modulation for Visual Grounding. arXiv preprint arXiv:2404.13400 (2024)

  67. [75]

    Linhui Xiao, Xiaoshan Yang, Fang Peng, Ming Yan, Yaowei Wang, and Chang- sheng Xu. 2023. CLIP-VG: Self-paced Curriculum Adapting of CLIP for Visual Grounding. IEEE Transactions on Multimedia (2023)

  68. [76]

    Sai Wang, Yutian Lin, and Yu Wu. 2024. Omni-Q: Omni-Directional Scene Understanding for Unsupervised Visual Grounding. InProceedings of the IEEE/CVF CVPR. 14261–14270

  69. [77]

    Yibo Wang, Ruiyuan Gao, Kai Chen, Kaiqiang Zhou, Yingjie Cai, Lanqing Hong, Zhenguo Li, Lihui Jiang, Dit-Yan Yeung, Qiang Xu, et al . 2024. Detdiffusion: Synergizing generative and perceptive models for enhanced data generation and perception. In Proceedings of the IEEE/CVF CV...

  70. [78]

    Yunyi Xuan, Weijie Chen, Shicai Yang, Di Xie, Luojun Lin, and Yueting Zhuang

  71. [79]

    In Proceedings of the 30th ACM International Conference on Multimedia

    SoftSkip: Empowering multi-modal dynamic pruning for single-stage referring comprehension. In Proceedings of the 30th ACM International Conference on Multimedia. 3608–3616

  72. [80]

    Siming Yan, Min Bai, Weifeng Chen, Xiong Zhou, Qixing Huang, and Li Erran Li

  73. [81]

    Kai Xiao, Logan Engstrom, Andrew Ilyas, and Aleksander Madry. 2020. Noise or signal: The role of image backgrounds in object recognition. arXiv preprint arXiv:2006.09994 (2020)

  74. [82]

    Linhui Xiao, Xiaoshan Yang, Xiangyuan Lan, Yaowei Wang, and Changsheng Xu. 2024. Towards Visual Grounding: A Survey. arXiv preprint arXiv:2412.20206 (2024)

  75. [83]

    Shuo Yang, Zeke Xie, Hanyu Peng, Min Xu, Mingming Sun, and Ping Li. 2022. Dataset pruning: Reducing training data by examining generalization influence. arXiv preprint arXiv:2205.09329 (2022)

  76. [84]

    Yue Yang, Wenlin Yao, Hongming Zhang, Xiaoyang Wang, Dong Yu, and Jianshu Chen. 2022. Z-LaVI: Zero-Shot Language Solver Fueled by Visual Imagination. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing. 1186–1203

  77. [85]

    Jiahao Xie, Wei Li, Xiangtai Li, Ziwei Liu, Yew Soon Ong, and Chen Change Loy

  78. [86]

    International Journal of Computer Vision (2024), 1–20

    Mosaicfusion: Diffusion models as data augmenters for large vocabulary instance segmentation. International Journal of Computer Vision (2024), 1–20

  79. [87]

    Sang Michael Xie, Shibani Santurkar, Tengyu Ma, and Percy S Liang. 2023. Data selection for language models via importance resampling. Advances in Neural Information Processing Systems 36 (2023), 34201–34227

  80. [88]

    Ruilin Yao, Shengwu Xiong, Yichen Zhao, and Yi Rong. 2024. Visual Ground- ing with Multi-modal Conditional Adaptation. In Proceedings of the 32nd ACM International Conference on Multimedia . 3877–3886

  81. [89]

    In Proceedings of the 31st ACM International Conference on Multimedia

    Distilling Vision-Language Foundation Models: A Data-Free Approach via Prompt Diversification. In Proceedings of the 31st ACM International Conference on Multimedia. 4928–4938

  82. [90]

    Han Xue, Zhiwu Huang, Qianru Sun, Li Song, and Wenjun Zhang. 2023. Freestyle layout-to-image synthesis. In Proceedings of the IEEE/CVF CVPR . 14256–14266

  83. [91]

    Licheng Yu, Zhe Lin, Xiaohui Shen, Jimei Yang, Xin Lu, Mohit Bansal, and Tamara L Berg. 2018. Mattnet: Modular attention network for referring ex- pression comprehension. In Proceedings of the IEEE CVPR . 1307–1315

  84. [92]

    Vigor: Improving visual grounding of large vision language models with fine-grained reward modeling. In ECCV. Springer, 37–53

  85. [93]

    Lihe Yang, Xiaogang Xu, Bingyi Kang, Yinghuan Shi, and Hengshuang Zhao. 2024. Freemask: Synthetic images with dense annotations make stronger segmentation models. Advances in Neural Information Processing Systems 36 (2024)

  86. [94]

    Li Yang, Yan Xu, Chunfeng Yuan, Wei Liu, Bing Li, and Weiming Hu. 2022. Improv- ing visual grounding with visual-linguistic verification and iterative reasoning. In Proceedings of the IEEE/CVF CVPR . 9499–9508

  87. [95]

    Yunan Zeng, Yan Huang, Jinjin Zhang, Zequn Jie, Zhenhua Chai, and Liang Wang. 2024. Investigating compositional challenges in vision-language models for visual grounding. In Proceedings of the IEEE/CVF CVPR . 14141–14151

  88. [96]

    Manlin Zhang, Jie Wu, Yuxi Ren, Ming Li, Jie Qin, Xuefeng Xiao, Wei Liu, Rui Wang, Min Zheng, and Andy J Ma. 2023. Diffusionengine: Diffusion model is scalable data engine for object detection. arXiv preprint arXiv:2309.03893 (2023)

  89. [97]

    Ziyan Yang, Kushal Kafle, Franck Dernoncourt, and Vicente Ordonez. 2023. Im- proving visual grounding by encouraging consistent gradient-based explanations. In Proceedings of the IEEE/CVF CVPR . 19165–19174

  90. [98]

    Zuhao Yang, Fangneng Zhan, Kunhao Liu, Muyu Xu, and Shijian Lu. 2023. AI- Generated Images as Data Source: The Dawn of Synthetic Era. arXiv preprint arXiv:2310.01830 (2023)

  91. [99]

    Quanming Yao, Hansi Yang, Bo Han, Gang Niu, and James Tin-Yau Kwok. 2020. Searching to exploit memorization effect in learning with noisy labels. In Inter- national Conference on Machine Learning . PMLR, 10789–10798

  92. [101]

    Fulong Ye, Yuxing Long, Fangxiang Feng, and Xiaojie Wang. 2023. Whether you can locate or not? Interactive Referring Expression Generation. In Proceedings of the 31st ACM International Conference on Multimedia . 4697–4706

  93. [102]

    Jiabo Ye, Junfeng Tian, Ming Yan, Xiaoshan Yang, Xuwu Wang, Ji Zhang, Liang He, and Xin Lin. 2022. Shifting more attention to visual backbone: Query- modulated refinement networks for end-to-end visual grounding. In Proceedings of the IEEE/CVF CVPR . 15502

  94. [104]

    Licheng Yu, Patrick Poirson, Shan Yang, Alexander C Berg, and Tamara L Berg

  95. [106]

    Xingrui Yu, Bo Han, Jiangchao Yao, Gang Niu, Ivor Tsang, and Masashi Sugiyama

  96. [108]

    Zhihan Yu and Ruifan Li. 2024. Revisiting counterfactual problems in referring expression comprehension. In Proceedings of the IEEE/CVF CVPR . 13438–13448. Paint Outside the Box: Synthesizing and Selecting Training Data for Visual Grounding Conference,

  97. [111]

    Haoyu Zhao, Wenhang Ge, and Ying-cong Chen. 2024. LLM-Optic: Unveiling the Capabilities of Large Language Models for Universal Visual Grounding. arXiv preprint arXiv:2405.17104 (2024)

  98. [112]

    Hanqing Zhao, Dianmo Sheng, Jianmin Bao, Dongdong Chen, Dong Chen, Fang Wen, Lu Yuan, Ce Liu, Wenbo Zhou, Qi Chu, et al . 2023. X-paste: Revisiting scalable copy-paste for instance segmentation using clip and stablediffusion. In International Conference on Machine Learning . P...

  99. [113]

    Minghang Zheng, Jiahua Zhang, Qingchao Chen, Yuxin Peng, and Yang Liu. 2024. ResVG: Enhancing Relation and Semantic Understanding in Multiple Instances for Visual Grounding. In Proceedings of the 32nd ACM International Conference on Multimedia. 1187–1196

  100. [2016]

    In ECCV 2016, Proceedings, Part II 14

    Modeling context in referring expressions. In ECCV 2016, Proceedings, Part II 14. Springer, 69–85

  101. [2019]

    In International Conference on Machine Learning

    How does disagreement help generalization against label corruption?. In International Conference on Machine Learning . PMLR, 7164

  102. [2021]

    In Proceedings of the IEEE/CVF ICCV

    Transvg: End-to-end visual grounding with transformers. In Proceedings of the IEEE/CVF ICCV. 1769–1779

  103. [2022]

    Advances in Neural Information Processing Systems 35 (2022), 19523–19536

    Beyond neural scaling laws: beating power law scaling via data pruning. Advances in Neural Information Processing Systems 35 (2022), 19523–19536

  104. [2023]

    arXiv preprint arXiv:2306.15195 (2023)

    Shikra: Unleashing multimodal llm’s referential dialogue magic. arXiv preprint arXiv:2306.15195 (2023)

  105. [2024]

    In Proceedings of the IEEE/CVF CVPR

    When visual grounding meets gigapixel-level large-scale scenes: benchmark and approach. In Proceedings of the IEEE/CVF CVPR . 22119–22128

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.