Latent diffusion models achieve state-of-the-art inpainting and competitive results on unconditional generation, scene synthesis, and super-resolution by performing the diffusion process in the latent space of pretrained autoencoders with cross-attention conditioning, while cutting computational and
The open images dataset v4: Unified image classification, object de- tection, and visual relationship detection at scale
6 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
dataset 1polarities
use dataset 1representative citing papers
Optuna introduces a define-by-run hyperparameter optimization framework with efficient searching, pruning, and support for distributed to interactive use cases.
Augmenting robot datasets via diffusion-based semantic inpainting enables manipulation policies to solve unseen tasks with new objects and improves robustness to novel distractors.
MHSA mitigates hallucinations in LVLMs by training an MLP to steer cross-modal attention, extending detection work to mitigation via attention replacement at inference.
Two data selection techniques (GMM visual similarity and bounding-box diversity) reduce required weakly labeled images by up to 100x on Open Images and 20x on Cityscapes while maintaining semantic segmentation performance.
Obj-GloVe is a contextual embedding for visual objects derived from scene co-occurrences using the GloVe method, shown useful for object detection and text-to-image synthesis.
citing papers explorer
-
High-Resolution Image Synthesis with Latent Diffusion Models
Latent diffusion models achieve state-of-the-art inpainting and competitive results on unconditional generation, scene synthesis, and super-resolution by performing the diffusion process in the latent space of pretrained autoencoders with cross-attention conditioning, while cutting computational and
-
Optuna: A Next-generation Hyperparameter Optimization Framework
Optuna introduces a define-by-run hyperparameter optimization framework with efficient searching, pruning, and support for distributed to interactive use cases.
-
Scaling Robot Learning with Semantically Imagined Experience
Augmenting robot datasets via diffusion-based semantic inpainting enables manipulation policies to solve unseen tasks with new objects and improves robustness to novel distractors.
-
MHSA: A Lightweight Framework for Mitigating Hallucinations via Steered Attention in LVLMs
MHSA mitigates hallucinations in LVLMs by training an MLP to steer cross-modal attention, extending detection work to mitigation via attention replacement at inference.
-
Data Selection for training Semantic Segmentation CNNs with cross-dataset weak supervision
Two data selection techniques (GMM visual similarity and bounding-box diversity) reduce required weakly labeled images by up to 100x on Open Images and 20x on Cityscapes while maintaining semantic segmentation performance.
-
Obj-GloVe: Scene-Based Contextual Object Embedding
Obj-GloVe is a contextual embedding for visual objects derived from scene co-occurrences using the GloVe method, shown useful for object detection and text-to-image synthesis.