PathVQA is the first public dataset of over 32,000 questions on nearly 5,000 pathology images for medical visual question answering.
Imagenet: A large-scale hierarchical image database
9 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
Dataset distillation creates a tiny synthetic training set that, when used with a fixed network initialization, produces models whose performance approximates that of models trained on the full original dataset.
VICReg prevents collapse in self-supervised image embeddings via explicit variance, invariance, and covariance regularization and matches state-of-the-art downstream performance.
Seer, a transformer-based PIDM pre-trained on large robotic datasets like DROID, outperforms prior methods on simulation and real-world robotic manipulation benchmarks with gains up to 43%.
Introduces the QEVD benchmark for asynchronous situated interaction in fitness coaching and proposes a streaming baseline to address limitations of existing vision-language models.
A two-stage pipeline generates pseudo masks from image-level labels to train Mask R-CNN, achieving state-of-the-art results on PASCAL VOC 2012 for weakly supervised instance segmentation.
PartCo improves generalized category discovery by incorporating part-level correspondence priors that capture finer semantic structures and integrate with existing GCD methods.
GPT-4V processes interleaved image-text inputs generically and supports visual referring prompting for new human-AI interaction.
Baidu-UTS won the EPIC-Kitchens challenge by guiding 3D CNN training with object detection features via a Gated Feature Aggregator to improve noun prediction.
citing papers explorer
-
PathVQA: 30000+ Questions for Medical Visual Question Answering
PathVQA is the first public dataset of over 32,000 questions on nearly 5,000 pathology images for medical visual question answering.
-
Dataset Distillation
Dataset distillation creates a tiny synthetic training set that, when used with a fixed network initialization, produces models whose performance approximates that of models trained on the full original dataset.
-
VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning
VICReg prevents collapse in self-supervised image embeddings via explicit variance, invariance, and covariance regularization and matches state-of-the-art downstream performance.
-
Predictive Inverse Dynamics Models are Scalable Learners for Robotic Manipulation
Seer, a transformer-based PIDM pre-trained on large robotic datasets like DROID, outperforms prior methods on simulation and real-world robotic manipulation benchmarks with gains up to 43%.
-
What to Say and When to Say it: Live Fitness Coaching as a Testbed for Situated Interaction
Introduces the QEVD benchmark for asynchronous situated interaction in fitness coaching and proposes a streaming baseline to address limitations of existing vision-language models.
-
Where are the Masks: Instance Segmentation with Image-level Supervision
A two-stage pipeline generates pseudo masks from image-level labels to train Mask R-CNN, achieving state-of-the-art results on PASCAL VOC 2012 for weakly supervised instance segmentation.
-
PartCo: Part-Level Correspondence Priors Enhance Category Discovery
PartCo improves generalized category discovery by incorporating part-level correspondence priors that capture finer semantic structures and integrate with existing GCD methods.
-
The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)
GPT-4V processes interleaved image-text inputs generically and supports visual referring prompting for new human-AI interaction.
-
Baidu-UTS Submission to the EPIC-Kitchens Action Recognition Challenge 2019
Baidu-UTS won the EPIC-Kitchens challenge by guiding 3D CNN training with object detection features via a Gated Feature Aggregator to improve noun prediction.