REVIEW 2 cited by
STRAP: Structured Object Affordance Segmentation with Point Supervision
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
With significant annotation savings, point supervision has been proven effective for numerous 2D and 3D scene understanding problems. This success is primarily attributed to the structured output space; i.e., samples with high spatial affinity tend to share the same labels. Sharing this spirit, we study affordance segmentation with point supervision, wherein the setting inherits an unexplored dual affinity-spatial affinity and label affinity. By label affinity, we refer to affordance segmentation as a multi-label prediction problem: A plate can be both holdable and containable. By spatial affinity, we refer to a universal prior that nearby pixels with similar visual features should share the same point annotation. To tackle label affinity, we devise a dense prediction network that enhances label relations by effectively densifying labels in a new domain (i.e., label co-occurrence). To address spatial affinity, we exploit a Transformer backbone for global patch interaction and a regularization loss. In experiments, we benchmark our method on the challenging CAD120 dataset, showing significant performance gains over prior methods.
Forward citations
Cited by 2 Pith papers
-
Zero-shot 2D Grounding with Novel Affordance Types
A new benchmark and pipeline show that grounding an action word like 'cut' on a familiar object can be done without ever training on that action word, outperforming prior affordance grounding methods by a large margin.
-
GEAL: Generalizable 3D Affordance Learning with Cross-Modal Consistency
A dual-branch 3D affordance model that renders point clouds with Gaussian splatting and aligns 2D foundation-model features to 3D beats prior methods on clean and corrupted benchmarks.
Discussion (0). Continue with ORCID to comment.