A dual-stream diffusion model that jointly denoises the output image and an insertion mask, trained on a new 3.16 million pair dataset, outperforms prior baselines on affordance-aware object insertion.
In Defense of the Direct Perception of Affordances
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
The field of functional recognition or affordance estimation from images has seen a revival in recent years. As originally proposed by Gibson, the affordances of a scene were directly perceived from the ambient light: in other words, functional properties like sittable were estimated directly from incoming pixels. Recent work, however, has taken a mediated approach in which affordances are derived by first estimating semantics or geometry and then reasoning about the affordances. In a tribute to Gibson, this paper explores his theory of affordances as originally proposed. We propose two approaches for direct perception of affordances and show that they obtain good results and can out-perform mediated approaches. We hope this paper can rekindle discussion around direct perception and its implications in the long term.
fields
cs.CV 1years
2024 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion
A dual-stream diffusion model that jointly denoises the output image and an insertion mask, trained on a new 3.16 million pair dataset, outperforms prior baselines on affordance-aware object insertion.