REVIEW 3 major objections 5 minor 94 references
Pix2Act learns 3D robot manipulation by predicting continuous 2D keypoint paths in stereo camera planes and recovering poses by triangulation, with per-camera equivariant augmentation that raises success over strong baselines.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.5
2026-07-14 06:32 UTC pith:PQ4U6KAE
load-bearing objection Solid robotics methods paper: continuous multi-view image actions plus n-equivariant augmentation deliver real gains; the “lossless” claim is a bit soft but not fatal. the 3 major comments →
Pix2Act: Image-Space Manipulation Policies with Equivariant Augmentation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
Pix2Act shows that representing SE(3) action chunks as continuous unbounded image-space keypoint trajectories on dual in-hand cameras, then recovering end-effector poses by triangulation, turns high-dimensional 3D control into a 2D prediction problem that still reconstructs full poses without discretization loss. Aligning observations and actions in image space enables equivariant per-camera transforms; with a multi-view transformer and shared diffusion heads that respect n-equivariance, the policy enlarges the effective training support, learns transform-invariant action structure, and outperforms strong image, voxel, point-cloud, and prior keypoint baselines in simulation and on a physical
What carries the argument
n-equivariance—the requirement that independent SE(2) transforms (and permutations) of n camera images must transform the matching image-space action chunks the same way—together with continuous unbounded image action chunks recovered by triangulation and Diffusion X-Net (multi-view attention plus shared per-view diffusion heads) that make that symmetry usable at train and test time.
Load-bearing premise
The method depends on a fixed, calibrated dual in-hand stereo pair and on patching near-plane blow-ups by holding the last valid image coordinate so triangulation stays accurate enough for precise control.
What would settle it
On high-precision insertion tasks, measure whether triangulated 3D trajectories systematically diverge from demonstration ground truth when 2D predictions look correct, especially under small camera miscalibration or frequent plane-crossing; if they do, and success falls to or below direct 3D action baselines, the lossless image-space claim fails.
If this is right
- Aligning actions with image coordinates makes joint observation–action augmentation routine without voxels or full 3D equivariant networks.
- Continuous (not discrete) multi-view keypoints plus cross-view fusion reduce out-of-frame and triangulation inconsistencies relative to prior pixel-keypoint policies.
- Independent head-camera augmentation can yield policies that stay usable when the agent view moves at deployment.
- Learning on orbits of multi-view transforms can improve sample efficiency by covering more of the data manifold from the same demos.
- The same image-action interface can carry multi-task language tokens with faster convergence, as the paper’s multi-task extension suggests.
Where Pith is reading between the lines
- Many closed-loop SE(3) policies may be spending capacity on a higher-dimensional action space than the image-plane geometry that actually drives contact.
- Extending the representation to multi-finger hands will need richer keypoints and will lose the simple 180° camera-swap symmetry of parallel jaws.
- If plane-crossing holds become common, hybrid 2D/3D action representations may be needed for free-space motions that leave the current camera plane.
- Independent per-camera augmentation could serve as a lightweight alternative to architectural equivariance for other multi-view control settings.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. Pix2Act reformulates closed-loop 3D manipulation as continuous, unbounded 2D keypoint trajectory prediction on dual in-hand camera planes, recovering SE(3) end-effector poses by triangulation. The image-space action representation enables n-equivariant data augmentation (independent per-camera SE(2) transforms of observations and actions, plus camera permutation equivariance; Props. 1–2) and is realized by Diffusion X-Net, a multi-view transformer with shared per-view DiT heads. On 10 MimicGen tasks with 100 demos, the method reports 75.2% average success versus 63.1% for the strongest baseline (EquiDiff-Voxel); ablations isolate equivariant augmentation, shared heads, and image-space vs. 3D actions (Table 2). Real UR5 experiments on four tasks outperform Diffusion Policy and remain robust under agent-view camera rotations. Limitations include a calibrated dual in-hand setup and a hold-last-valid heuristic for out-of-plane projections.
Significance. If the results hold under the stated hardware assumptions, the paper offers a practical and conceptually clean alternative to SE(3) action chunks: continuous image-space keypoints that are both more learnable and naturally aligned for equivariant augmentation. The n-equivariance framing (SE(2) ≀ S_n), the shared-head Diffusion X-Net design that respects it, and the consistent gains on high-precision MimicGen variants and real-robot camera-perturbation tests are concrete contributions. Strengths include a multi-task extension, real-world transfer with measured reconstruction error ~10^{-6}, and ablations that separate representation, architecture, and augmentation effects. The work is of clear interest to the imitation-learning and equivariant-robotics communities, provided the projection pipeline’s practical losslessness is better quantified.
major comments (3)
- Abstract and §4.1 claim the projection–triangulation pipeline is “lossless” / “fully invertible” and recovers “full precision.” Appendix 7.7, however, replaces image coordinates with |pix| > ρ ≈ 1.5s by the previous action to avoid out-of-plane blow-ups. This heuristic is load-bearing for high-precision tasks (Square, Threading) yet is neither quantified (how often it fires on MimicGen demos/rollouts) nor ablated. Without that evidence the “lossless” claim overstates the method’s guarantees under the fixed dual in-hand geometry.
- Table 1 reports best-checkpoint mean success over 50 tests with no error bars or seed statistics in the main table (Appendix 7.8 mentions two seeds only for simulation evaluation). Given that the central claim is a 12-point average gain over EquiDiff-Voxel and larger gains on high-precision tasks, the absence of variability measures makes it hard to judge whether the ranking is stable. At minimum, means ± std over seeds (or confidence intervals) should be added for the main comparison and the key ablations in Table 2.
- §5.1 / Table 1 baselines: EquiDiff (Img/Voxel), ACT, and DP3 results are taken from Wang et al. [30] and use different observation setups (single in-hand or multi-depth + one in-hand) than Pix2Act’s dual in-hand + agent view. Motion Track and Diffusion Policy share the image setup, but the cross-paper numbers for the 3D baselines are not re-run under matched cameras. The paper should either re-evaluate the strongest 3D baselines under the same dual in-hand + agent-view configuration or clearly qualify that those comparisons are not hardware-matched.
minor comments (5)
- Figure 1 and the abstract emphasize “lossless” recovery; a short quantitative statement of triangulation residual on held-out demos (already ~10^{-6} in real-world text) should appear in the main body, not only in §5.2.
- Proposition 1–2 and the wreath-product claim are clear, but a one-line statement that n-equivariance of π implies SE(2)^n ⋊ S_n invariance of the triangulated 3D policy would help readers who skip the appendix.
- Table 2 “w/o shared head” shows a 42-point drop on Square; a brief discussion of why permutation equivariance matters more on insertion tasks would strengthen the architecture claim.
- Typos / polish: “V oxel” spacing in Table 1; “Equidiff” vs “EquiDiff” inconsistency; “n-equivarianceto” missing space in §4.2.
- Code and trained checkpoints are not mentioned; releasing them would substantially raise reproducibility given the custom X-Net and dual-camera setup.
Circularity Check
No circular derivation: empirical imitation-learning method with geometric action representation and held-out success metrics; self-citations are related-work baselines only.
full rationale
Pix2Act is an empirical robotics/imitation-learning paper. Its load-bearing chain is (i) a geometric reparameterization of SE(3) action chunks as continuous unbounded multi-view image keypoints with triangulation recovery (P,T), (ii) n-equivariance (Props. 1–2) enforced by per-camera SE(2) data augmentation and a shared-head Diffusion X-Net, and (iii) success rates on held-out MimicGen episodes and real UR5 trials versus external baselines. Nothing in that chain reduces a claimed prediction to a fitted constant or to a self-defined quantity: training minimizes a standard diffusion noise objective on demonstration image-action chunks; evaluation reports task success on unseen tests (Tables 1–3). Projection–triangulation invertibility is a camera-geometry claim under stated calibration assumptions, not a fit renamed as a result. Self-citations (e.g. EquiDiff [30], prior equivariant transporter/grasp work by overlapping authors) appear as related work and baselines, not as uniqueness theorems or load-bearing premises that force the performance claim. No fitted-input-as-prediction, no ansatz smuggled as derivation, no renaming of a known empirical law. Circularity score is therefore 0.
Axiom & Free-Parameter Ledger
free parameters (5)
- SE(2) augmentation rotation range =
[-30°, +30°]
- SE(2) augmentation translation range =
±h/8, ±w/8
- Out-of-plane stability threshold ρ =
≈1.5 × image size
- Action chunk horizon and execute length =
h=12 exec=8 (sim); horizon 10 (real)
- Diffusion training schedule and network widths =
as listed in Appendix 7.7–7.8
axioms (5)
- domain assumption Known camera matrices and fixed dual in-hand offset relative to the end-effector enable invertible projection/triangulation of gripper keypoints.
- domain assumption Four gripper-finger keypoints plus scalar width bijectively encode parallel-jaw SE(3) pose for the tasks considered.
- ad hoc to paper Independent per-camera SE(2) transforms of observations should transform image-space action chunks equivariantly (Prop. 1), and camera permutations should permute action chunks (Prop. 2).
- domain assumption Imitation from finite expert demos with diffusion denoising yields a policy that generalizes under MimicGen spatial variants and real initial-condition variation.
- standard math Standard multiview geometry: triangulation of corresponding continuous image points recovers 3D points (up to calibration error).
invented entities (3)
-
image action chunk (continuous unbounded multi-view keypoint trajectories)
no independent evidence
-
n-equivariance (SE(2)≀Sn symmetry of multi-camera image-action policies)
no independent evidence
-
Diffusion X-Net
no independent evidence
read the original abstract
Representing manipulation actions as 2D trajectories in the camera plane provides a compact and interpretable basis for learning complex 3D manipulation policies. However, it also creates challenges from out-of-frame trajectories and limited precision. We propose Pix2Act, an imitation learning method that addresses these challenges by generating continuous image-space keypoint trajectories in each camera plane and losslessly recovering end-effector poses via triangulation. This reformulates high-dimensional 3D control as a simpler, more learnable 2D prediction problem. Crucially, it aligns observations and actions in the same coordinate space, enabling equivariant transformations to jointly rotate individual camera images together with their image-space actions. We analyze the symmetry properties of this augmentation and design a network architecture that can fuse multiple camera views while respecting their per-view rotations. As a result, Pix2Act implicitly enlarges the support of the data distribution and learns invariant action structures across transformations, yielding improved generalization and overall performance. Across diverse simulated and real-world manipulation tasks, Pix2Act outperforms state-of-the-art baselines and remains robust under camera perturbations.
Figures
Reference graph
Works this paper leans on
-
[1]
J. Ren, P. Sundaresan, D. Sadigh, S. Choudhury, and J. Bohg. Motion tracks: A unified represen- tation for human-robot transfer in few-shot imitation learning.arXiv preprint arXiv:2501.06994, 2025
arXiv 2025
-
[2]
Song and S
Y . Song and S. Ermon. Generative modeling by estimating gradients of the data distribution. Advances in neural information processing systems, 32, 2019
2019
-
[3]
Y . Song, J. Sohl-Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole. Score-based generative modeling through stochastic differential equations.arXiv preprint arXiv:2011.13456, 2020
Pith/arXiv arXiv 2011
-
[4]
Alain and Y
G. Alain and Y . Bengio. What regularized auto-encoders learn from the data-generating distribution.The Journal of Machine Learning Research, 15(1):3563–3593, 2014
2014
-
[5]
D. Wang, B. Hu, S. Song, R. Walters, and R. Platt. A practical guide for incorporating symmetry in diffusion policy. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. URLhttps://openreview.net/forum?id=e0Dn7dg395
2025
-
[6]
P. Y . Simard, D. Steinkraus, J. C. Platt, et al. Best practices for convolutional neural networks applied to visual document analysis. InIcdar, volume 3. Edinburgh, 2003
2003
-
[7]
Laskin, K
M. Laskin, K. Lee, A. Stooke, L. Pinto, P. Abbeel, and A. Srinivas. Reinforcement learning with augmented data.Advances in neural information processing systems, 33:19884–19895, 2020
2020
-
[8]
Radford, J
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al. Learning transferable visual models from natural language supervision. InInternational conference on machine learning, pages 8748–8763. PMLR, 2021
2021
-
[9]
J. Li, D. Li, C. Xiong, and S. Hoi. Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation. InInternational conference on machine learning, pages 12888–12900. PMLR, 2022
2022
-
[10]
Kafle, M
K. Kafle, M. Yousefhussien, and C. Kanan. Data augmentation for visual question answering. InProceedings of the 10th international conference on natural language generation, pages 198–202, 2017
2017
-
[11]
B. Xiao, H. Wu, W. Xu, X. Dai, H. Hu, Y . Lu, M. Zeng, C. Liu, and L. Yuan. Florence-2: Advancing a unified representation for a variety of vision tasks. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 4818–4829, 2024
2024
-
[12]
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhou, et al. Chain-of- thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022
2022
-
[13]
Shorten, T
C. Shorten, T. M. Khoshgoftaar, and B. Furht. Text data augmentation for deep learning.Journal of big Data, 8(1):101, 2021
2021
-
[14]
Sennrich, B
R. Sennrich, B. Haddow, and A. Birch. Improving neural machine translation models with mono- lingual data. InProceedings of the 54th annual meeting of the association for computational linguistics (volume 1: long papers), pages 86–96, 2016
2016
-
[15]
J. Wei, M. Bosma, V . Y . Zhao, K. Guu, A. W. Yu, B. Lester, N. Du, A. M. Dai, and Q. V . Le. Finetuned language models are zero-shot learners.arXiv preprint arXiv:2109.01652, 2021. 9
Pith/arXiv arXiv 2021
-
[16]
Tobin, R
J. Tobin, R. Fong, A. Ray, J. Schneider, W. Zaremba, and P. Abbeel. Domain randomization for transferring deep neural networks from simulation to the real world. In2017 IEEE/RSJ international conference on intelligent robots and systems (IROS), pages 23–30. IEEE, 2017
2017
-
[17]
S. Dasari, F. Ebert, S. Tian, S. Nair, B. Bucher, K. Schmeckpeper, S. Singh, S. Levine, and C. Finn. Robonet: Large-scale multi-robot learning.arXiv preprint arXiv:1910.11215, 2019
Pith/arXiv arXiv 1910
-
[18]
H.-S. Fang, C. Wang, H. Fang, M. Gou, J. Liu, H. Yan, W. Liu, Y . Xie, and C. Lu. Anygrasp: Robust and efficient grasp perception in spatial and temporal domains.IEEE Transactions on Robotics, 39(5):3929–3945, 2023
2023
-
[19]
Z. Jiang, Y . Zhu, M. Svetlik, K. Fang, and Y . Zhu. Synergies between affordance and geometry: 6-dof grasp detection via implicit representations.arXiv preprint arXiv:2104.01542, 2021
Pith/arXiv arXiv 2021
-
[20]
Mandlekar, D
A. Mandlekar, D. Xu, J. Wong, S. Nasiriany, C. Wang, R. Kulkarni, L. Fei-Fei, S. Savarese, Y . Zhu, and R. Mart´ın-Mart´ın. What matters in learning from offline human demonstrations for robot manipulation. InConference on Robot Learning (CoRL), 2021
2021
-
[21]
N. Hansen, R. Jangir, Y . Sun, G. Aleny `a, P. Abbeel, A. A. Efros, L. Pinto, and X. Wang. Self-supervised policy adaptation during deployment.arXiv preprint arXiv:2007.04309, 2020
Pith/arXiv arXiv 2007
-
[22]
Shridhar, L
M. Shridhar, L. Manuelli, and D. Fox. Cliport: What and where pathways for robotic manipula- tion. InConference on robot learning, pages 894–906. PMLR, 2022
2022
-
[23]
A. Zeng, P. Florence, J. Tompson, S. Welker, J. Chien, M. Attarian, T. Armstrong, I. Krasin, D. Duong, V . Sindhwani, et al. Transporter networks: Rearranging the visual world for robotic manipulation. InConference on Robot Learning, pages 726–747. PMLR, 2021
2021
-
[24]
D. Wang, C. Kohler, and R. Platt. Policy learning in se (3) action spaces.arXiv preprint arXiv:2010.02798, 2020
Pith/arXiv arXiv 2010
-
[25]
Shridhar, L
M. Shridhar, L. Manuelli, and D. Fox. Perceiver-actor: A multi-task transformer for robotic manipulation. InConference on Robot Learning, pages 785–799. PMLR, 2023
2023
-
[26]
Goyal, J
A. Goyal, J. Xu, Y . Guo, V . Blukis, Y .-W. Chao, and D. Fox. Rvt: Robotic view transformer for 3d object manipulation. InConference on Robot Learning, pages 694–710. PMLR, 2023
2023
-
[27]
A. Goyal, V . Blukis, J. Xu, Y . Guo, Y .-W. Chao, and D. Fox. Rvt-2: Learning precise manipula- tion from few demonstrations.arXiv preprint arXiv:2406.08545, 2024
Pith/arXiv arXiv 2024
-
[28]
H. Huang, D. Wang, R. Walters, and R. Platt. Equivariant transporter network.arXiv preprint arXiv:2202.09400, 2022
Pith/arXiv arXiv 2022
-
[29]
M. Jia, D. Wang, G. Su, D. Klee, X. Zhu, R. Walters, and R. Platt. Seil: Simulation-augmented equivariant imitation learning.arXiv preprint arXiv:2211.00194, 2022
Pith/arXiv arXiv 2022
-
[30]
D. Wang, S. Hart, D. Surovik, T. Kelestemur, H. Huang, H. Zhao, M. Yeatman, J. Wang, R. Walters, and R. Platt. Equivariant diffusion policy.arXiv preprint arXiv:2407.01812, 2024
Pith/arXiv arXiv 2024
-
[31]
B. Hu, D. Wang, D. Klee, H. Tian, X. Zhu, H. Huang, R. Platt, and R. Walters. 3d equivariant visuomotor policy learning via spherical projection.arXiv preprint arXiv:2505.16969, 2025
arXiv 2025
-
[32]
T. Zhao, Y . Wang, W. Sun, Y . Chen, G. Niu, and M. Sugiyama. Representation learning for continuous action spaces is beneficial for efficient policy learning.Neural Networks, 159: 137–152, 2023
2023
-
[33]
Zitkovich, T
B. Zitkovich, T. Yu, S. Xu, P. Xu, T. Xiao, F. Xia, J. Wu, P. Wohlhart, S. Welker, A. Wahid, et al. Rt-2: Vision-language-action models transfer web knowledge to robotic control. InConference on Robot Learning, pages 2165–2183. PMLR, 2023. 10
2023
-
[34]
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, et al. Openvla: An open-source vision-language-action model.arXiv preprint arXiv:2406.09246, 2024
Pith/arXiv arXiv 2024
-
[35]
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y . Du, B. Burchfiel, R. Tedrake, and S. Song. Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025
2025
-
[36]
Levine, C
S. Levine, C. Finn, T. Darrell, and P. Abbeel. End-to-end training of deep visuomotor policies. Journal of Machine Learning Research, 17(39):1–40, 2016
2016
-
[37]
Y . Ze, G. Zhang, K. Zhang, C. Hu, M. Wang, and H. Xu. 3d diffusion policy.arXiv preprint arXiv:2403.03954, 2024
Pith/arXiv arXiv 2024
-
[38]
Y . Ze, Z. Chen, W. Wang, T. Chen, X. He, Y . Yuan, X. B. Peng, and J. Wu. Generalizable humanoid manipulation with 3d diffusion policies. In2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 2873–2880. IEEE, 2025
2025
-
[39]
Y . Zhou, C. Barnes, J. Lu, J. Yang, and H. Li. On the continuity of rotation representations in neural networks. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 5745–5753, 2019
2019
-
[40]
P. Zech, E. Renaudo, S. Haller, X. Zhang, and J. Piater. Action representations in robotics: A taxonomy and systematic classification.The International Journal of Robotics Research, 38(5): 518–562, 2019
2019
-
[41]
C. Pan, B. Okorn, H. Zhang, B. Eisner, and D. Held. Tax-pose: Task-specific cross-pose estimation for robot manipulation. InConference on Robot Learning, pages 1783–1792. PMLR, 2023
2023
-
[42]
H. Huang, K. Schmeckpeper, D. Wang, O. Biza, Y . Qian, H. Liu, M. Jia, R. Platt, and R. Walters. Imagination policy: Using generative point cloud models for learning manipulation policies. arXiv preprint arXiv:2406.11740, 2024
Pith/arXiv arXiv 2024
-
[43]
Huang, H
H. Huang, H. Liu, D. Wang, R. Walters, and R. Platt. Match policy: A simple pipeline from point cloud registration to manipulation policies. In2025 IEEE International Conference on Robotics and Automation (ICRA), pages 16907–16914. IEEE, 2025
2025
-
[44]
Choi and H
C. Choi and H. I. Christensen. Real-time 3d model-based tracking using edge and keypoint features for robotic manipulation. In2010 IEEE international conference on robotics and automation, pages 4048–4055. IEEE, 2010
2010
-
[45]
W. Huang, C. Wang, Y . Li, R. Zhang, and L. Fei-Fei. Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation.arXiv preprint arXiv:2409.01652, 2024
Pith/arXiv arXiv 2024
-
[46]
S. Patel, X. Yin, W. Huang, S. Garg, H. Nayyeri, L. Fei-Fei, S. Lazebnik, and Y . Li. A real-to- sim-to-real approach to robotic manipulation with vlm-generated iterative keypoint rewards. arXiv preprint arXiv:2502.08643, 2025
Pith/arXiv arXiv 2025
-
[47]
Z. Liu, M. Zhang, and Y . Li. Kuda: Keypoints to unify dynamics learning and visual prompting for open-vocabulary robotic manipulation.arXiv preprint arXiv:2503.10546, 2025
Pith/arXiv arXiv 2025
-
[48]
Y . Tian, J. Zhang, G. Huang, B. Wang, P. Wang, J. Pang, and H. Dong. Robokeygen: robot pose and joint angles estimation via diffusion-based 3d keypoint generation. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 5375–5381. IEEE, 2024
2024
-
[49]
S. Bahl, R. Mendonca, L. Chen, U. Jain, and D. Pathak. Affordances from human videos as a versatile representation for robotics. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 13778–13790, 2023. 11
2023
-
[50]
Montesano and M
L. Montesano and M. Lopes. Learning grasping affordances from local visual descriptors. In 2009 IEEE 8th international conference on development and learning, pages 1–6. IEEE, 2009
2009
-
[51]
Manuelli, W
L. Manuelli, W. Gao, P. Florence, and R. Tedrake. kpam: Keypoint affordances for category- level robotic manipulation. InThe International Symposium of Robotics Research, pages 132–157. Springer, 2019
2019
-
[52]
M. Jia, H. Huang, Z. Zhang, C. Wang, L. Zhao, D. Wang, J. X. Liu, R. Walters, R. Platt, and S. Tellex. Learning efficient and robust language-conditioned manipulation using textual-visual relevancy and equivariant language mapping.IEEE Robotics and Automation Letters, 2025
2025
-
[53]
Weiler and G
M. Weiler and G. Cesa. General E(2)-Equivariant Steerable CNNs. InConference on Neural Information Processing Systems (NeurIPS), 2019
2019
-
[54]
C. Deng, O. Litany, Y . Duan, A. Poulenard, A. Tagliasacchi, and L. J. Guibas. Vector neu- rons: A general framework for so (3)-equivariant networks. InProceedings of the IEEE/CVF International Conference on Computer Vision, pages 12200–12209, 2021
2021
-
[55]
G. Cesa, L. Lang, and M. Weiler. A program to build E(N)-equivariant steerable CNNs. In International Conference on Learning Representations, 2022. URL https://openreview. net/forum?id=WE4qe9xlnQw
2022
-
[56]
Y .-L. Liao and T. Smidt. Equiformer: Equivariant graph attention transformer for 3d atomistic graphs.arXiv preprint arXiv:2206.11990, 2022
Pith/arXiv arXiv 2022
-
[57]
L. He, Y . Chen, Y . Dong, Y . Wang, Z. Lin, et al. Efficient equivariant network.Advances in Neural Information Processing Systems, 34:5290–5302, 2021
2021
-
[58]
X. Zhu, D. Wang, O. Biza, G. Su, R. Walters, and R. Platt. Sample efficient grasp learning using equivariant models.Proceedings of Robotics: Science and Systems (RSS), 2022
2022
-
[59]
X. Zhu, D. Wang, G. Su, O. Biza, R. Walters, and R. Platt. On robot grasp learning using equivariant models.Autonomous Robots, 47(8):1175–1193, 2023
2023
-
[60]
Huang, D
H. Huang, D. Wang, X. Zhu, R. Walters, and R. Platt. Edge grasp network: A graph-based se (3)-invariant approach to grasp detection. In2023 IEEE International Conference on Robotics and Automation (ICRA), pages 3882–3888. IEEE, 2023
2023
-
[61]
B. Hu, X. Zhu, D. Wang, Z. Dong, H. Huang, C. Wang, R. Walters, and R. Platt. Orbitgrasp: SE (3)-equivariant grasp learning. In8th Annual Conference on Robot Learning, 2024
2024
-
[62]
J. Yang, C. Deng, J. Wu, R. Antonova, L. Guibas, and J. Bohg. Equivact: Sim (3)-equivariant visuomotor policies beyond rigid object manipulation. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 9249–9255. IEEE, 2024
2024
-
[63]
L. Zhao, X. Zhu, L. Kong, R. Walters, and L. L. Wong. Integrating symmetry into differentiable planning with steerable convolutions.arXiv preprint arXiv:2206.03674, 2022
Pith/arXiv arXiv 2022
-
[64]
L. Zhao, H. Li, T. Padır, H. Jiang, and L. L. Wong. E(2) equivariant graph planning for navigation.IEEE Robotics and Automation Letters, 2024
2024
-
[65]
D. Wang, R. Walters, X. Zhu, and R. Platt. Equivariant Q Learning in Spatial Action Spaces. In 5th Annual Conference on Robot Learning, 2021
2021
-
[66]
Simeonov, Y
A. Simeonov, Y . Du, A. Tagliasacchi, J. B. Tenenbaum, A. Rodriguez, P. Agrawal, and V . Sitz- mann. Neural descriptor fields: Se (3)-equivariant object representations for manipulation. In 2022 International Conference on Robotics and Automation (ICRA), pages 6394–6400. IEEE, 2022. 12
2022
-
[67]
Simeonov, Y
A. Simeonov, Y . Du, Y .-C. Lin, A. R. Garcia, L. P. Kaelbling, T. Lozano-P´erez, and P. Agrawal. Se (3)-equivariant relational rearrangement with neural descriptor fields. InConference on Robot Learning, pages 835–846. PMLR, 2023
2023
-
[68]
H. Huang, D. Wang, R. Walters, and R. Platt. Equivariant Transporter Network. InProceedings of Robotics: Science and Systems, New York City, NY , USA, June 2022. doi:10.15607/RSS. 2022.XVIII.007
doi:10.15607/rss 2022
-
[69]
Huang, D
H. Huang, D. Wang, A. Tangri, R. Walters, and R. Platt. Leveraging symmetries in pick and place.The International Journal of Robotics Research, page 02783649231225775, 2024
2024
-
[70]
Huang, O
H. Huang, O. L. Howell, D. Wang, X. Zhu, R. Platt, and R. Walters. Fourier transporter: Bi-equivariant robotic manipulation in 3d. InThe Twelfth International Conference on Learning Representations, 2024. URLhttps://openreview.net/forum?id=UulwvAU1W0
2024
-
[71]
H. Ryu, H.-i. Lee, J.-H. Lee, and J. Choi. Equivariant descriptor fields: Se (3)-equivariant energy-based models for end-to-end visual robotic manipulation learning.arXiv preprint arXiv:2206.08321, 2022
Pith/arXiv arXiv 2022
-
[72]
H. Ryu, J. Kim, J. Chang, H. S. Ahn, J. Seo, T. Kim, J. Choi, and R. Horowitz. Diffusion-edfs: Bi-equivariant denoising generative modeling on se (3) for visual robotic manipulation.arXiv preprint arXiv:2309.02685, 2023
Pith/arXiv arXiv 2023
-
[73]
B. Eisner, Y . Yang, T. Davchev, M. Vecerik, J. Scholz, and D. Held. Deep se (3)-equivariant geometric reasoning for precise placement tasks.arXiv preprint arXiv:2404.13478, 2024
Pith/arXiv arXiv 2024
-
[74]
D. Wang, R. Walters, and R. Platt. SO(2)-equivariant reinforcement learning. InInternational Conference on Learning Representations, 2022. URL https://openreview.net/forum? id=7F9cOhdvfk_
2022
-
[75]
M. Jia, D. Wang, G. Su, D. Klee, X. Zhu, R. Walters, and R. Platt. Seil: Simulation-augmented equivariant imitation learning. In2023 IEEE International Conference on Robotics and Au- tomation (ICRA), pages 1845–1851. IEEE, 2023
2023
-
[76]
D. Wang, M. Jia, X. Zhu, R. Walters, and R. Platt. On-robot learning with equivariant models. In 6th Annual Conference on Robot Learning, 2022. URLhttps://openreview.net/forum? id=K8W6ObPZQyh
2022
-
[77]
S. Liu, M. Xu, P. Huang, X. Zhang, Y . Liu, K. Oguchi, and D. Zhao. Continual Vision-based Reinforcement Learning with Group Symmetries. InConference on Robot Learning, pages 222–240. PMLR, 2023
2023
-
[78]
C. Kohler, A. S. Srikanth, E. Arora, and R. Platt. Symmetric models for visual force policy learning.arXiv preprint arXiv:2308.14670, 2023
Pith/arXiv arXiv 2023
-
[79]
H. H. Nguyen, A. Baisero, D. Klee, D. Wang, R. Platt, and C. Amato. Equivariant reinforcement learning under partial observability. InConference on Robot Learning, pages 3309–3320. PMLR, 2023
2023
-
[80]
H. Nguyen, T. Kozuno, C. C. Beltran-Hernandez, and M. Hamaya. Symmetry-aware reinforce- ment learning for robotic assembly under partial observability with a soft wrist.arXiv preprint arXiv:2402.18002, 2024
Pith/arXiv arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.