REVIEW 4 major objections 79 references
OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies
T0 review · 4 major / 0 minor · reviewed 2026-07-12 · grok-4.5
Pith's one-line read OmniTacTune lifts weak visual robot policies to high contact-rich success by learning tactile residual corrections in 40–80 minutes of real-world practice.
desk verdict Solid systems paper: residual tactile RL on frozen visual policies works across four real contact tasks and several base policies; the efficiency claim is real but rests on small-N human-labeled success rates. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two-stage tactile residual RL: Stage 1 warm-starts a flow-tactile critic and tactile encoder from frozen base-policy rollouts (contact-only encoder updates plus trajectory-level force-conditioned tactile augmentation); Stage 2 freezes the base and trains a residual actor a = a_base + s_t a_residual, gated by contact, conditioned on proprioception, object-centric flow goals, tactile features, and the base action chunk, with SAC under a multi-sensory reward of reaching, grasp, flow-subgoal, and safety terms.
What would settle it
On the same four contact-rich tasks and wall-clock budgets, residual training after the paper’s warm-start fails to raise final success above the visual-only residual baseline and above the frozen base across several base-policy architectures and tactile encodings.
Extended reading notes
Core claim
OmniTacTune shows that tactile sensing can be adapted to frozen, architecturally diverse visual base policies through residual real-world RL without offline tactile demonstrations. Autonomous base-policy rollouts warm-start a flow-tactile critic and task-adapted tactile encoder; online residual RL then learns lightweight contact-aware corrections on top of the base actions under object-centric multi-sensory reward shaping. Across four contact-rich real-world tasks this raises success from 5–40% to 85–100% in 40–80 minutes and generalizes across base policies and tactile representations.
Load-bearing premise
Autonomous rollouts of a still-weak visual policy, plus synthetic tactile augmentation and a hand-designed multi-sensory reward with human success labels, are enough to bootstrap a stable critic and residual policy without offline tactile demos.
Editorial extensions
If this is right
- Scalable visual policies can be made contact-aware without retraining them or collecting large paired visuo-tactile datasets.
- Tens of minutes of real-world residual practice can close the last-mile contact gap that pure imitation leaves open.
- One residual interface can attach to flow, ACT, diffusion, and VLA-style bases, so tactile adaptation need not be architecture-specific.
- Both compact marker signals and pretrained tactile image encoders can serve as the tactile stream for residual correction.
- Warm-start of critic and tactile encoder, plus multi-sensory reward shaping, is material to sample-efficient tactile residual RL.
Reading between the lines
- If the base policy rarely reaches near-contact states, residual learning can stall; the method implicitly needs a prior that already lands in a useful contact neighborhood.
- The vision-for-planning, touch-for-refinement split may extend to multi-finger dexterity and bimanual assembly where tactile data remain scarce relative to vision.
- Manual resets and human terminal success labels remain practical bottlenecks; automating both would be a direct stress test of the same residual recipe.
- Dynamic levering failures (slip, edge miss, pose tilt) suggest residual scale and force-aware rewards may need richer contact profiles than the current scheduler alone provides.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. OmniTacTune proposes a policy-agnostic two-stage real-world RL pipeline that adapts tactile feedback to frozen visual base policies via residual correction, without offline tactile demonstrations. Stage 1 warm-starts a flow-tactile critic and tactile encoder from autonomous base-policy rollouts (with ControlTac trajectory-level tactile augmentation); Stage 2 learns a lightweight residual actor that adds contact-gated corrections on top of the base action chunk, guided by a multi-sensory reward combining object-centric flow subgoals, tactile grasp/safety terms, and human terminal success labels. On four contact-rich tasks (peg-in-hole, charger insertion, cap opening, box opening), the method reports improving base success from 5–40% to 85–100% in 40–80 minutes, with generalization across base policies (human/teleop flow, ACT, DP, π0.5) and tactile representations (AnyTouch2, Sparsh, T3, markers), outperforming adapted PLD and ViTAL baselines and visuo-tactile imitation under a matched time budget.
Significance. If the reported efficiency and generality hold under stronger evaluation, this is a practically important contribution: it offers a concrete path to attach tactile residual practice to scalable visual priors rather than collecting large paired visuo-tactile datasets. Strengths include real hardware results on four genuinely contact-rich tasks, systematic compatibility tests across policy architectures and tactile encoders, relevant residual-RL and imitation baselines, and ablations of reward components, residual design, warm-start, and action scaling (Sec. 4, App. C). The residual interface (shared flow goals + base action chunks + contact gate) is a clean systems idea for policy-agnostic adaptation. These are falsifiable empirical claims with public project-page materials, not circular constructions.
major comments (4)
- Table 1 and Fig. 4 (also Sec. 4.1): the central efficiency claim (5–40% → 85–100% in 40–80 min) rests on final success from 20 trials and intermediate checkpoints from 10 trials, with no variance, confidence intervals, multi-seed runs, or multi-operator evaluation. For Cap Opening and Box Opening the base starts at 5%, so absolute margins over PLD*/ViTAL are large but statistically under-specified. Please report binomial CIs or bootstrap intervals at minimum, and preferably repeated training runs or denser evaluation; without this the headline numbers are not fully secured.
- Sec. 3.4 and App. A.6: terminal success is a human-assigned reward of 1, and the same operator performs resets. This couples the learning signal and the evaluation metric. Clarify whether evaluation success uses the same human judgment as training, whether any automatic success detector exists, and how label consistency was controlled across methods. If evaluation is fully human-labeled, consider a blinded protocol or automatic geometric/contact criteria so reported gains cannot be attributed to label drift.
- Sec. 3.3–3.4 and the weakest operating regime (5% base policies): residual learning assumes warm-start rollouts plus flow/tactile shaping produce enough near-contact experience for a usable critic. The paper shows warm-start ablations (Fig. 17) but does not quantify contact-state coverage or success of warm-start trajectories on Cap/Box Opening. Please report how often base rollouts reach contact/near-goal states, and discuss failure modes when the base rarely enters the residual’s useful region—this is load-bearing for the claim of no offline tactile demos.
- Sec. 4.3 / Table 2: imitation baselines receive 40 extra teleop demos under a 50-minute budget, while OmniTacTune uses online interaction with human terminal labels and dense shaped rewards. The comparison is informative but not fully matched in supervision type. Explicitly state what human effort (resets, success labeling, interventions) is required for OmniTacTune versus teleop collection, so data-efficiency claims are not overstated relative to pure demonstration methods.
Circularity Check
No significant circularity: empirical real-world RL systems paper; success rates are measured outcomes, not quantities forced by definition or self-citation.
full rationale
OmniTacTune is a methods-and-experiments robotics paper. Its load-bearing claim is measured task success (5–40% → 85–100% in 40–80 min on four contact-rich tasks; Table 1, Fig. 4), obtained from real robot trials after residual SAC training. That metric is not defined in terms of the reward weights, residual scale schedule, ControlTac force perturbations, or flow subgoals; those are free design choices that could have failed. Warm-start (autonomous base rollouts + critic/encoder bootstrap) and residual action a = a_base + s·a_residual are standard residual-RL constructions, not self-definitional identities that force high success. Self-citations (GenFlowRL/Im2Flow2Act for object flow, ControlTac for tactile augmentation, PLD/ViTAL as baselines) supply components or comparison points; they do not import a uniqueness theorem or rename a known empirical law as a new prediction. There is no fitted constant re-presented as a first-principles forecast, and no derivation chain that reduces the headline result to its inputs by construction. Evaluation sparsity and human terminal labels are statistical/correctness concerns, not circularity.
Assumptions & free parameters
free parameters (4)
- Reward weights (wr, wg, wf, ws) =
1.0 / 0.5 / 2.0 / 1.0
- Residual action scale schedule st =
0.05 to 0.15
- Contact/safety thresholds (εcontact, εsafety, εdepth, εflow) =
e.g. 1.5 px, 8.0 px, 0.10, 0.03
- Warm-start duration and ControlTac force range =
12 min; Δf in [-3,10]
assumptions (4)
- domain assumption Frozen visual base policies provide useful near-contact motion priors that residual tactile corrections can refine rather than replace.
- domain assumption Object-centric keypoint flow from a fine-tuned generator yields a dense, embodiment-agnostic reward and shared residual conditioning across policy architectures.
- domain assumption SAC with twin critics, contact-gated tactile features, and reconstruction-regularized encoder adaptation is a valid online learner for residual actions on hardware.
- ad hoc to paper Marker displacement / tactile depth thresholds detect contact and unsafe force reliably enough for gating and safety resets.
invented entities (2)
-
OmniTacTune two-stage tactile residual RL pipeline
-
Object-centric multi-sensory reward (flow + tactile grasp/safety)
Cite this review
Pith. "Pith review of OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies." pith.science (2026). https://pith.science/paper/7WBPX5ZN
@misc{pith2026260703723,
author = {Pith},
title = {Pith review of: OmniTacTune: Policy-Agnostic Real-World RL for Tactile Residual Adaptation of Visual Policies},
year = {2026},
howpublished = {\url{https://pith.science/paper/7WBPX5ZN}},
note = {Machine review of arXiv:2607.03723}
}
read the original abstract
Visual policies learned from human videos, teleoperation, and robot demonstrations offer scalable motion priors, but often fail in contact-rich manipulation, where success significantly depends on local force and contact geometry. Tactile sensing provides these complementary signals, yet tactile data remain costly to collect and hard to generalize across sensors, robots, and tasks. We introduce OmniTacTune, a policy-agnostic real-world RL pipeline that adapts tactile feedback to pretrained visual policies through residual correction. OmniTacTune uses a two-stage design: it first bootstraps tactile-aware learning from autonomous base-policy rollouts, then learns a lightweight tactile residual policy through online interaction. Extensive experiments show that OmniTacTune generalizes across diverse contact-rich tasks, visual base policies, and tactile representations. Across four real-world contact-rich tasks, it improves visual base policies from 5-40% success to 85-100% within 40-80 minutes, demonstrating an efficient path for adapting tactile feedback to scalable visual robot policies. Project page: https://colinyu1.github.io/omnitactune-site/
Figures
Figures from the paper (18 more)
Reference graph
Works this paper leans on
-
[1]
E. Collaboration, A. O’Neill, A. Rehman, A. Gupta, A. Maddukuri, A. Gupta, A. Padalkar, A. Lee, A. Pooley, A. Gupta, A. Mandlekar, A. Jain, A. Tung, A. Bewley, A. Herzog, A. Ir- pan, A. Khazatsky, A. Rai, A. Gupta, A. Wang, A. Kolobov, A. Singh, A. Garg, A. Kembhavi, A. Xie, A. Brohan, A. Raffin, A. Sharma, A. Yavary, A. Jain, A. Balakrishna, A. Wahid, B....
arXiv 2025
-
[2]
A. Khazatsky, K. Pertsch, S. Nair, A. Balakrishna, S. Dasari, S. Karamcheti, S. Nasiriany, M. K. Srirama, L. Y . Chen, K. Ellis, P. D. Fagan, J. Hejna, M. Itkina, M. Lepert, Y . J. Ma, P. T. Miller, J. Wu, S. Belkhale, S. Dass, H. Ha, A. Jain, A. Lee, Y . Lee, M. Memmel, S. Park, I. Radosavovic, K. Wang, A. Zhan, K. Black, C. Chi, K. B. Hatch, S. Lin, J. ...
arXiv 2025
- [3]
- [4]
-
[5]
R. Punamiya, S. Kareer, Z. Liu, J. Citron, R.-Z. Qiu, X. Cai, A. Gavryushin, J. Chen, D. Li- conti, L. Y . Zhu, P. Aphiwetsa, B. Li, A. Cheluva, P. Kuppili, Y . Liu, D. Patel, A. Gao, H.-Y . Chung, R. Co, R. Zbizika, J. Liu, X. Xu, H. Xiong, G. Chen, S. Oliani, C. Yang, X. Wang, J. Fort, R. Newcombe, J. Gao, J. Chong, G. Matsuda, A. Doriwala, M. Pollefeys...
arXiv 2026
-
[6]
W. Yuan, S. Dong, and E. H. Adelson. Gelsight: High-resolution robot tactile sensors for estimating geometry and force.Sensors, 17(12):2762, 2017
2017
-
[7]
K. Yu, Y . Han, Q. Wang, V . Saxena, D. Xu, and Y . Zhao. Mimictouch: Leveraging multi- modal human tactile demonstrations for contact-rich manipulation, 2025. URLhttps:// arxiv.org/abs/2310.16917
arXiv 2025
- [8]
Show all 79 references
-
[9]
H. Li, Y . Zhang, J. Zhu, S. Wang, M. A. Lee, H. Xu, E. Adelson, L. Fei-Fei, R. Gao, and J. Wu. See, hear, and feel: Smart sensory fusion for robotic manipulation, 2022. URLhttps: //arxiv.org/abs/2212.03858
2022 arXiv
-
[10]
H. Xue, J. Ren, W. Chen, G. Zhang, Y . Fang, G. Gu, H. Xu, and C. Lu. Reactive diffu- sion policy: Slow-fast visual-tactile policy learning for contact-rich manipulation, 2025. URL https://arxiv.org/abs/2503.02881
2025 arXiv
-
[11]
Y . Dou, F. Yang, Y . Liu, A. Loquercio, and A. Owens. Tactile-augmented radiance fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 26529–26539, 2024
2024
-
[12]
X. Zhu, B. Huang, and Y . Li. Touch in the wild: Learning fine-grained manipulation with a portable visuo-tactile gripper. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. URLhttps://openreview.net/forum?id=WabVVQKTUF
2025
-
[13]
L. Fu, G. Datta, H. Huang, W. C.-H. Panitch, J. Drake, J. Ortiz, M. Mukadam, M. Lam- beta, R. Calandra, and K. Goldberg. A touch, vision, and language dataset for multi- modal alignment. InForty-first International Conference on Machine Learning, 2024. URL https://openreview.n...
2024
-
[14]
R. Gao, Y . Dou, H. Li, T. Agarwal, J. Bohg, Y . Li, L. Fei-Fei, and J. Wu. The object- folder benchmark: Multisensory learning with neural and real objects. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17276–17286, 2023
2023
-
[15]
F. Liu, C. Li, Y . Qin, J. Xu, P. Abbeel, and R. Chen. Vitamin: Learning contact-rich tasks through robot-free visuo-tactile manipulation interface, 2025. URLhttps://arxiv.org/ abs/2504.06156
2025 arXiv
-
[16]
Hoque, P
R. Hoque, P. Huang, D. J. Yoon, M. Sivapurapu, and J. Zhang. Egodex: Learning dexter- ous manipulation from large-scale egocentric video, 2026. URLhttps://arxiv.org/abs/ 2505.11709. 10
2026 arXiv
-
[17]
Huang, J
B. Huang, J. Xu, I. Akinola, W. Yang, B. Sundaralingam, R. O’Flaherty, D. Fox, X. Wang, A. Mousavian, Y .-W. Chao, and Y . Li. VT-refine: Learning bimanual assembly with visuo- tactile feedback via simulation fine-tuning. In9th Annual Conference on Robot Learning,
-
[18]
URLhttps://openreview.net/forum?id=bOVF8Rj33i
-
[19]
P. Dan, K. Kedia, A. Chao, E. W. Duan, M. A. Pace, W.-C. Ma, and S. Choudhury. X-sim: Cross-embodiment learning via real-to-sim-to-real, 2025. URLhttps://arxiv.org/abs/ 2505.07096
2025
-
[20]
W. Wan, J. Fu, X. Yuan, Y . Zhu, and H. Su. Lodestar: Long-horizon dexterity via synthetic data augmentation from human demonstrations, 2025. URLhttps://arxiv.org/abs/2508. 17547
2025
-
[21]
S. Ross, G. J. Gordon, and J. A. Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning, 2011. URLhttps://arxiv.org/abs/1011.0686
2011 arXiv
-
[22]
J. Luo, C. Xu, J. Wu, and S. Levine. Precise and dexterous robotic manipulation via human- in-the-loop reinforcement learning, 2025. URLhttps://arxiv.org/abs/2410.21845
2025 arXiv
-
[23]
W. Xiao, H. Lin, A. Peng, H. Xue, T. He, Y . Xie, F. Hu, J. Wu, Z. Luo, L. J. Fan, G. Shi, and Y . Zhu. Self-improving vision-language-action models with data generation via residual rl, 2025. URLhttps://arxiv.org/abs/2511.00091
2025
-
[24]
Z. Zhou, A. Peng, Q. Li, S. Levine, and A. Kumar. Efficient online reinforcement learning fine-tuning need not retain offline data, 2025. URLhttps://arxiv.org/abs/2412.07762
2025 arXiv
-
[25]
C. Xu, J. T. Springenberg, M. Equi, A. Amin, A. Esmail, S. Levine, and L. Ke. Rl token: Bootstrapping online rl with vision-language-action models, 2026. URLhttps://arxiv. org/abs/2604.23073
2026 arXiv
-
[26]
T. Z. Zhao, V . Kumar, S. Levine, and C. Finn. Learning fine-grained bimanual manipulation with low-cost hardware, 2023. URLhttps://arxiv.org/abs/2304.13705
2023 arXiv
-
[27]
C. Chi, Z. Xu, S. Feng, E. Cousineau, Y . Du, B. Burchfiel, R. Tedrake, and S. Song. Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 44(10-11):1684–1704, 2025
2025
-
[28]
M. Xu, Z. Xu, Y . Xu, C. Chi, G. Wetzstein, M. Veloso, and S. Song. Flow as the cross-domain manipulation interface, 2024. URLhttps://arxiv.org/abs/2407.15208
2024 arXiv
-
[29]
Intelligence, K
P. Intelligence, K. Black, N. Brown, J. Darpinian, K. Dhabalia, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, M. Y . Galliker, D. Ghosh, L. Groom, K. Hausman, B. Ichter, S. Jakubczak, T. Jones, L. Ke, D. LeBlanc, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, A....
2025 arXiv
-
[30]
Intelligence, A
P. Intelligence, A. Amin, R. Aniceto, A. Balakrishna, K. Black, K. Conley, G. Connors, J. Darpinian, K. Dhabalia, J. DiCarlo, D. Driess, M. Equi, A. Esmail, Y . Fang, C. Finn, C. Glos- sop, T. Godden, I. Goryachev, L. Groom, H. Hancock, K. Hausman, G. Hussein, B. Ichter, S. Ja...
2025 arXiv
-
[31]
M. J. Kim, K. Pertsch, S. Karamcheti, T. Xiao, A. Balakrishna, S. Nair, R. Rafailov, E. Foster, G. Lam, P. Sanketi, Q. Vuong, T. Kollar, B. Burchfiel, R. Tedrake, D. Sadigh, S. Levine, P. Liang, and C. Finn. Openvla: An open-source vision-language-action model, 2024. URL https...
2024 arXiv
-
[32]
L. Y . Zhu, P. Kuppili, R. Punamiya, P. Aphiwetsa, D. Patel, S. Kareer, S. Ha, and D. Xu. Emma: Scaling mobile manipulation via egocentric human data, 2025. URLhttps://arxiv.org/ abs/2509.04443
2025
-
[33]
Y . Liu, W. C. Shin, Y . Han, Z. Chen, H. Ravichandar, and D. Xu. Immimic: Cross-domain imitation from human videos via mapping and interpolation, 2025. URLhttps://arxiv. org/abs/2509.10952
2025 arXiv
-
[34]
K. Yu, S. Zhang, H. Soora, F. Huang, H. Huang, P. Tokekar, and R. Gao. Genflowrl: Shap- ing rewards with generative object-centric flow in visual reinforcement learning, 2025. URL https://arxiv.org/abs/2508.11049
2025 arXiv
-
[35]
Y . Han, Z. Chen, K. A. Williams, and H. Ravichandar. Learning prehensile dexterity by im- itating and emulating state-only observations.IEEE Robotics and Automation Letters, 9(10): 8266–8273, 2024
2024
-
[36]
Z. Wang, B. He, K. Yu, S. Lee, R. Gao, F. Huang, and Y . Aloimonos. Humanego: Zero-shot robot learning from minutes of human egocentric videos, 2026
2026
-
[37]
Huang, C
W. Huang, C. Wang, Y . Li, R. Zhang, and L. Fei-Fei. Rekep: Spatio-temporal reasoning of relational keypoint constraints for robotic manipulation.arXiv preprint arXiv:2409.01652, 2024
2024 arXiv
-
[38]
Singh, K
A. Singh, K. Torshizi, K. Habib, K. Yu, R. Gao, and P. Tokekar. Afford2act: Affordance- guided automatic keypoint selection for generalizable and lightweight robotic manipulation,
-
[39]
URLhttps://arxiv.org/abs/2510.01433
-
[40]
S. Bahl, A. Gupta, and D. Pathak. Human-to-robot imitation in the wild, 2022. URLhttps: //arxiv.org/abs/2207.09450
2022 arXiv
-
[41]
Bhirangi, V
R. Bhirangi, V . Pattabiraman, E. Erciyes, Y . Cao, T. Hellebrekers, and L. Pinto. Anyskin: Plug- and-play skin sensing for robotic touch, 2024. URLhttps://arxiv.org/abs/2409.08276
2024 arXiv
-
[42]
Lambeta, P.-W
M. Lambeta, P.-W. Chou, S. Tian, B. Yang, B. Maloon, V . R. Most, D. Stroud, R. Santos, A. Byagowi, G. Kammerer, D. Jayaraman, and R. Calandra. Digit: A novel design for a low- cost compact high-resolution tactile sensor with application to in-hand manipulation.IEEE Robotics a...
2020 doi
-
[43]
Kim and A
S. Kim and A. Rodriguez. Active extrinsic contact sensing: Application to general peg-in-hole insertion, 2022. URLhttps://arxiv.org/abs/2110.03555
2022 arXiv
-
[44]
Y . Han, K. Yu, R. Batra, N. Boyd, C. Mehta, T. Zhao, Y . She, S. Hutchinson, and Y . Zhao. Learning generalizable vision-tactile robotic grasping strategy for deformable objects via trans- former.IEEE/ASME Transactions on Mechatronics, 30(1):554–566, 2025. doi:10.1109/ TMECH....
2025
-
[45]
Xu and Y
Z. Xu and Y . She. Letac-mpc: Learning model predictive control for tactile-reactive grasping,
-
[46]
URLhttps://arxiv.org/abs/2403.04934
-
[47]
Z.-H. Yin, B. Huang, Y . Qin, Q. Chen, and X. Wang. Rotating without seeing: Towards in-hand dexterity through touch, 2023. URLhttps://arxiv.org/abs/2303.10880
2023 arXiv
-
[48]
Guzey, B
I. Guzey, B. Evans, S. Chintala, and L. Pinto. Dexterity from touch: Self-supervised pre- training of tactile representations with robotic play, 2023. URLhttps://arxiv.org/abs/ 2303.12076
2023 arXiv
-
[49]
J. J. Liu, Y . Li, K. Shaw, T. Tao, R. Salakhutdinov, and D. Pathak. Factr: Force-attending curriculum training for contact-rich policy learning, 2025. URLhttps://arxiv.org/abs/ 2502.17432. 12
2025 arXiv
-
[50]
Sferrazza, Y
C. Sferrazza, Y . Seo, H. Liu, Y . Lee, and P. Abbeel. The power of the senses: Generalizable manipulation from vision and touch through masked multimodal learning, 2023. URLhttps: //arxiv.org/abs/2311.00924
2023 arXiv
-
[51]
Kelly, C
M. Kelly, C. Sidrane, K. Driggs-Campbell, and M. J. Kochenderfer. Hg-dagger: Interactive imitation learning with human experts, 2019. URLhttps://arxiv.org/abs/1810.02890
2019 arXiv
-
[52]
X. Xu, Y . Hou, Z. Liu, and S. Song. Compliant residual DAgger: Improving real-world contact- rich manipulation with human corrections. InThe Thirty-ninth Annual Conference on Neural Information Processing Systems (NeurIPS), 2025. URLhttps://openreview.net/forum? id=cjcm5LYVWm
2025
-
[53]
Z. Chen, S. Chen, E. Arlaud, I. Laptev, and C. Schmid. Vividex: Learning vision-based dexter- ous manipulation from human videos, 2025. URLhttps://arxiv.org/abs/2404.15709
2025 arXiv
-
[54]
Levine, A
S. Levine, A. Kumar, G. Tucker, and J. Fu. Offline reinforcement learning: Tutorial, review, and perspectives on open problems, 2020. URLhttps://arxiv.org/abs/2005.01643
2020 arXiv
-
[55]
Kostrikov, A
I. Kostrikov, A. Nair, and S. Levine. Offline reinforcement learning with implicit q-learning,
-
[56]
URLhttps://arxiv.org/abs/2110.06169
-
[57]
Nakamoto, Y
M. Nakamoto, Y . Zhai, A. Singh, M. S. Mark, Y . Ma, C. Finn, A. Kumar, and S. Levine. Cal-ql: Calibrated offline rl pre-training for efficient online fine-tuning, 2024. URLhttps: //arxiv.org/abs/2303.05479
2024 arXiv
-
[58]
P. J. Ball, L. Smith, I. Kostrikov, and S. Levine. Efficient online reinforcement learning with offline data, 2023. URLhttps://arxiv.org/abs/2302.02948
2023 arXiv
-
[59]
Ankile, A
L. Ankile, A. Simeonov, I. Shenfeld, M. Torne, and P. Agrawal. From imitation to refinement – residual rl for precise assembly, 2024. URLhttps://arxiv.org/abs/2407.16677
2024 arXiv
-
[60]
Haldar, J
S. Haldar, J. Pari, A. Rai, and L. Pinto. Teach a robot to fish: Versatile imitation from one minute of demonstrations, 2023. URLhttps://arxiv.org/abs/2303.01497
2023 arXiv
-
[61]
Guzey, Y
I. Guzey, Y . Dai, G. Savva, R. Bhirangi, and L. Pinto. Bridging the human to robot dexterity gap through object-oriented rewards, 2024. URLhttps://arxiv.org/abs/2410.23289
2024 arXiv
-
[62]
A. Iyer, Z. Peng, Y . Dai, I. Guzey, S. Haldar, S. Chintala, and L. Pinto. Open teach: A versatile teleoperation system for robotic manipulation, 2024. URLhttps://arxiv.org/abs/2403. 07870
2024
-
[63]
Kuang, S
Y . Kuang, S. Park, K. Fragkiadaki, and S. Tulsiani. Dex4d: Task-agnostic point track policy for sim-to-real dexterous manipulation, 2026. URLhttps://arxiv.org/abs/2602.15828
2026
-
[64]
Oquab, T
M. Oquab, T. Darcet, T. Moutakanni, H. V o, M. Szafraniec, V . Khalidov, P. Fernandez, D. Haz- iza, F. Massa, A. El-Nouby, M. Assran, N. Ballas, W. Galuba, R. Howes, P.-Y . Huang, S.-W. Li, I. Misra, M. Rabbat, V . Sharma, G. Synnaeve, H. Xu, H. Jegou, J. Mairal, P. Labatut, A...
2024 arXiv
-
[65]
Kirillov, E
A. Kirillov, E. Mintun, N. Ravi, H. Mao, C. Rolland, L. Gustafson, T. Xiao, S. Whitehead, A. C. Berg, W.-Y . Lo, P. Doll ´ar, and R. Girshick. Segment anything, 2023. URLhttps: //arxiv.org/abs/2304.02643
2023 arXiv
-
[66]
Karaev, I
N. Karaev, I. Makarov, J. Wang, N. Neverova, A. Vedaldi, and C. Rupprecht. Cotracker3: Simpler and better point tracking by pseudo-labelling real videos, 2024. URLhttps: //arxiv.org/abs/2410.11831. 13
2024 arXiv
-
[67]
Z. Zhao, S. Haldar, J. Cui, L. Pinto, and R. Bhirangi. Touch begins where vision ends: General- izable policies for contact-rich manipulation, 2025. URLhttps://arxiv.org/abs/2506. 13762
2025
-
[68]
R. Feng, Y . Zhou, S. Mei, D. Zhou, P. Wang, S. Cui, B. Fang, G. Yao, and D. Hu. Anytouch 2: General optical tactile representation learning for dynamic tactile perception, 2026. URL https://arxiv.org/abs/2602.09617
2026
-
[69]
Yarats, R
D. Yarats, R. Fergus, A. Lazaric, and L. Pinto. Mastering visual continuous control: Im- proved data-augmented reinforcement learning, 2021. URLhttps://arxiv.org/abs/ 2107.09645
2021 arXiv
-
[70]
D. Luo, K. Yu, A.-H. Shahidzadeh, C. Ferm ¨uller, Y . Aloimonos, and R. Gao. Controltac: Force- and position-controlled tactile data augmentation with a single reference image, 2025. URLhttps://arxiv.org/abs/2505.20498
2025 arXiv
-
[71]
Huang and Y
B. Huang and Y . Li. Flexitac: A low-cost, open-source, scalable tactile sensing solution for robotic systems, 2026. URLhttps://arxiv.org/abs/2604.28156
2026 arXiv
-
[72]
Higuera, A
C. Higuera, A. Sharma, C. K. Bodduluri, T. Fan, P. Lancaster, M. Kalakrishnan, M. Kaess, B. Boots, M. Lambeta, T. Wu, and M. Mukadam. Sparsh: Self-supervised touch representa- tions for vision-based tactile sensing, 2024. URLhttps://arxiv.org/abs/2410.24090
2024 arXiv
-
[73]
J. Zhao, Y . Ma, L. Wang, and E. H. Adelson. Transferable tactile transformers for representa- tion learning across diverse sensors and tasks, 2024. URLhttps://arxiv.org/abs/2406. 13640
2024
-
[74]
Y . Guo, T. Lee, L. X. Shi, J. Chen, P. Liang, and C. Finn. Vlaw: Iterative co-improvement of vision-language-action policy and world model, 2026. URLhttps://arxiv.org/abs/ 2602.12063
2026
-
[75]
S. Yang, Y . Du, K. Ghasemipour, J. Tompson, L. Kaelbling, D. Schuurmans, and P. Abbeel. Learning interactive real-world simulators, 2024. URLhttps://arxiv.org/abs/2310. 06114
2024
-
[76]
Hafner, J
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap. Mastering diverse domains through world models, 2024. URLhttps://arxiv.org/abs/2301.04104
2024 arXiv
-
[77]
E. J. Hu, Y . Shen, P. Wallis, Z. Allen-Zhu, Y . Li, S. Wang, L. Wang, and W. Chen. Lora: Low- rank adaptation of large language models, 2021. URLhttps://arxiv.org/abs/2106. 09685
2021
-
[78]
A. . Team, J. Aldaco, T. Armstrong, R. Baruch, J. Bingham, S. Chan, K. Draper, D. Dwibedi, C. Finn, P. Florence, S. Goodrich, W. Gramlich, T. Hage, A. Herzog, J. Hoech, T. Nguyen, I. Storz, B. Tabanpour, L. Takayama, J. Tompson, A. Wahid, T. Wahrburg, S. Xu, S. Yaroshenko, K. ...
2024 arXiv
-
[79]
insert the peg into the hole
N. R. Arachchige, Z. Chen, W. Jung, W. C. Shin, R. Bansal, P. Barroso, Y . H. He, Y . C. Lin, B. Joffe, S. Kousik, and D. Xu. Sail: Faster-than-demonstration execution of imitation learning policies, 2025. URLhttps://arxiv.org/abs/2506.11948. 14 Appendix A Implementation Detai...
2025 arXiv
Reviewed July 12, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.