REVIEW 4 major objections 6 minor 49 references
OPA-Pack: Object-Property-Aware Robotic Bin Packing
T0 review · 4 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read OPA-Pack is the first object-property-aware robotic bin packing framework, raising incompatible-pair avoidance from 52% to 95% and cutting pressure on fragile objects by 29.4% while keeping compactness roughly unchanged.
desk verdict First property-aware bin packing system with real-robot validation, but the headline safety gains are tied to the same ChatGPT-4 labels that drive the training rewards. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central machinery is the OPA-Net packing policy built on dueling deep Q-learning. It has four encoders: a point-cloud shape encoder, a pose encoder, a property-embedding layer, and convolutional encoders over three container heightmaps—occupancy, fragility, and avoidance. The fragility heightmap records where fragile objects already sit so the placement predictor keeps heavy items off them; the avoidance heightmap is rendered per candidate object to mark locations close to that object's incompatible pairs. The property-embedding layer is what lets the network tell a steel hammer from a wax candle at decision time. Training is driven by the OPA reward, a weighted sum of compactness, a fragility penalty proportional to the number of fragile objects covered times the object's estimated weight, and a binary avoidance penalty.
What would settle it
Have several independent human raters annotate fragility and density for a random sample of the 1,032 OPA objects and compare against the model's labels; if agreement is low on multi-material objects like the drill, screwdriver, and vial, then the reported 29.4% pressure reduction is being measured against labels that may not reflect real fragility. A second check is to rerun the 200-sequence benchmark with human expert labels as ground truth; if the avoidance advantage over the geometry baseline collapses, the improvement is an artifact of the label source.
Extended reading notes
Core claim
The paper's central claim is that object properties belong inside the packing decision itself: the long-term value of a packing action should include not just how much space the arrangement uses but whether it puts weight on fragile items and whether it places incompatible pairs side by side. OPA-Net, a dueling deep Q-network, takes as input a property vector for each candidate object plus three container heightmaps—occupancy, fragility, and avoidance—and is trained with a reward that is a weighted sum of compactness, a fragility penalty scaled by the object's estimated weight, and a binary avoidance penalty. In 200 random packing sequences in simulation, the policy separates incompatible pairs in 95.0% of cases versus 52.3% for the geometry-only baseline and cuts mean pressure on fragile objects from 5.07 to 3.58, with compactness of 0.413 versus 0.420. The same network, applied to real-scanned supermarket objects and a physical robot arm, still keeps compactness near the baseline while reducing squeezed fragile objects by more than five times.
Load-bearing premise
The load-bearing premise is that the object property annotations—fragility, density level, and avoidance relations—are correct; they are generated by a vision-language model guided by a hand-written material table and checked by a single human on only 100 of the 1,032 objects.
Editorial extensions
If this is right
- Safety can be added to packing without a large space penalty: the property-aware policy stays within about 2% compactness of the geometry-only baseline while improving avoidance and fragility metrics substantially.
- Fragile objects can be protected by learned placement decisions rather than explicit post-hoc rules, cutting pressure on fragile objects by 29.4% in simulation and squeezed fragile objects by more than five times on the physical platform.
- Object property labels can be produced at scale by a vision-language model combined with retrieval-augmented generation and chain-of-thought reasoning, yielding a released dataset of 1,032 annotated everyday objects.
- The safety-efficiency trade-off is controllable through the reward weights: raising the fragility penalty further lowers pressure but at a noticeable cost in compactness.
- The policy transfers from synthetic objects to real-scanned supermarket items, so the learned property-aware behavior is not confined to the training set's meshes.
Reading between the lines
- Inference: the four avoidance relations (sharp-soft, medicine-edible, chemical-edible, ignition-flammable) form a small reusable ontology; new pairwise constraints such as odor or moisture sensitivity could likely be added by extending the property vector and avoidance heightmap without changing the network or training loop.
- Inference: the explicit reward weights mean the packing preference could be adjusted at test time by reweighting, which is exactly the on-the-fly user-preference modification the paper lists as future work.
- Inference: since the packing policy is trained on the generated labels, the framework's ceiling is set by label accuracy; improving the property recognition stage, or validating it on more than 100 objects with more than one rater, would likely change the safety numbers.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes OPA-Pack, a two-stage robotic bin packing framework that first recognizes object properties (fragility, softness, sharpness, density level, and semantic categories) using a vision-language model with retrieval-augmented generation and chain-of-thought reasoning, and then trains a deep Q-learning packing policy (OPA-Net) with a reward that penalizes placing heavy objects on fragile ones and packing incompatible object pairs closely, while maintaining compactness. The authors contribute a dataset of 1,032 everyday objects with property annotations, report simulation results on 200 random packing sequences showing an improvement in avoidance accuracy from 52.33% to 95.00% and a 29.4% reduction in pressure on fragile objects relative to the IR-BPP baseline, and demonstrate the approach on a physical robot platform over 50 packing sequences.
Significance. If the reported gains are genuine, the paper addresses a real gap in robotic bin packing, which has largely treated packing as a purely geometric problem. The proposed framework, the new dataset, and the real-robot demonstration are useful contributions. The work is also commendable for releasing the dataset and for including ablations of both the property-recognition pipeline and the network design. However, the central claims of improved safety depend critically on the quality of the ChatGPT-4-generated property labels and on the independence of the evaluation metrics from the training reward, both of which currently have important weaknesses. The contribution is therefore significant but not yet fully established.
major comments (4)
- [§V-C, §VII-A, §VII-D] The evaluation metrics Avoid. Acc. and Press. on Frag. are essentially the negative of the reward terms R_avoidance and R_fragility defined in Eqs. (11)-(12), and the DQN is trained to maximize the weighted sum in Eq. (13). Consequently, the reported 42.7% improvement in avoidance accuracy and 29.4% reduction in pressure on fragile objects partly reflect reward optimization rather than an independent measure of real-world safety. I would like to see an evaluation that does not use the same object-property annotations in both the training reward and the evaluation metric, for example by using a manually corrected label set on a held-out subset and recomputing Avoid. Acc. and Press. on Frag. with those corrected labels.
- [§VII-D, Fig. 7] The property-recognition accuracy of the full pipeline is validated on only 100 of the 1,032 objects, by a single human evaluator, and no inter-annotator agreement or confidence intervals are reported. Because the same ChatGPT-4-generated labels define both the training signal and the evaluation metrics, a systematic error on the remaining 932 objects would shift the training and evaluation together, making the headline safety gains potentially an artifact of the label source. The authors should either validate a substantially larger and randomly selected subset with multiple annotators, or report the sensitivity of the packing metrics to plausible label noise.
- [§VII-E, Table IV] The hyperparameters λ and β in Eq. (13) are chosen after inspecting the evaluation-set results (Table IV), with the chosen values (λ=20, β=0.2) reported as the main configuration. This introduces a selection effect: the headline numbers in Table II are produced by a configuration selected on the same data on which it is evaluated. I recommend reporting results on a separate held-out set, or at least disclosing the selection procedure and its potential overfitting effect.
- [§VIII-B, Table VI] The real-robot results are reported as averages over 50 sequences without error bars or statistical significance tests. Since the real-world metrics (# Close Avoid. Pairs and # Squeeze Fragile) also rely on the same ChatGPT-4-based property annotations, the 42.6% reduction in close avoidance pairs and the roughly 5x reduction in squeezed fragile objects need to be accompanied by variance measures and, ideally, independent human verification of at least a subset of the physical outcomes (e.g., visual inspection of whether the persimmon is actually crushed).
minor comments (6)
- [§II] There is a typo on the first line: 'geoemtric' should be 'geometric'.
- [§IV] The text says 'the prediced inner and outer materials' — 'prediced' should be 'predicted'.
- [§VIII-A] The subsection title 'Hardward' should be 'Hardware'.
- [Fig. 6] The JSON snippets in Figure 6 contain curly quotes (e.g., “4” and “yes”) that should be straight quotes for consistency and readability.
- [Table III] The table formatting is unclear: the first row in each block appears to be the baseline, but the checkmark columns are not labeled for the baseline row; using an explicit 'Baseline' label and consistent row separators would improve readability.
- [§VII-C] The claim that the recognition scheme 'achieves a high accuracy of over 90%' is supported only by the 100-object subset in Section VII-D; please report the per-property accuracies and the sample size in the main text where the claim is made.
Circularity Check
No significant circularity found; the reported gains are empirical comparisons against baselines under shared evaluation metrics.
full rationale
OPA-Pack's central claims are empirical: on 200 random virtual sequences and 50 real-world cases, OPA-Net is compared against five baselines under the same object-property annotations (Tables II, V, VI). The reward terms R_fragility and R_avoidance (Eqs. 11-12) are indeed aligned with the evaluation metrics 'Press. on Frag.' and 'Avoid. Acc.', but this is a reward-design choice, not a definitional reduction: the metrics are computed from the final packing state for all methods, including baselines that never saw the reward, and the reported gains are relative to those baselines. The property labels are generated by ChatGPT-4 and only human-checked on 100 of 1,032 objects, which is a data-quality and validity concern, not a circularity of the derivation chain. No load-bearing step relies on a self-citation; the cited baseline and architecture works are external (IR-BPP, Dueling DQN, PointNet). No uniqueness theorem or ansatz is imported from the authors' prior work. The real-robot experiments provide an independent, physically grounded check. Therefore no circular step is exhibited.
Assumptions & free parameters
free parameters (4)
- lambda (fragility reward weight) =
20
- beta (avoidance reward weight) =
0.2
- Material property table (density, density level, fragility) =
15 hand-defined entries
- Estimated density per density level =
average density of materials in the level
assumptions (5)
- domain assumption ChatGPT-4 with RAG and CoT can infer object properties from images, class name, size, retrieved examples, and the material table.
- ad hoc to paper The four avoidance relation types are sufficient for safe packing.
- domain assumption Damage to a fragile object scales with the total weight of objects placed above it.
- domain assumption Object properties in the OPA dataset are correct labels usable as ground truth for training and evaluation.
- domain assumption The 1,032 Objaverse objects are representative of real everyday packing items.
Cite this review
Pith. "Pith review of OPA-Pack: Object-Property-Aware Robotic Bin Packing." pith.science (2026). https://pith.science/paper/ZDBPJ6GH
@misc{pith2026250513339,
author = {Pith},
title = {Pith review of: OPA-Pack: Object-Property-Aware Robotic Bin Packing},
year = {2026},
howpublished = {\url{https://pith.science/paper/ZDBPJ6GH}},
note = {Machine review of arXiv:2505.13339}
}
read the original abstract
Robotic bin packing aids in a wide range of real-world scenarios such as e-commerce and warehouses. Yet, existing works focus mainly on considering the shape of objects to optimize packing compactness and neglect object properties such as fragility, edibility, and chemistry that humans typically consider when packing objects. This paper presents OPA-Pack (Object-Property-Aware Packing framework), the first framework that equips the robot with object property considerations in planning the object packing. Technical-wise, we develop a novel object property recognition scheme with retrieval-augmented generation and chain-of-thought reasoning, and build a dataset with object property annotations for 1,032 everyday objects. Also, we formulate OPA-Net, aiming to jointly separate incompatible object pairs and reduce pressure on fragile objects, while compacting the packing. Further, OPA-Net consists of a property embedding layer to encode the property of candidate objects to be packed, together with a fragility heightmap and an avoidance heightmap to keep track of the packed objects. Then, we design a reward function and adopt a deep Q-learning scheme to train OPA-Net. Experimental results manifest that OPA-Pack greatly improves the accuracy of separating incompatible object pairs (from 52% to 95%) and largely reduces pressure on fragile objects (by 29.4%), while maintaining good packing compactness. Besides, we demonstrate the effectiveness of OPA-Pack on a real packing platform, showcasing its practicality in real-world scenarios.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
Learning physically realizable skills for online packing of general 3D shapes,
H. Zhao, Z. Pan, Y . Yu, and K. Xu, “Learning physically realizable skills for online packing of general 3D shapes,” ACM Transactions on Graphics, vol. 42, no. 5, pp. 1–21, 2023
work page 2023
-
[2]
Packing irregular objects in 3D space via hybrid optimization,
Y . Ma, Z. Chen, W. Hu, and W. Wang, “Packing irregular objects in 3D space via hybrid optimization,” in Computer Graphics Forum , vol. 37, no. 5. Wiley Online Library, 2018, pp. 49–59
work page 2018
-
[3]
Heuristics integrated deep reinforcement learning for online 3D bin packing,
S. Yang, S. Song, S. Chu, R. Song, J. Cheng, Y . Li, and W. Zhang, “Heuristics integrated deep reinforcement learning for online 3D bin packing,” IEEE Transactions on Automation Science and Engineering , 2023
work page 2023
-
[4]
Robot online 3D bin packing strategy based on deep reinforcement learning and 3D vision,
J. Jia, H. Shang, and X. Chen, “Robot online 3D bin packing strategy based on deep reinforcement learning and 3D vision,” in IEEE Interna- tional Conference on Networking, Sensing and Control . IEEE, 2022, pp. 1–6
work page 2022
-
[5]
Online 3D bin packing reinforcement learning solution with buffer,
A. V . Puche and S. Lee, “Online 3D bin packing reinforcement learning solution with buffer,” in IEEE/RSJ International Conference on Intelli- gent Robots and Systems . IEEE, 2022, pp. 8902–8909
work page 2022
-
[6]
Planning irregular object packing via hierarchical reinforcement learning,
S. Huang, Z. Wang, J. Zhou, and J. Lu, “Planning irregular object packing via hierarchical reinforcement learning,” IEEE Robotics and Automation Letters, vol. 8, no. 1, pp. 81–88, 2022
2022
-
[7]
TAP-Net: transport-and-pack using reinforcement learning,
R. Hu, J. Xu, B. Chen, M. Gong, H. Zhang, and H. Huang, “TAP-Net: transport-and-pack using reinforcement learning,” ACM Transactions on Graphics, vol. 39, no. 6, pp. 1–15, 2020
work page 2020
-
[8]
Stable bin packing of non-convex 3D objects with a robot manipulator,
F. Wang and K. Hauser, “Stable bin packing of non-convex 3D objects with a robot manipulator,” in 2019 International Conference on Robotics and Automation. IEEE, 2019, pp. 8698–8704
work page 2019
Show all 49 references
-
[9]
Sdf-pack: Towards compact bin packing with signed-distance- field minimization,
J.-H. Pan, K.-H. Hui, X. Gao, S. Zhu, Y .-H. Liu, P.-A. Heng, and C.- W. Fu, “Sdf-pack: Towards compact bin packing with signed-distance- field minimization,” in 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2023, pp. 10 612–10 619
2023
-
[10]
PPN-Pack: Placement proposal network for efficient robotic bin packing,
J.-H. Pan, X. Gao, K.-H. Hui, S. Zhu, Y .-H. Liu, P.-A. Heng, and C.-W. Fu, “PPN-Pack: Placement proposal network for efficient robotic bin packing,” IEEE Robotics and Automation Letters , 2024. 15
2024
-
[11]
Dense robotic packing of irregular and novel 3d objects,
F. Wang and K. Hauser, “Dense robotic packing of irregular and novel 3d objects,” IEEE Transactions on Robotics , vol. 38, no. 2, pp. 1160– 1173, 2021
2021
-
[12]
A container loading algorithm with static mechanical equilibrium stability constraints,
A. G. Ramos, J. F. Oliveira, J. F. Gonçalves, and M. P. Lopes, “A container loading algorithm with static mechanical equilibrium stability constraints,” Transportation Research Part B: Methodological , vol. 91, pp. 565–581, 2016
2016
-
[13]
A hybrid genetic algorithm for pack- ing in 3D with deepest bottom left with fill method,
K. Karabulut and M. M. ˙Inceo˘glu, “A hybrid genetic algorithm for pack- ing in 3D with deepest bottom left with fill method,” in International Conference on Advances in Information Systems . Springer, 2004, pp. 441–450
2004
-
[14]
A hybrid genetic algorithm with a new packing strategy for the three-dimensional bin packing problem,
K. Kang, I. Moon, and H. Wang, “A hybrid genetic algorithm with a new packing strategy for the three-dimensional bin packing problem,” Applied Mathematics and Computation , vol. 219, no. 3, pp. 1287–1299, 2012
2012
-
[15]
Extreme point-based heuristics for three-dimensional bin packing,
T. G. Crainic, G. Perboli, and R. Tadei, “Extreme point-based heuristics for three-dimensional bin packing,” Informs Journal on computing , vol. 20, no. 3, pp. 368–384, 2008
2008
-
[16]
HAPE3D—a new constructive algorithm for the 3D irregular packing problem,
X. Liu, J.-M. Liu, A.-X. Cao, and Z.-L. Yao, “HAPE3D—a new constructive algorithm for the 3D irregular packing problem,” Frontiers of Information Technology & Electronic Engineering, vol. 16, no. 5, pp. 380–390, 2015
2015
-
[17]
V oxel- based solution approaches to the three-dimensional irregular packing problem,
C. Lamas-Fernandez, J. A. Bennell, and A. Martinez-Sykora, “V oxel- based solution approaches to the three-dimensional irregular packing problem,” Operations Research, 2022
2022
-
[18]
Learning to place new objects in a scene,
Y . Jiang, M. Lim, C. Zheng, and A. Saxena, “Learning to place new objects in a scene,” The International Journal of Robotics Research , vol. 31, no. 9, pp. 1021–1043, 2012
2012
-
[19]
Online 3D bin packing with constrained deep reinforcement learning,
H. Zhao, Q. She, C. Zhu, Y . Yang, and K. Xu, “Online 3D bin packing with constrained deep reinforcement learning,” in Proceedings of the AAAI Conference on Artificial Intelligence , vol. 35, no. 1, 2021, pp. 741–749
2021
-
[20]
A generalized reinforcement learning algorithm for online 3d bin-packing,
R. Verma, A. Singhal, H. Khadilkar, A. Basumatary, S. Nayak, H. V . Singh, S. Kumar, and R. Sinha, “A generalized reinforcement learning algorithm for online 3d bin-packing,” arXiv preprint arXiv:2007.00463, 2020
2007 arXiv
-
[21]
Attend2Pack: Bin packing through deep reinforcement learning with attention,
J. Zhang, B. Zi, and X. Ge, “Attend2Pack: Bin packing through deep reinforcement learning with attention,” in 2021 International Conference on Machine Learning Workshops , 2021
2021
-
[22]
Adjustable robust reinforcement learning for online 3d bin packing,
Y . Pan, Y . Chen, and F. Lin, “Adjustable robust reinforcement learning for online 3d bin packing,” Advances in Neural Information Processing Systems, vol. 36, pp. 51 926–51 954, 2023
2023
-
[23]
PackerBot: Variable-sized product packing with heuristic deep rein- forcement learning,
Z. Yang, S. Yang, S. Song, W. Zhang, R. Song, J. Cheng, and Y . Li, “PackerBot: Variable-sized product packing with heuristic deep rein- forcement learning,” in IEEE/RSJ International Conference on Intelli- gent Robots and Systems . IEEE, 2021, pp. 5002–5008
2021
-
[24]
Container packing problem with balance constraints,
I. Moon and T. V . L. Nguyen, “Container packing problem with balance constraints,” OR spectrum, vol. 36, no. 4, pp. 837–878, 2014
2014
-
[25]
Two-dimensional strip packing problem with load balancing, load bearing and multi-drop constraints,
T. A. de Queiroz and F. K. Miyazawa, “Two-dimensional strip packing problem with load balancing, load bearing and multi-drop constraints,” International Journal of Production Economics , vol. 145, no. 2, pp. 511–530, 2013
2013
-
[26]
Foundation models in robotics: Applications, challenges, and the future,
R. Firoozi, J. Tucker, S. Tian, A. Majumdar, J. Sun, W. Liu, Y . Zhu, S. Song, A. Kapoor, K. Hausman, et al., “Foundation models in robotics: Applications, challenges, and the future,” The International Journal of Robotics Research, p. 02783649241281508, 2023
2023
-
[27]
Manipllm: Embodied multimodal large language model for object-centric robotic manipulation,
X. Li, M. Zhang, Y . Geng, H. Geng, Y . Long, Y . Shen, R. Zhang, J. Liu, and H. Dong, “Manipllm: Embodied multimodal large language model for object-centric robotic manipulation,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2024, pp....
2024
-
[28]
Generate subgoal images before act: Unlocking the chain-of-thought reasoning in diffusion model for robot manipulation with multimodal prompts,
F. Ni, J. Hao, S. Wu, L. Kou, J. Liu, Y . Zheng, B. Wang, and Y . Zhuang, “Generate subgoal images before act: Unlocking the chain-of-thought reasoning in diffusion model for robot manipulation with multimodal prompts,” in Proceedings of the IEEE/CVF Conference on Computer Vis...
2024
-
[29]
V oxposer: Composable 3d value maps for robotic manipulation with language models,
W. Huang, C. Wang, R. Zhang, Y . Li, J. Wu, and L. Fei-Fei, “V oxposer: Composable 3d value maps for robotic manipulation with language models,” arXiv preprint arXiv:2307.05973 , 2023
2023 arXiv
-
[30]
Octopus: Embodied vision-language pro- grammer from environmental feedback,
J. Yang, Y . Dong, S. Liu, B. Li, Z. Wang, H. Tan, C. Jiang, J. Kang, Y . Zhang, K. Zhou, et al. , “Octopus: Embodied vision-language pro- grammer from environmental feedback,” in European Conference on Computer Vision. Springer, 2025, pp. 20–38
2025
-
[31]
Rekep: Spatio- temporal reasoning of relational keypoint constraints for robotic manip- ulation,
W. Huang, C. Wang, Y . Li, R. Zhang, and L. Fei-Fei, “Rekep: Spatio- temporal reasoning of relational keypoint constraints for robotic manip- ulation,” arXiv preprint arXiv:2409.01652 , 2024
2024 arXiv
-
[32]
Keypoint action tokens enable in-context imitation learning in robotics,
N. Di Palo and E. Johns, “Keypoint action tokens enable in-context imitation learning in robotics,” arXiv preprint arXiv:2403.19578 , 2024
2024 arXiv
-
[33]
On the opportunities and risks of foundation models,
R. Bommasani, D. A. Hudson, E. Adeli, R. Altman, S. Arora, S. von Arx, M. S. Bernstein, J. Bohg, A. Bosselut, E. Brunskill, et al. , “On the opportunities and risks of foundation models,” arXiv preprint arXiv:2108.07258, 2021
2021 arXiv
-
[34]
Gpt-4 technical report,
J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, et al. , “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774 , 2023
2023 arXiv
-
[35]
Manifoundation model for general-purpose robotic manipulation of contact synthesis with arbitrary objects and robots,
Z. Xu, C. Gao, Z. Liu, G. Yang, C. Tie, H. Zheng, H. Zhou, W. Peng, D. Wang, T. Chen, et al. , “Manifoundation model for general-purpose robotic manipulation of contact synthesis with arbitrary objects and robots,” arXiv preprint arXiv:2405.06964 , 2024
2024 arXiv
-
[36]
Retrieval-augmented generation for ai-generated content: A survey,
P. Zhao, H. Zhang, Q. Yu, Z. Wang, Y . Geng, F. Fu, L. Yang, W. Zhang, and B. Cui, “Retrieval-augmented generation for ai-generated content: A survey,” arXiv preprint arXiv:2402.19473 , 2024
2024 arXiv
-
[37]
Retrieval- augmented generation for knowledge-intensive nlp tasks,
P. Lewis, E. Perez, A. Piktus, F. Petroni, V . Karpukhin, N. Goyal, H. Küttler, M. Lewis, W.-t. Yih, T. Rocktäschel, et al. , “Retrieval- augmented generation for knowledge-intensive nlp tasks,” Advances in Neural Information Processing Systems , vol. 33, pp. 9459–9474, 2020
2020
-
[38]
Chain-of-thought prompting elicits reasoning in large language models,
J. Wei, X. Wang, D. Schuurmans, M. Bosma, F. Xia, E. Chi, Q. V . Le, D. Zhou, et al., “Chain-of-thought prompting elicits reasoning in large language models,” Advances in neural information processing systems , vol. 35, pp. 24 824–24 837, 2022
2022
-
[39]
Towards revealing the mystery behind chain of thought: a theoretical perspective,
G. Feng, B. Zhang, Y . Gu, H. Ye, D. He, and L. Wang, “Towards revealing the mystery behind chain of thought: a theoretical perspective,” Advances in Neural Information Processing Systems , vol. 36, 2024
2024
-
[40]
Automatic chain of thought prompting in large language models,
Z. Zhang, A. Zhang, M. Li, and A. Smola, “Automatic chain of thought prompting in large language models,” in The Eleventh International Conference on Learning Representations
-
[41]
A survey of chain of thought reasoning: Advances, frontiers and future,
Z. Chu, J. Chen, Q. Chen, W. Yu, T. He, H. Wang, W. Peng, M. Liu, B. Qin, and T. Liu, “A survey of chain of thought reasoning: Advances, frontiers and future,” arXiv preprint arXiv:2309.15402 , 2023
2023 arXiv
-
[42]
Learning transferable visual models from natural language supervision,
A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark, et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning . PMLR, 2021, pp. 8748–8763
2021
-
[43]
Chatgpt-4,
OpenAI, “Chatgpt-4,” https://chat.openai.com/chat, 2023
2023
-
[44]
Benchmarking in manipulation research: Using the Yale-CMU- Berkeley object and model set,
B. Calli, A. Walsman, A. Singh, S. Srinivasa, P. Abbeel, and A. M. Dollar, “Benchmarking in manipulation research: Using the Yale-CMU- Berkeley object and model set,” IEEE Robotics & Automation Magazine, vol. 22, no. 3, pp. 36–52, 2015
2015
-
[45]
Objaverse: A universe of annotated 3d objects,
M. Deitke, D. Schwenk, J. Salvador, L. Weihs, O. Michel, E. VanderBilt, L. Schmidt, K. Ehsani, A. Kembhavi, and A. Farhadi, “Objaverse: A universe of annotated 3d objects,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , 2023, pp. 13 142–13 153
2023
-
[46]
Learning efficient online 3D bin packing on packing configuration trees,
H. Zhao, Y . Yu, and K. Xu, “Learning efficient online 3D bin packing on packing configuration trees,” in International Conference on Learning Representations, 2021
2021
-
[47]
Q-learning,
C. J. Watkins and P. Dayan, “Q-learning,” Machine learning, vol. 8, pp. 279–292, 1992
1992
-
[48]
Dueling network architectures for deep reinforcement learning,
Z. Wang, T. Schaul, M. Hessel, H. Hasselt, M. Lanctot, and N. Freitas, “Dueling network architectures for deep reinforcement learning,” in International conference on machine learning. PMLR, 2016, pp. 1995– 2003
2016
-
[49]
PointNet: Deep learning on point sets for 3d classification and segmentation,
C. R. Qi, H. Su, K. Mo, and L. J. Guibas, “PointNet: Deep learning on point sets for 3d classification and segmentation,” in Proceedings of the IEEE conference on computer vision and pattern recognition , 2017, pp. 652–660
2017
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.