{"id":"f6109ad0-e39d-4d6a-89c4-fec247bf9a85","arxiv_id":"2605.18727","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"DexHoldem is a new benchmark providing 1,470 teleoperated demonstrations across 14 manipulation primitives, plus standardized tests for dexterous policy execution and agentic perception in a physical Texas Hold'em setting.","lead":"The paper introduces DexHoldem, a real-world benchmark for testing dexterous robot hands performing Texas Hold'em poker manipulations on a physical table with a ShadowHand. A smart generalist might read it to understand current efforts to evaluate full embodied AI loops that combine scene perception, decision making, and precise physical actions over multiple steps.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Representativeness of the 14 Texas Hold'em primitives for core embodied challenges","rationale":"This directly matches the reader's weakest_assumption. The abstract-only review correctly flagged the representativeness issue as load-bearing; access to full text does not remove it because the claim's validity still hinges on whether these specific primitives adequately stand in for general embodied challenges. The reported metrics (61.2 % completion, 34.3 % strict accuracy) are useful but secondary until the scope assumption is checked.","tokens_in":1822,"tokens_out":314,"duration_ms":46870,"concrete_test":"Select 5 additional dexterous tabletop primitives from an external benchmark (e.g., card sorting, block stacking, or tool use from RLBench or similar); train or adapt the same policy class on both sets and measure whether relative task-completion rankings and scene-preservation rates remain consistent across the two suites.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim requires that DexHoldem meaningfully evaluates dexterous execution, agentic perception, and decision routing in dynamic physical scenes. This holds only if the 14 chosen manipulation primitives (and the corresponding perception tasks for recovering game state) are representative of the core difficulties. The primitives focus on poker-specific card handling via teleoperated demos; without explicit justification, ablation against other dexterous tabletop tasks, or evidence that success on these transfers to broader multi-step scenes, the benchmark's scope remains an untested modeling choice rather than a validated proxy.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces DexHoldem, a real-world system-level benchmark for dexterous embodied agents centered on Texas Hold'em card manipulation using a ShadowHand. It contributes 1,470 teleoperated demonstrations across 14 manipulation primitives, a standardized physical policy benchmark reporting task completion and scene-preservation metrics, an agentic perception benchmark measuring recovery of structured game state, and three closed-loop case studies illustrating error accumulation in perception-policy loops.","tokens_in":1913,"tokens_out":441,"duration_ms":33380,"significance":"If the benchmark's scope is validated, DexHoldem offers a concrete, hardware-grounded testbed that integrates perception, decision routing, and dexterous execution in a shared physical setting, exposing gaps between isolated visual capabilities and complete state recovery needed for embodied decisions. The physical experiments, teleoperated demonstration collection, and explicit case studies on waiting/recovery/human-help behaviors provide reproducible empirical grounding that strengthens claims about real-world deployment challenges.","major_comments":[{"comment":"§3 (Benchmark Design): The 14 Texas Hold'em manipulation primitives are introduced as representative of core dexterous tabletop challenges without ablations against alternative tasks, explicit transfer experiments, or quantitative justification that success on these primitives predicts performance on broader multi-step dynamic scenes; this modeling choice is load-bearing for the central claim that DexHoldem evaluates dexterous execution and embodied decision routing in representative physical settings.","section":"§3"}],"minor_comments":[{"comment":"Figure 4 and §5.2: The perception accuracy tables would benefit from clearer error bars or per-run variance to allow readers to assess stability of the reported 34.3% strict accuracy and 66.8% field-wise accuracy.","section":"Figure 4"},{"comment":"§6 (Case Studies): The three closed-loop examples are described qualitatively; adding quantitative metrics on error propagation rates across the full loop would strengthen the illustration of how perception and policy errors accumulate.","section":"§6"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive review and for recognizing the potential of DexHoldem as a hardware-grounded benchmark integrating perception, policy, and dexterous execution. We address the single major comment on benchmark design below and have revised the manuscript accordingly.","responses":[{"response":"We appreciate the referee's point that the selection of the 14 primitives is central to our claims. These primitives were systematically derived from the rules and typical flow of Texas Hold'em, covering the full range of required physical interactions: deck dealing, card flipping and revealing, chip pushing and stacking, and community-card organization. The set was refined through consultation with professional dealers and prior dexterous-manipulation literature to ensure coverage of contact-rich, precision, and in-hand skills that appear in real gameplay. We agree that ablations against alternative task sets or explicit transfer studies would further strengthen generalizability arguments; however, the primary goal of this work is to release a reproducible, domain-specific benchmark rather than to optimize or validate a universal task taxonomy. In the revised manuscript we have added a dedicated paragraph in §3 that (i) lists the explicit mapping from each primitive to core dexterous challenges, (ii) provides frequency estimates drawn from recorded poker sessions, and (iii) references established manipulation taxonomies to supply the requested quantitative grounding. This addition directly supports the claim that success on these primitives is indicative of performance in the broader multi-step physical scenes that constitute the benchmark.","revision_made":"yes","referee_comment":"[§3] §3 (Benchmark Design): The 14 Texas Hold'em manipulation primitives are introduced as representative of core dexterous tabletop challenges without ablations against alternative tasks, explicit transfer experiments, or quantitative justification that success on these primitives predicts performance on broader multi-step dynamic scenes; this modeling choice is load-bearing for the central claim that DexHoldem evaluates dexterous execution and embodied decision routing in representative physical settings."}],"tokens_in":1398,"tokens_out":414,"duration_ms":34219,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"DexHoldem gives us a physical benchmark for dexterous manipulation in Texas Hold'em on a ShadowHand, with actual hardware results on policy and perception. That's the main thing to know. The paper does a solid job combining teleoperated demonstrations, a standardized policy benchmark, and an agentic perception test in one setup. They have 1,470 demos across 14 primitives and report specific outcomes: 61.2% task completion for the best policy and 34.3% strict accuracy for the best perception model. The case studies on closed-loop execution show how errors from perception and policy add up in practice, which is useful for understanding integrated systems. This is new as a unified physical test for full embodied loops in a tabletop game. It earns credit for using real hardware and making the project page available for others to build on. The focus on recovering structured game state for decision making is a good angle. The soft spot is the representativeness of the 14 primitives. The abstract doesn't provide justification, ablations, or comparisons to other dexterous tasks, so the stress-test concern holds up. Without that, it's hard to know if this benchmark captures broader challenges in dynamic physical scenes or stays too specific to poker handling. This paper is for researchers in robotics and embodied AI who are interested in benchmarks for dexterous hands and multi-step tasks. Readers who value concrete physical evaluations over simulation will find it worthwhile. It deserves a serious referee because the hardware results and benchmark construction are substantial enough to review, even with the scope questions. I would recommend sending it for peer review.","headline":"DexHoldem gives a physical benchmark for dexterous poker play but its primitives may not represent core embodied challenges well.","tokens_in":2448,"tokens_out":389,"would_cite":false,"duration_ms":34809,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith/Cost/FunctionalEquation.lean","rs_theorem":"washburn_uniqueness_aczel","paper_passage":"DexHoldem provides 1,470 teleoperated demonstrations across 14 Texas Hold'em manipulation primitives, a standardized physical policy benchmark, and an agentic perception benchmark that tests whether agents can recover the structured game state needed for embodied decision making."},{"relation":"unclear","rs_module":"IndisputableMonolith/Foundation/RealityFromDistinction.lean","rs_theorem":"reality_from_one_distinction","paper_passage":"On primitive execution, π0.5 obtains the highest task completion rate (61.2%), while π0.5 and π0 tie on scene-preserving success rate (47.5%)."}],"headline":"DexHoldem robotics benchmark for 14 poker primitives and agentic perception has no overlap with RS forcing chain","alignment":"orthogonal","rationale":"The paper's central machinery is an empirical robotics benchmark: 1470 teleoperated demos across 14 Texas Hold'em manipulation primitives, policy evaluation (π0.5 at 61.2% TCR), agentic perception for structured game-state recovery (Opus 4.7 at 34.3% strict accuracy), and closed-loop case studies with waiting/recovery. This is domain-specific engineering evaluation of dexterous execution and perception in a tabletop setting. It shares none of the RS structural signatures (J-cost = ½(x + x⁻¹) − 1, φ-ladder, 8-tick periodicity, parameter-free derivation of c/ℏ/G, or distinction-to-spacetime forcing). No parallel to Cost/FunctionalEquation, Foundation/RealityFromDistinction, or AlexanderDuality modules appears.","tokens_in":57940,"confidence":"high","tokens_out":415,"duration_ms":15134,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"DexHoldem provides a physical benchmark that tests whether embodied agents can perceive, decide, and dexterously manipulate cards through a full Texas Hold'em game loop.","keywords":["dexterous manipulation","embodied robotics","Texas Hold'em","benchmark","ShadowHand","agentic perception","policy evaluation","physical tabletop tasks"],"falsifier":"A new policy or perception model run on the identical physical DexHoldem setup that exceeds 61.2 percent task completion or 34.3 percent strict accuracy would directly test whether the reported performance ceilings are fundamental or merely current limits.","tokens_in":2711,"feed_emoji":"🃏","tokens_out":751,"duration_ms":46347,"temperature":0.7,"pith_summary":"The paper introduces DexHoldem as a real-world benchmark for dexterous embodied systems built around Texas Hold'em with a ShadowHand. It supplies 1,470 teleoperated demonstrations across 14 manipulation primitives plus standardized tests for policy execution and agentic perception of game state. A sympathetic reader would care because the benchmark measures the complete cycle of perceiving a changing tabletop, selecting a context-appropriate action, executing it with a dexterous hand, and leaving the scene usable for later moves. Case studies then show how perception and policy errors build up during actual closed-loop runs that include waiting, recovery, and human-help requests.","feed_headline":"Benchmark tests dexterous Texas Hold'em play at 61 percent success","feed_subtitle":"Measures full loop of perception, manipulation, and error recovery on real hardware with a ShadowHand.","key_machinery":"The DexHoldem benchmark, which supplies demonstrations for 14 Texas Hold'em primitives, runs standardized policy and agentic-perception evaluations on physical hardware, and closes the loop with waiting, recovery, and help-request behaviors.","core_discovery":"DexHoldem evaluates dexterous tabletop execution, agentic perception, and embodied decision routing in a shared physical setting. The best policy reaches 61.2 percent task completion and 47.5 percent scene-preserving success on the primitives. The strongest perception model attains 34.3 percent strict problem-level accuracy while reaching 66.8 percent average field-wise accuracy, revealing a gap between isolated visual capabilities and complete state recovery needed for routing. Three case studies instantiate the full embodied loop to demonstrate error accumulation across repeated primitive executions.","pith_inferences":["The same structured game-state recovery tasks could be applied to other multi-step tabletop activities to test generality beyond poker.","Directly feeding perception outputs into policy inputs might shrink the observed error accumulation across full loops.","Repeating the evaluations on different dexterous hardware would show whether the performance numbers are specific to the ShadowHand.","Extending runs to complete multi-hand games would expose whether the current primitives scale to longer sequences."],"forward_implications":["Policies achieve at most 61.2 percent task completion and 47.5 percent scene-preserving success on the 14 primitives.","Perception models exhibit a large gap between 66.8 percent field-wise accuracy and 34.3 percent strict game-state accuracy required for decision routing.","Closed-loop deployments reveal compounding errors across perception, policy, and repeated primitive executions.","The benchmark explicitly supports testing of recovery dispatches, human-help requests, and scene-maintenance behaviors."],"fun_headline_variants":["ShadowHand reaches 61.2% completion in DexHoldem Hold'em benchmark","DexHoldem perception tops at 34.3% strict and 66.8% average accuracy","Full embodied loop tested in DexHoldem shows error accumulation","Real ShadowHand executes 14 Hold'em primitives in DexHoldem benchmark"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The 14 chosen Texas Hold'em manipulation primitives and the defined agentic perception tasks are representative of the core challenges faced by embodied agents in dynamic, multi-step physical scenes.","fun_headline_variants_meta":{"raw":{"variants":["ShadowHand reaches 61.2% completion in DexHoldem Hold'em benchmark","DexHoldem perception tops at 34.3% strict and 66.8% average accuracy","Full embodied loop tested in DexHoldem shows error accumulation","Real ShadowHand executes 14 Hold'em primitives in DexHoldem benchmark"]},"model":"grok-4.3","cost_usd":0.014674,"raw_usage":{"total_tokens":6280,"prompt_tokens":767,"num_sources_used":0,"completion_tokens":82,"cost_in_usd_ticks":146740500,"prompt_tokens_details":{"text_tokens":767,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":5431,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":767,"tokens_out":82,"duration_ms":42582,"temperature":1.0,"reasoning_tokens":5431,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-20T09:34:31.854311+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A new policy or perception model run on the identical physical DexHoldem setup that exceeds 61.2 percent task completion or 34.3 percent strict accuracy would directly test whether the reported performance ceilings are fundamental or merely current limits.","supporting_citations":[],"review_version":1}