{"id":"01232be0-95ee-4302-92a5-d822c4eed634","arxiv_id":"2607.14443","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Adding Tactile, an MCP tool layer that grounds agent actions in macOS accessibility semantics, OCR, and visual fallback, raised Codex Success@100 from 41.1% to 50.0% on macOSWorld-style tasks.","lead":"Tactile is an open-source tool layer that gives desktop-using AI agents structured 'hands and feet': it turns macOS accessibility semantics, OCR text, and visual regions into ranked, verifiable targets instead of raw screen coordinates. Adding it raised Codex's success rate from 41% to 50% on macOSWorld-style tasks, with smaller reported gains across three other agents — though those cross-agent numbers are best-of-three upper bounds.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-agent 'With Tactile' is a best-of-three oracle upper bound; if the headline Codex comparison uses the same rule, the abstract's causal claim is unsupported.","rationale":"The reader's weakest_assumption is right: the comparison protocol between baseline and treatment is the load-bearing point. Section 5 explicitly defines the cross-agent With Tactile score as an upper bound over no-skill, tactile-implicit, and tactile-explicit per task, while the baseline is only no-skill. That makes the cross-agent table (Figure 4 right) a comparison of an oracle max against a single policy. Since the abstract states 'adding TACTILE improves' without mentioning this, the headline causal claim is overstated if the rule applies to the main Codex result. The paper is transparent in §5.3 about the upper bound and missing uncertainty intervals, but the abstract does not carry these caveats. The most load-bearing issue is therefore not the absence of error bars per se (which is a reporting gap) but the definition of the treatment condition: unless the 'With Tactile' condition is a fixed, executable policy, the measured gain does not match the claim that using TACTILE improves success. My recommended verdict is unchanged: CONDITIONAL, with the concrete test above as a condition. I agree with the reader's identification of this concern.","tokens_in":9577,"tokens_out":5383,"duration_ms":48906,"concrete_test":"Run the 96-task cross-agent set under three separate fixed policies—no-skill, tactile-implicit, tactile-explicit—and report Success@100 with bootstrap 95% CIs per agent per policy. Also state explicitly in the evaluation section whether the graded Codex comparison (41.06%→50.00%) is from a single fixed policy or from a best-of-three upper bound. If tactile-explicit alone does not beat no-skill for Codex overall, or if the headline Codex number is actually the upper bound, the causal claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is causal: adding TACTILE as an execution substrate improves Success@100. The evidence for the cross-agent claim in §5 is explicitly not a single policy: 'The reported cross-agent With Tactile score is a skill-optional upper bound computed from the best per-task result among no-skill, tactile-implicit, and tactile-explicit settings.' This is an oracle bound: for each task it selects whichever of three settings scored best, while the baseline is a single no-skill condition. Comparing best-of-three against one inflates apparent gains by construction, and a deployed agent could not realize this value without knowing the per-task optimum in advance. The Goose gain is +2.08pp on N=96 (≈2 tasks) and the overall cross-agent 'consistent gains' claim rests on this kind of noisy delta. The paper does not state whether the headline Codex comparison (41.06%→50.00% in §5.1) also uses the upper-bound rule; if it does, the abstract's strongest number is an oracle upper bound. The system design and the AX-first ladder are plausible, but the empirical attribution is not established until a fixed-policy comparison with uncertainty is reported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces TACTILE, an open-source MCP-compatible tool layer for macOS desktop agents. It converts accessibility-tree data, OCR text, and visual regions into a compact, action-grounded interface state, and wraps execution in an observe-ground-act-verify loop that prefers native semantic actions over coordinate clicks. The authors evaluate on macOSWorld-style tasks, reporting Success@100 gains for Codex (41.1% to 50.0% overall; 45.2% to 55.3% on accessibility-adapted tasks) and for a 96-task cross-agent subset across Codex, Claude Code, OpenCode, and Goose. The paper is candid that the cross-agent 'With Tactile' score is a skill-optional upper bound, and §5.3 and §7 list limitations including missing uncertainty estimates and the need for stronger baselines.","tokens_in":9764,"tokens_out":3494,"duration_ms":37347,"significance":"If the central causal claim is established, TACTILE would be a useful reusable execution substrate: it makes targetability, actionability, verifiability, and auditability first-class properties, and the accessibility-first ladder is a sensible design principle. The paper has real strengths: it is open-source, evaluated against an external benchmark, includes trace-based examples, and is unusually candid about its limitations. However, the current evidence does not yet support the strength of the abstract's claims because the cross-agent comparison uses an oracle-style upper bound and the main results lack uncertainty quantification.","major_comments":[{"comment":"The abstract and §5.1 present 'consistent gains across Codex, Claude Code, OpenCode, and Goose' as if they describe a deployable policy. But §5 states: 'The reported cross-agent With Tactile score is a skill-optional upper bound computed from the best per-task result among no-skill, tactile-implicit, and tactile-explicit settings.' Comparing this best-of-three score against a single no-skill baseline inflates the apparent gain by construction, because a real agent must commit to one policy before seeing the task. Please report a fixed-policy comparison and state explicitly whether the headline Codex numbers in §5.1 also use the upper-bound rule; if they do, the abstract's lead numbers are oracle upper bounds and must be labeled as such.","section":"§5, cross-agent comparisons"},{"comment":"No uncertainty estimates are reported anywhere. On the 96-task cross-agent subset, the Goose gain is +2.08 percentage points, which is approximately two tasks; without confidence intervals, a paired per-task analysis, or multiple seeds, the 'consistent gains' claim is not robustly supported. §5.3 defers uncertainty to future work, but the abstract states the improvements as definitive. At minimum, report Wilson intervals or bootstrap CIs for each Success@100 estimate and for the deltas, and state the number of tasks in each split.","section":"§5.1 / §5.3"},{"comment":"The partition into AX-adapted and Limited-AX tasks is introduced post hoc and is central to the mechanistic interpretation that accessibility semantics drive the gain. No independent criteria, annotation protocol, or reliability measure are given for this split. Because the headline effect is largest on the AX-adapted subset, the split should be specified before evaluating, or at least independently validated; otherwise the interpretation risks being shaped by the authors' own categorization of the same tasks used in the comparison.","section":"§5, task splits"}],"minor_comments":[{"comment":"There is an unresolved 'Figure??' reference; the evidence-compiler figure is missing or its cross-reference is broken.","section":"§4.2"},{"comment":"The abstract rounds 41.06% to 41.1% and 50.00% to 50.0%; this is fine, but the abstract omits the upper-bound caveat that the body applies to the cross-agent comparison. A one-sentence qualifier ('under a skill-optional upper-bound evaluation') would prevent over-reading.","section":"Abstract / §5.1"},{"comment":"The Zoom trace example is informative, but no success/failure was reported for that specific trace. Stating whether the recorded trajectory completed the task would strengthen the illustrative value.","section":"§5.2"}],"recommendation":"major_revision","confidential_remarks":"The core idea is interesting and the system design is plausible, but the empirical section as written supports a more modest claim than the abstract. The authors should be asked to either run a fixed-policy comparison (or at least clearly separate oracle-upper-bound numbers from policy-realizable numbers) and to add uncertainty quantification. With those changes, the paper could be acceptable; without them, the headline causal claim is not supported."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nRead Tactile. It is a solid systems paper, not a breakthrough. What's new is the assembly: an MCP-compatible execution layer that turns accessibility trees, OCR, and visual fallback into a ranked candidate list with provenance, normalized coordinates, and verification cues, plus an observe-ground-act-verify loop. None of the ingredients are new — the paper cites the accessibility APIs, OCR grounding, MCP, and even the macOS MCP server that informed its code — but the integration is coherent and the open-source implementation matters. The four runtime requirements in Table 1 are a sensible framing.\n\nThe paper is also unusually honest. Section 5.3 says the cross-agent score is an upper bound, and Section 7 lists the real limitations. That candor is creditworthy.\n\nThe soft spots are exactly where the reader put them. The cross-agent 'With Tactile' number is best-of-three per task against a single no-skill baseline — that is an oracle bound, not a deployable policy. The Goose gain is +2.08pp on N=96, roughly two tasks. No error bars anywhere. The paper does not state whether the headline Codex comparison (41.06% to 50.00%) also uses the upper-bound rule; if it does, the abstract's strongest number is unsupported as stated. The AX-adapted vs. Limited-AX partition is post hoc and drives the mechanistic story, so it needs to be treated as a hypothesis, not a conclusion.\n\nNone of this is fatal. The direction of the effect is consistent across all four agents, the system design is plausible, and the body openly lists most of the missing pieces. The fix is straightforward: report a single fixed-policy condition, add uncertainty intervals, clarify the main comparison, and ship pinned artifacts. The broken figure reference is minor but should be cleaned up.\n\nThis deserves a serious referee. It is a useful contribution to the computer-use subfield, and the open-source layer will likely be reused. I would send it to review with a request for the fixed-policy evaluation, not desk reject it.","headline":"A genuinely useful systems integration paper whose headline effect is real but whose cross-agent numbers are an oracle upper bound — worth refereeing, but only after a fixed-policy comparison is reported.","tokens_in":10362,"tokens_out":1722,"would_cite":true,"duration_ms":17705,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that desktop computer-use agents fail less from lack of model intelligence and more from a brittle motor interface, and introduces Tactile, an open tool layer that gives agents semantic, verifiable 'hands and feet' for des","keywords":["computer-use agents","desktop automation","accessibility API","GUI grounding","action-grounded interface state","observe-ground-act-verify","tool layer","Success@100"],"falsifier":"Re-run the 96-task cross-agent benchmark with a single pre-committed 'always use Tactile' policy and report confidence intervals; if the Codex gain collapses below noise or another agent's gain inverts, the claim of consistent gains from the semantic substrate would fail. Also verify whether the main Codex graded comparison uses the same upper-bound rule.","tokens_in":9343,"feed_emoji":"🖱️","tokens_out":4067,"duration_ms":38019,"temperature":0.7,"pith_summary":"This paper argues that desktop computer-use agents fail less because of model intelligence and more because their interface to software is a brittle motor layer: screenshot, predict coordinates, click, hope. It introduces Tactile, an open tool layer that turns UI evidence (accessibility semantics, OCR text, visual regions) into action-grounded interface states—candidates with roles, state, geometry, affordances, and verification cues—and runs an observe-ground-act-verify loop that prefers native semantic actions. On macOSWorld-style tasks, adding Tactile raises Success@100 from 41.1% to 50.0% overall, with larger gains on accessibility-adapted tasks, and a 96-task cross-agent subset shows gains across four agents. The paper's central claim is that a reusable execution substrate—exposing actions as semantic, verifiable, auditable objects—is a necessary complement to stronger models.","feed_headline":"Semantic tool layer lifts desktop-agent success to 50%","feed_subtitle":"Giving agents semantic targets instead of raw screen coordinates raises Success@100 from 41% to 50% on macOS-style tasks.","key_machinery":"The action-grounded interface state is the central object: a ranked set of target candidates built from accessibility elements, OCR lines, and visual regions, each with source labels, role/text, state, geometry, executable affordances, and provenance. The accessibility-first operating ladder orders evidence—Level 1 native accessibility semantics, Level 2 OCR-grounded coordinates, Level 3 visual fallback—so each step uses the richest available signal. The observe-ground-act-verify loop separates the four decisions that screenshot-first control collapses: collect evidence, select a candidate, execute the safest primitive (semantic action when available, coordinate click otherwise), and re-obse","core_discovery":"The central discovery is that separating target grounding, action execution, and outcome verification—and preferring accessibility semantics before OCR coordinates before visual fallback—improves grounded desktop operation across agents. Tactile compiles heterogeneous UI evidence into compact target candidates: each candidate carries source labels, role or text, state, geometry, executable affordances, and verification cues; agents choose candidates, and the runtime executes the safest primitive and re-observes to verify. The authors report that this raises Codex Success@100 from 41.1% to 50.0% overall, from 45.2% to 55.3% on accessibility-adapted tasks, and produces consistent gains on a 96","pith_inferences":["The cross-agent 'With Tactile' score is a best-of-several-settings upper bound compared with a single no-skill baseline; a deployed agent that must commit to one policy in advance could see smaller, or in the weakest case negligible, gains.","A direct test of the mechanism would compare Tactile against a control that normalizes coordinates and adds verification without semantic targets; if that control matches Tactile's gains, the accessibility-first ladder is not the active ingredient.","The largest untested upside is policy learning: a router that decides per-step whether to invoke semantic, OCR, or visual tools could beat both the no-skill and always-Tactile baselines.","Because the implementation is strongest on macOS with only early Windows support, the cross-platform claim is a design commitment, not yet an empirical result."],"forward_implications":["Adding Tactile raises Codex Success@100 from 41.1% to 50.0% overall and from 45.2% to 55.3% on accessibility-adapted tasks.","Gains generalize across four different agents on a 96-task subset, with the largest improvements on tasks where applications expose useful accessibility metadata.","Accessibility semantics provide the strongest action and verification contracts; OCR and visual fallback remain necessary for semantically opaque interfaces.","Separating grounding from execution makes failures attributable: traces can show whether a target was absent, filtered, mis-ranked, mis-executed, or unverifiable.","Because the same accessibility metadata serves screen-reader users and agents, improving human accessibility also improves agent operability."],"fun_headline_variants":["Semantic grounding lifts agent success to 50%","Hands and feet for agents: 50% success on desktop","UI semantics beat pixels: agent hits 50% success","New tool layer raises agent desktop to 50% success","From coordinates to semantics: agents jump to 50%"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The reported gains are attributed to the semantic operating layer, but the cross-agent comparison uses a best-of-three upper bound with no uncertainty estimates, so the improvement could shrink if an agent must commit to a single policy in advance.","fun_headline_variants_meta":{"raw":{"variants":["Semantic grounding lifts agent success to 50%","Hands and feet for agents: 50% success on desktop","UI semantics beat pixels: agent hits 50% success","New tool layer raises agent desktop to 50% success","From coordinates to semantics: agents jump to 50%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000251,"raw_usage":{"total_tokens":1417,"prompt_tokens":793,"completion_tokens":624,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":537,"completion_tokens_details":{"reasoning_tokens":542}},"tokens_in":537,"tokens_out":624,"duration_ms":6405,"temperature":1.0,"reasoning_tokens":542,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T02:04:07.004881+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the 96-task cross-agent benchmark with a single pre-committed 'always use Tactile' policy and report confidence intervals; if the Codex gain collapses below noise or another agent's gain inverts, the claim of consistent gains from the semantic substrate would fail. Also verify whether the main Codex graded comparison uses the same upper-bound rule.","supporting_citations":[],"review_version":1}