{"id":"a3e3ca65-16b9-4e3b-a457-ca4afb05c2c5","arxiv_id":"2607.23389","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Controlled WebGL render-timing cleanly separates software bots (~5× slower) and still distinguishes headless real-GPU automation from interactive use on Intel GPUs via jitter and timer quantization.","lead":"GPU render timing under a fixed WebGL workload separates software bots from real GPUs by about 5×, and even headless automation on real silicon leaves a measurable jitter signature. It offers a passive, non-identifying bot filter that AI puzzle-solvers do not address.","discovery_kind":"new_method","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The hard-negative separation may be a device/launch-configuration effect rather than a headless-execution effect: the “differing only” claim is not yet supported because Intel human and bot sessions are unpaired and hard negatives span only 2 GPUs.","rationale":"I agree with the reader that cross-architecture replication is open, but I think that is downstream. The more immediate load-bearing issue is whether the Intel hard-negative signature is even internally established as a headless-mode effect rather than a two-device/configuration effect. The paper deserves credit for honest pilot scoping, single-endpoint collection, trustworthy negative labels via keyed automation, self-gating hard negatives to genuine hardware, and stating device-held-out evaluation as the correct discipline. Those choices reduce, but do not eliminate, the pairing/confound problem because the actual Table 2 comparison is not paired and has too few hard-negative devices to apply that discipline. I would not reject: the render-free wild adversary result and the large software-renderer gap are useful and plausibly robust. I would not accept unconditionally: the “differing only in headless vs interactive execution” phrase is stronger than the design shown. Keep the verdict CONDITIONAL, but make the gating condition more specific than “more hardware”: paired same-device, flag/backend-controlled, order-randomized replication with per-device and held-out-device statistics must preserve the jitter/quantization/CV separation before the hard-negative claim is relied on.","tokens_in":6107,"tokens_out":3138,"duration_ms":57639,"concrete_test":"Run a paired randomized A/B on the same physical machines: for at least 8–10 Intel iGPU systems, collect ≥30 interleaved sessions per arm with identical Chromium binary/profile, recorded flags, fixed ANGLE backend, fixed workload, randomized headed/headless order, logged timer resolution and thermal/power state. Analyze feature ~ mode + flags/backend/timer-resolution + (1|device), report per-device paired deltas and bootstrap CIs, then hold out whole devices. If the headless coefficient remains positive and excludes 0 across held-out devices after covariates, the concern does not land; if within-device deltas approach zero, change sign, or vanish after flags/backend are controlled, the strongest claim fails.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"For the strongest claim to hold, toggling the same silicon/browser from interactive to headless must causally raise timer quantization, jitter, and CV via loss of compositor/refresh pacing. Table 2 instead compares n=14 interactive human Intel sessions with n=20 headless hard-negative sessions across only 2 hard-negative GPUs; it is not a within-device paired toggle. “Same GPU family” is not the same GPU, driver, OS power state, thermal condition, ANGLE backend, or Chromium flag set. This matters because common automation launches use flags such as --disable-gpu-vsync, --disable-frame-rate-limit, --use-angle=*, feature disables, or virtual-time/begin-frame control that directly remove pacing or change timer behavior; the measured roughness could detect that configuration rather than automation-on-real-hardware generally. With two hard-negative devices, the paper’s own device-held-out discipline cannot be meaningfully applied, and no CIs, mixed-effects model, or per-device deltas are reported. The direction is physically plausible and the mean-render overlap is honestly noted, but the 75–106% univariate gaps could still be dominated by the identity/configuration of the two headless Intel machines. There is also a presentational inconsistency: the text says 3.3×/2.6×/2.2× while Table 2 labels 106/89/75 as “Sep.”, which looks like log-ratio or rounded-effect reporting rather than percent-higher; this is secondary but should be clarified before thresholds are trusted.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The paper proposes GPU render-timing under a controlled WebGL workload as a passive bot-detection signal that classifies (human vs. automated) rather than identifies, distinguishing it from WebGL pixel-hash fingerprinting. It reports three empirical components: (i) a 12-hour passive deployment characterizing unsolicited traffic (207 requests, 86% judged automated, none reaching a rendering path); (ii) a single-endpoint collection methodology in which consenting real browsers (26 sessions, 13 GPUs) and keyed headless automation (software and hardware backends) flow through the same harness, giving trustworthy negative labels; and (iii) a pilot analysis showing software-rendered automation ~5.2× slower in mean render time (26.3 ms vs 5.1 ms), and — in an Intel/Chromium-matched comparison — headless automation on real GPUs showing higher timer-quantization ratio, frame jitter, and CV than interactive human sessions (Table 2). The authors explicitly scope the hard-negative result as pilot-scale, single-architecture, and make no cross-architecture or trained-classifier claim.","tokens_in":6457,"tokens_out":5507,"duration_ms":120569,"significance":"If the signal generalizes, a sub-second, enrollment-free, identifier-free physical measurement that is orthogonal to the deployed fingerprint-spoofing ecosystem (renderer strings, pixel-hash noise, antidetect browsers) would be a genuinely useful addition to multi-signal bot detection, and the paper's positioning against visual CAPTCHAs, behavioral biometrics, PoW, and WebGL fingerprinting is clear and accurate. Credit is due for several methodological choices that are uncommon at this scale: a single shared harness for both classes (eliminating measurement confounds between them), server-side keying for trustworthy negative labels, an explicit device-held-out evaluation discipline (§4.5), an attempt at engine/GPU-family confound control, honest reporting that hard negatives overlap humans on mean render time, a physically motivated (and counterintuitive) direction for the headless signature — automation is *rougher*, not flatter — and an explicit, falsifiable enumeration of the cross-architecture collection needed next. The work is honestly framed as a pilot; the central risks are that the pilot's two headline comparisons are weaker than the language describing them, and that the","major_comments":[{"comment":"The phrase \"identical GPU family and browser engine, differing only in headless vs. interactive execution\" overstates the design. Table 2 compares n=14 interactive human Intel sessions against n=20 headless sessions drawn from only 2 GPUs; this is an unpaired, cross-device comparison, not a within-device toggle of the same silicon between headed and headless modes. \"Same GPU family\" does not control for the specific GPU, driver version, OS power/thermal state, ANGLE backend, or — critically — the Chromium launch configuration. Common automation flags (--disable-gpu-vsync, --disable-frame-rate-limit, --use-angle=*, virtual-time/begin-frame control) directly remove compositor/refresh pacing or alter timer behavior, so the measured 2.2–3.3× elevation in quantization ratio, jitter, and CV may detect the launch configuration of the authors' sweep tool rather than headless-execution-on-real-ha","section":"§5.2 / Table 2"},{"comment":"No confidence intervals, significance tests, or per-GPU aggregates are reported anywhere. The 20 hard-negative sessions span only 2 GPUs, so the effective sample size for the headless signature is closer to 2 devices than 20 sessions; sessions within a GPU are strongly correlated, and treating them as independent inflates the apparent separation. The same concern applies to the human class (26 sessions / 13 GPUs) and the software-bot class (273 sessions / 4 GPUs — the 5.2× mean-time gap, though large, rests on 4 devices). Relatedly, §4.5 states a device-held-out evaluation discipline, but no classifier or threshold is ever evaluated under it; with 2 hard-negative GPUs a leave-one-GPU-out split is degenerate. At minimum: report per-GPU summary statistics, cluster-bootstrapped confidence intervals (clustering on GPU) for the Table 1 and Table 2 effects, and either demonstrate the §4.5 disc","section":"§4.5, §5.1–5.2"},{"comment":"The threat-model generalization is not supported by the sample. A single cloud endpoint with no deployed challenge, passively observed for 12 hours, received 207 requests dominated by secret- and configuration-discovery probes; exactly one touched a data-collection path and none reached a rendering endpoint. This measures generic background scanning noise, not the population of adversaries who would attempt to bypass a CAPTCHA deployed at a valuable target. The inferences that \"the operative adversary is overwhelmingly render-free\" (§3) and that software-rendered automation is \"empirically the dominant real-world adversary\" (§5.1, also Abstract) should be scoped accordingly — e.g., \"the unsolicited traffic reaching an unadvertised endpoint\" — or backed by citation to larger measurements of CAPTCHA-targeted traffic. Additionally, the 86% automated figure derives from the authors' own head","section":"§3, §5.1"},{"comment":"The title and framing claim the signal is \"AI-resistant\" and (§6) \"hard-to-spoof,\" but the manuscript contains no evasion analysis beyond acknowledging that an adversary \"can render on real hardware or inject synthetic jitter.\" Since the hard-negative samples are produced by the authors' own non-adaptive sweep tool, the measured signature is a signature of that tool's default configuration; an adaptive adversary could plausibly erase it (headed-mode automation under a virtual display, relaxed launch flags, begin-frame control to impose human-like pacing, or distribution-matched noise injection into the frame series). For a security venue, at least a preliminary robustness assessment is needed — e.g., show whether simple synthetic jitter or pacing control moves the hard-negative feature values in Table 2 into the human range — or the title and abstract should be tempered to \"passive GPU-t","section":"Title / §6 (Adversary adaptation)"}],"minor_comments":[{"comment":"Effect-size reporting is inconsistent. §5.2 states \"3.3× higher quantization ratio, 2.6× higher jitter, and 2.2× higher CV,\" while Table 2's \"Sep.\" column reports 106%/89%/75% and the Abstract uses 75–106%. The Table 2 values appear to be a symmetric percent difference (|b−h|/((b+h)/2)), not the ratio multiples quoted in the text. Define the metric in the table caption and use one convention consistently across Abstract, §5.2, and Table 2; \"3.3× higher\" is also ambiguous (3.3× as high vs. 3.3× higher = 4.3×) and should be replaced by an unambiguous ratio.","section":"Table 2 vs. §5.2 / Abstract"},{"comment":"Table 1's human jitter (0.42) and Table 2's human jitter (0.351) differ without explanation; presumably Table 2 restricts to the Intel subset (n=14 of 26). State this explicitly in both captions, and give units for the jitter column (ms).","section":"Table 1 / Table 2"},{"comment":"Figure 4's caption says hard negatives form a band of \"elevated quantization at low jitter,\" but Table 1 shows hard-negative jitter (0.91) roughly double the human value (0.42). \"Low\" apparently means relative to software bots; clarify the caption so it does not appear to contradict the tables.","section":"Figure 4"},{"comment":"The timer-quantization feature is entangled with browser timer granularity (coarsening / anti-fingerprinting jitter), which the paper notes in §6 but does not analyze. Since the harness already records observed timing-source resolution per session (§4.2), report the per-class distribution of observed resolutions and show that the Table 2 quantization gap survives conditioning on equal timer granularity.","section":"§4.3 / §6 (Timer hardening)"},{"comment":"The harness is under-specified for replication: shader/workload intensity, frame-window length, warm-up duration, and the header-consistency rule set are not given, and no code or data release is mentioned. Given that workload intensity is a free parameter that likely modulates every reported gap, release the harness and the raw per-frame series (which §4.2 says are retained), or at minimum specify these parameters.","section":"§4.2"},{"comment":"Several load-bearing empirical claims rest on vendor or blog sources: the 96–99.8% solver-accuracy range cites [1] (ScopeDesign blog) alongside [2]; header-based detection practice cites [6] (KnowledgeSDK); the WebGL fingerprinting and spoofing discussion cites [11, 12] (Spidra, Browserless). Where peer-reviewed alternatives exist (e.g., for WebGL fingerprinting, the literature has canonical references; for CAPTCHA solver accuracy, [2] and [5] suffice), prefer them; the survey [3] and academic works already cited are appropriate.","section":"References"},{"comment":"The 12-hour deployment's header-consistency rule (presence of Sec-Fetch-*, Accept-Language, content-specific Accept) should be stated as an explicit rule set, since \"85% of browser-claiming clients failed\" depends entirely on it; note also that Sec-Fetch-* absence is expected from non-Chromium clients and first-party navigations, which could misclassify some genuine traffic.","section":"§3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is honest about being a pilot and is a clean, well-motivated measurement study, but the two headline numbers (5.2× and 75–106%) currently rest on 4 and 2 devices respectively, with language (\\\"confound-controlled, differing only in headless vs. interactive execution\\\") stronger than the design. The fixes I request are mostly re-scoping plus a modest paired experiment and cluster-level statistics — within the paper's scope but beyond cosmetic. If the venue expects deployed-system evidence, this may fit better as a workshop/short paper; as a research contribution it is publishable once the claims are brought in line with the sample. The citation pattern leans on vendor blogs for a few empirical claims; not disqualifying, but worth the authors' attention."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful bit is simple: they treat WebGL render timing as a classification signal, not a fingerprint, and they actually measured what hits an open endpoint. Wild traffic is mostly render-free junk (86% automated, header-inconsistent browser claims). Against that adversary, software-rendered bots sit ~5× slower in mean frame time. That alone is a cheap filter, and the single shared harness plus keyed negatives is the right collection design.\n\nWhat is actually new is the hard-negative angle—headless Chromium on real silicon vs interactive humans—and the claim that missing compositor/refresh pacing leaves higher jitter, CV, and timer-quantization. On Intel + Chromium/ANGLE they report 75–106% separations. They scope it as pilot, one architecture, no cross-GPU claim. Citation pattern is fine: fingerprinting, PoW, behavioral biometrics, and offensive GPU timing are in the right places; they are not reinventing those.\n\nSoft spots are real but proportional. The stress note lands: Table 2 is same family, not paired same-device toggles. Hard negatives are 20 sessions on 2 GPUs, so their own device-held-out rule barely applies, and launch flags (vsync, frame-rate limit, ANGLE backend) could drive the roughness as much as “headless.” No CIs, no mixed model, no released harness/data, no trained classifier metrics. Text multipliers (3.3×/2.6×/2.2×) and table “Sep.” percents need a one-line reconcile. None of that kills the software-bot result or the methodological idea.\n\nThis is for people building multi-signal anti-abuse stacks who want an orthogonal, sub-second, non-identity probe. Not a CAPTCHA theory paper. I would send it to referees; it is sharp enough and honest enough to deserve that time, with the obvious ask for paired within-device toggles, more architectures, and artifacts. Worth a skim if you work bots or browser side channels; not mandatory otherwise.","headline":"Honest pilot with a clean software-bot gap and a real but under-controlled hard-negative claim on Intel only.","tokens_in":7674,"tokens_out":513,"would_cite":false,"duration_ms":23000,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Real GPU render timing under a controlled WebGL load can separate genuine browsers from automated clients without puzzles, enrollment, or a persistent device ID.","keywords":["CAPTCHA","GPU timing","WebGL","bot detection","headless browsers","hardware attestation","render timing","passive signals"],"falsifier":"Collect matched human and hard-negative (real-GPU headless) sessions on discrete NVIDIA, AMD, and Apple Silicon under the same engine and harness; if jitter, quantization ratio, and CV no longer separate the two classes, the central claim fails to generalize.","tokens_in":7548,"feed_emoji":"🖥️","tokens_out":915,"duration_ms":30055,"temperature":0.7,"pith_summary":"This paper argues that CAPTCHA-style bot defense can rest on a physical property rather than a puzzle: how a client’s GPU actually times a short, controlled WebGL rendering workload. Unlike WebGL fingerprinting, which turns pixel output into a lasting device identifier, the method keeps only session timing statistics and aims to classify human versus automated clients. The authors first show that most unsolicited traffic never renders at all, then collect matched timing samples from real browsers and from headless automation. Software-rendered bots are about five times slower than real GPUs; even headless automation on real Intel hardware still looks rougher than interactive use on the same chip. If the pattern holds more broadly, sites could add a fast, privacy-light hardware check that AI solvers and header-spoofing bots do not naturally supply.","feed_headline":"GPU render timing spots bots without puzzles or tracking IDs","feed_subtitle":"Software bots run ~5× slower; even headless real GPUs leave a measurable jitter signature.","key_machinery":"A single-endpoint WebGL harness that runs a deterministic, fragment-shader-bound workload, forces pipeline completion with synchronous pixel read-back, and records per-frame render times; offline features (mean, jitter, CV, timer-quantization ratio) turn that series into a human-versus-bot score without retaining a persistent identifier.","core_discovery":"On pilot data, software-rendered automation separates from genuine GPUs by roughly 5× in mean render time, and on a confound-controlled Intel integrated GPU comparison with the same browser engine, headless automation still separates from interactive human sessions by 75–106% on timer-quantization ratio, frame jitter, and coefficient of variation. The paper presents this as evidence that physical render-timing dynamics are a usable classification signal, not merely another fingerprint hash.","pith_inferences":["If discrete GPUs compress the gap, defenders may need harder shader workloads or multi-frame probes tuned per GPU tier rather than one global threshold.","Combining render-timing with existing header-consistency checks would catch both non-rendering scrapers and the smaller set of headless clients that do acquire a GPU.","Browser vendors that further coarsen timers would weaken this signal and any defensive use of it, creating a quiet policy tension with anti-bot needs.","The same harness could later test whether mobile GPUs and power-saving modes still leave a stable interactive-versus-automated gap."],"forward_implications":["A render-gated challenge would filter almost all of the render-free automated traffic that dominates open endpoints.","Sites could score sessions with a sub-second physical signal that does not require enrollment, TPM/WebAuthn, or a stable device hash.","The signal is orthogonal to pixel-hash WebGL fingerprinting and to behavioral biometrics, so it can sit inside multi-signal risk scores.","Evading it forces the adversary onto real GPU hardware or synthetic-jitter injection, raising cost rather than enabling free spoofing of a puzzle answer.","Cross-architecture collection with the same single-endpoint method becomes the direct next measurement program."],"fun_headline_variants":["GPU render timing separates software bots by ~5× mean time","Render-timing dynamics flag headless GPUs without IDs or puzzles","Pilot data: headless real GPUs lag humans 75–106% on jitter metrics","WebGL render timing classifies bots, leaks no persistent identifier","Software automation ~5× slower than genuine GPU render timing"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The timing roughness that separates headless from interactive use on Intel integrated GPUs will still be measurable on faster discrete and Apple GPUs, rather than shrinking until it is useless.","fun_headline_variants_meta":{"raw":{"variants":["GPU render timing separates software bots by ~5× mean time","Render-timing dynamics flag headless GPUs without IDs or puzzles","Pilot data: headless real GPUs lag humans 75–106% on jitter metrics","WebGL render timing classifies bots, leaks no persistent identifier","Software automation ~5× slower than genuine GPU render timing"]},"model":"grok-4.5","effort":"low","cost_usd":0.002031,"raw_usage":{"total_tokens":967,"prompt_tokens":834,"num_sources_used":0,"completion_tokens":81,"cost_in_usd_ticks":20308000,"prompt_tokens_details":{"text_tokens":834,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":52,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":834,"tokens_out":81,"duration_ms":3471,"temperature":1.0,"reasoning_tokens":52,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T23:29:52.224222+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Collect matched human and hard-negative (real-GPU headless) sessions on discrete NVIDIA, AMD, and Apple Silicon under the same engine and harness; if jitter, quantization ratio, and CV no longer separate the two classes, the central claim fails to generalize.","supporting_citations":[],"review_version":1}