{"id":"12f85682-4b42-4d22-99ac-19a389a03d7c","arxiv_id":"2608.10194","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A new AI auditing method, the contextual audit, treats each measurement modality as provisionally true and is demonstrated on a motion-capture skeleton inference case study.","lead":"This paper introduces 'contextual audits,' a way to test AI systems by examining them inside the real-world practices that produce their data. The authors then present a motion-capture case study to show how treating every measurement method as provisionally correct can guide audits when true values are unknown.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Low-power null results underpin the 'mostly holds' findings; absence of significance is not evidence of reliability at N=24.","rationale":"The reader's weakest assumption is also the most load-bearing concern I can identify. The conceptual contribution can stand without the case study, but contribution 3 and the claim that symmetry enables audits in practice rest on the empirical demonstration. The power issue is concrete and checkable: null results from 24 participants, 12 of whom returned, cannot support 'mostly holds' as evidence of reliability. I also considered whether the case study only demonstrates agreement or reliability rather than accuracy; this is a real limitation of the demonstration, but the paper explicitly frames the audit as exploratory and the qualitative framework is still illustrated. The fix is evidentiary rather than conceptual: report equivalence bounds, effect sizes, or a power analysis, and soften the Section 4.5 conclusions. Thus the reader's CONDITIONAL verdict is appropriate and no change is needed.","tokens_in":23673,"tokens_out":7178,"duration_ms":79133,"concrete_test":"Compute equivalence bounds and power for the facet coefficients in Table 2. Use the 12 two-session repeats to estimate the within-subject repeatability of each BSP (e.g., within-subject SD, or the repeatability coefficient of the tape measure); set the equivalence bound for each regression slope to a value corresponding to that repeatability. Then re-test each facet coefficient with a 90% confidence interval (two one-sided tests). If any facet coefficient in the H2 or H3 models has a CI that exceeds the bound, or if a simulation with N=24/12 shows less than 80% power to detect a slope equal to that bound, the 'mostly holds' language is unsupported and should be replaced with 'no significant difference detected'.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that contextual auditing plus symmetry makes audits possible without agreed ground truth. The case study is the only empirical demonstration of this claim, and the quantitative core of that demonstration (Section 4.4) defines reliability as 'regression coefficients that are not statistically significantly different from zero (p≥0.05, corrected)'. The audit used 24 participants, with only 12 participants measured in a second session, and then fit nested regressions across six body parts and multiple facets. In this design the statistical power to detect a slope of the size that would matter for BSP measurement is low; for the sex facet there are 11 vs 13 participants, and for the time facet the effective sample is 12. A null result in such a design does not establish that measurements are stable; it is also consistent with measurements being unreliable but the study being unable to detect it. Because Section 4.5 concludes that H0-H3 'mostly hold' or 'hold', this assumption is directly load-bearing: if it fails, the case study does not demonstrate the value of the proposed method. The paper's own Limitations section acknowledges the small sample, but the interpretative claims in Section 4.5 still rely on the fragile null-as-reliability inference.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces 'contextual audits,' a method for auditing AI systems that make measurements of people, grounded in social practice theory, and proposes that the STS concept of 'symmetry' lets auditors proceed when ground truth is unknown, unknowable, or contested. The authors argue that audits should evaluate system outputs within the specific practices that produce them, making explicit what serves as ground truth rather than assuming a single objective standard. They demonstrate the method through a case study of skeleton inference in a marker-based motion capture system, comparing mocap-derived body segment parameters with tape-measure anthropometry across facets of body size, sex, and time, using Bland-Altman-style nested regressions. The paper reports that the hypotheses that measurements are reliable across these facets 'mostly hold' or 'hold,' while emphasizing that the results are exploratory and the sample is small.","tokens_in":23883,"tokens_out":4143,"duration_ms":45854,"significance":"If the conceptual framework holds, the paper makes a useful contribution to AI auditing by connecting auditing practice to social practice theory and STS symmetry, and by offering a worked example in a domain where ground truth is genuinely contested. The authors deserve credit for pre-registering the analysis, releasing code and data, documenting their qualitative fieldwork in detail, and stating limitations candidly, including the non-expert status of the measurers. The case study is valuable as an illustration of how an audit can be structured when no single ground truth exists. However, the empirical demonstration is substantially weakened by the operationalization of reliability as non-significance in a low-powered design, and the paper's central claim that the case study demonstrates the method's value depends on this fragile inference.","major_comments":[{"comment":"The operationalization of reliability as 'regression coefficients that are not statistically significantly different from zero (p≥0.05, corrected)' treats a null result as evidence of stability. With 24 participants and only 12 repeat sessions, the statistical power to detect meaningful deviations is low, particularly for the time facet and for the sex facet with 11 vs. 13 participants. A non-significant coefficient is compatible with both reliable and unreliable measurements, so the conclusions in Section 4.5 that H0-H3 'mostly hold' or 'hold' are not supported as stated. Because the case study is the only empirical demonstration of the proposed method, this issue is load-bearing; the authors should either add equivalence testing with pre-specified bounds, report effect sizes and confidence intervals, provide a power analysis, or explicitly recast the quantitative results as merely illustrative of the workflow rather than as evidence for reliability.","section":"Section 4.4 (Data Analysis), Table 2, and Section 4.5 (Audit Results)"},{"comment":"The role of symmetry in the quantitative analysis is underspecified. Section 4.2 states that neither the tape measure nor the mocap system should be taken as ground truth, while Section 5 says the authors used the tape measure 'as a provisional ground truth' and that under symmetry 'each measurement modality can serve as ground truth for the other.' If symmetry is a matter of interpretation rather than a change to the statistical model, the paper should say so explicitly, because the regression models in Section 4.4 are standard Bland-Altman limits-of-agreement models with added covariates and are indistinguishable from a conventional concurrent validity analysis. Without this clarification, the distinct methodological contribution of symmetry relative to existing practice is not clear.","section":"Section 4.2 (Incorporating Symmetry) and Section 5 (Discussion)"},{"comment":"The case study's claim to demonstrate a contextual audit is limited by the fact that the researchers, not expert practitioners, operated the mocap system. The Limitations section acknowledges that marker placement errors produced by non-experts are likely larger than those of the practitioners whose practices the study sought to replicate, which means the audit interrogates 'the system and its operators (us)' rather than the system in its real context of use. This is an honest and important admission, but it undercuts the strength of the demonstration; the paper should temper the contribution claim accordingly or include a comparison with expert operators to bound the effect of operator skill.","section":"Section 4.3 (Audit Protocol) and Limitations"}],"minor_comments":[{"comment":"Several rows report p-values and adjusted p-values only as ranges (e.g., 'all vars > 0.4'), which makes it difficult to assess the magnitude of the null results; exact values for all coefficients and their uncorrected and corrected p-values should be reported.","section":"Appendix D, Table 2"},{"comment":"The Benjamini-Hochberg correction is applied to a set of nested, non-independent regression models; the paper should justify this or acknowledge that the correction is approximate, since the assumption of independence is not met.","section":"Section 4.4"},{"comment":"The terms 'nominal outputs,' 'ground truth,' and 'provisional ground truth' are used in ways that may confuse readers; a concise glossary or a sentence clarifying how they relate would improve readability.","section":"Section 2.1 and Section 5"}],"recommendation":"major_revision","confidential_remarks":"The conceptual contribution is promising and the case study is thoughtfully designed in its qualitative components, but the quantitative demonstration rests on a null-as-reliability inference that the sample cannot support. I would encourage the editor to seek a revision that either supplies appropriate inferential methods (equivalence bounds, confidence intervals) or explicitly downgrades the empirical conclusions to a process illustration."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this paper introduces a real method, not just a critique. The contextual audit—five qualitative steps before any quantitative testing, using the socio-technical matrix to surface assumptions and facets—is a practical response to the contested-ground-truth problem that Jacobs and Wallach flagged. The adaptation of STS symmetry to treat measurement modalities as provisionally true is also new in this literature. The writing is clear and the framework is usable.\n\nWhat it does well: the qualitative grounding is genuinely ethnographic, not decorative. They trained with practitioners, followed local protocols, identified assumptions, and made those assumptions explicit in a matrix. The statistical analysis is pre-registered, code and data are public, and the authors repeatedly flag that results are exploratory. That is honest, reproducible work and deserves credit.\n\nThe soft spot is the empirical demonstration. Reliability is operationalized as regression coefficients not significantly different from zero after correction. With 24 participants, 12 of whom returned for a second session, the power to detect a meaningful slope is low. So the 'mostly holds' conclusions in Section 4.5 are consistent with unreliable measurements that the study simply couldn't detect. Section 4.5 is the only place where the method's value is demonstrated, so this is not a minor detail. The authors do acknowledge small sample and non-expert measurers in the Limitations, and they call the results exploratory—that softens the issue, but it doesn't fix the inference from null to reliability. The reader's stress test is right about this.\n\nWhere I push back a little: the conceptual argument does not depend on these specific results. The method could still be valuable even if the case study only shows how to run such an audit. So I would not treat the low power as fatal to the contribution. But a serious revision should report effect sizes, confidence intervals, or a power analysis, and should stop using 'holds' language for null results.\n\nBottom line: this is a paper for the AI auditing and FAccT community. It deserves peer review, not desk rejection. I'd want the authors to fix the reliability operationalization before publication. The citation pattern looks fine; heavy self-citation is justified because the framework extends their prior work.","headline":"A genuinely useful conceptual contribution to AI auditing—contextual audit with STS symmetry—but the case study's quantitative support leans on low-powered null results.","tokens_in":24410,"tokens_out":1515,"would_cite":true,"duration_ms":15842,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces the contextual audit, a method for testing AI systems within the practices that produce their outputs, and shows that the science-and-technology-studies principle of symmetry lets auditors proceed when ground truth is…","keywords":["AI auditing","contextual audit","symmetry","ground truth","motion capture","skeleton inference","body segment parameters","measurement validity"],"falsifier":"Run a confirmatory audit of the same kind of marker-based motion capture system with a much larger, more diverse sample and an independent ground truth such as MRI-derived bone lengths; if the system is then shown to be systematically inaccurate for narrower shoulders or larger bodies where this audit found non-significant coefficients, the null-result-as-reliability convention is the load-bearing premise that fails.","tokens_in":1494,"feed_emoji":"🦴","tokens_out":2002,"duration_ms":92219,"temperature":0.7,"pith_summary":"The paper argues that AI audits only mean something when they test systems in the context of the actual practices that produce the outputs, and it introduces the contextual audit as a method for doing so. It claims that when the thing being measured has no agreed, directly observable ground truth—as with body segment lengths inferred from motion capture—an auditor can still work by treating each measurement modality as provisionally true and examining where the two modalities diverge, a stance borrowed from the science-and-technology-studies principle of symmetry. To show the method works, the authors audit a marker-based motion capture system against tape-measure anthropometry, testing whether inferred body segment lengths are stable across body size, sex, and time. A sympathetic reader would care because most audit methods assume a ground truth that many measurement-focused AI systems simply do not have, and this paper offers a way to hold such systems accountable anyway.","feed_headline":"Symmetry lets audits of AI run without a ground truth","feed_subtitle":"A motion-capture case study shows contextual audits surface hidden assumptions in how systems measure bodies.","key_machinery":"The central mechanism is symmetry, which lets each measurement modality serve as a provisional ground truth for the other; the audit protocol is organized by an adapted socio-technical matrix that directs the auditor to identify the objects of capture and inference, how objects are made legible to the system, how ground truth has historically been established, what assumptions those ground truths encode, and how those assumptions could cause harm. The quantitative engine is a nested-regression extension of the standard limits-of-agreement approach, regressing the difference between paired measurements on the average measurement and on facets such as BMI, weight, sex, and session, with reliability defined as coefficients not significantly different from zero.","core_discovery":"The central claim is that audits of AI systems that make measurements should be contextual: they should make claims about accuracy within the context of the practices that produce the outputs, and they should be explicit about what serves as ground truth. When ground truth is unknown, unknowable, or contested, auditors can apply symmetry, treating the outputs of each measurement modality as provisionally true and then interrogating the implications of that choice, rather than elevating one modality to ground truth. The demonstration shows that a marker-based motion capture system's body segment parameters are mostly stable across body size, sex, and time when compared with tape measurements, but tensions appear for shoulder breadth and standing height; symmetry, not a claimed ground truth, is what turns those tensions into audit findings.","pith_inferences":["The symmetry principle could generalize: any two imperfect measurement instruments probing the same unobservable quantity can be cross-audited without appointing a champion, provided their assumptions are disjoint enough that they do not silently share the same error.","A testable extension is to audit the same system against several different provisional ground truths; the tensions should differ, and the pattern of tensions would map the system's assumption landscape.","The case study's non-significant coefficients probably understate instability, because non-significance with a small sample is weak evidence of reliability; a larger, more diverse sample is the natural confirmatory test.","If the paper's reflexive point is taken seriously, audit reports should routinely separate system error from operator error, since non-expert auditors may inflate the apparent instability of a system built for expert use."],"forward_implications":["Auditors of measurement-based AI systems can state explicitly what they take as ground truth, and when none is available, they can audit by comparing modalities symmetrically instead of halting the audit.","The approach makes qualitative fieldwork a prerequisite for quantitative audit design, because the facets and ideals to test are derived from context, not from the system's specifications alone.","For motion capture, the case study identifies body size—narrow shoulders and larger bodies in particular—as the facet where tensions between mocap and tape measurements emerged and where future audits should focus.","Following the method would push AI auditing away from synthetic benchmark inputs toward inputs produced by real users following real protocols, which the paper argues is necessary for ecological validity."],"supporting_citations":[{"why":"Supplies the limits-of-agreement regression that the audit extends with nested facet variables.","marker":"Altman and Bland 1983"},{"why":"Origin of the symmetry principle used to avoid designating a single ground truth.","marker":"Bloor 1991"},{"why":"Documents the thin, mostly male body data behind mocap body models and motivates the body-size facet.","marker":"Harvey et al. 2024"},{"why":"Provides the anthropometric measurement protocol used as the tape-measure comparison.","marker":"DHHS 1994"},{"why":"Defines AI auditing as comparing actual to nominal outputs, the framing the contextual audit revises.","marker":"Sandvig et al. 2014"},{"why":"Supplies the measurement-theoretic account of ground truth and concurrent validity the paper engages.","marker":"Jacobs and Wallach 2021"},{"why":"Contributes the socio-technical matrix the paper adapts for measurement contexts.","marker":"Sloane, Moss, and Chowdhury 2022"},{"why":"Grounds the argument that inputs and nominal outputs must be meaningful in context.","marker":"Sloane et al. 2023"},{"why":"Correction for multiple comparisons used in the significance tests of regression coefficients.","marker":"Benjamini and Hochberg 2000"}],"fun_headline_variants":["Audit AI measurements with symmetry, not ground truth","Symmetry lets AI audits proceed despite unknown ground truth","Contextual audits reveal assumptions in motion capture AI","Skeleton inference audit uses symmetry when truth is elusive","No ground truth? Symmetry is the auditor's workaround"],"cache_read_input_tokens":26624,"weakest_assumption_plain":"The demonstration treats a regression coefficient that is not statistically significant (corrected $p\\geq 0.05$) as evidence that measurements are reliable, but with 24 participants and only 12 repeat sessions the tests are low-powered, so a null result cannot confirm stability across body size, sex, or time.","fun_headline_variants_meta":{"raw":{"variants":["Audit AI measurements with symmetry, not ground truth","Symmetry lets AI audits proceed despite unknown ground truth","Contextual audits reveal assumptions in motion capture AI","Skeleton inference audit uses symmetry when truth is elusive","No ground truth? Symmetry is the auditor's workaround"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000375,"raw_usage":{"total_tokens":1979,"prompt_tokens":903,"completion_tokens":1076,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":519,"completion_tokens_details":{"reasoning_tokens":999}},"tokens_in":519,"tokens_out":1076,"duration_ms":10517,"temperature":1.0,"reasoning_tokens":999,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:10:18.086257+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a confirmatory audit of the same kind of marker-based motion capture system with a much larger, more diverse sample and an independent ground truth such as MRI-derived bone lengths; if the system is then shown to be systematically inaccurate for narrower shoulders or larger bodies where this audit found non-significant coefficients, the null-result-as-reliability convention is the load-bearing premise that fails.","supporting_citations":[],"review_version":1}