{"id":"4d30377b-bcbd-4c7e-9c71-388fbdfcfd92","arxiv_id":"2512.15783","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A Logia grammar is introduced that turns expert–AI interactions into eight standardised fields, enabling population-level surveillance of AI output failures without access to model internals.","lead":"This paper introduces a standardised framework for recording expert–AI interactions, compressing them into structured fields so that alignment and accuracy scores can be tracked across deployments. If the reliability claims survive larger validation, institutions could detect risky AI outputs from observed patterns of expert overrides rather than from model internals.","discovery_kind":"new_method","skeptic_critique":null,"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes 'AI Epidemiology', a governance framework that treats expert–AI interactions as population-level surveillance data. It introduces the Logia Grammar (mission, conclusion, justification, risk level, alignment score, accuracy score, override, corrective option) and Tracelayer as a pattern-analysis layer. Exposure variables (alignment and accuracy) are meant to predict output failure, operationalized as expert override or adverse outcome. The paper reports a feasibility study in which one ophthalmologist reviewed Logia's automated (NotebookLM/RAG-based) analysis of three GPT-5-generated ophthalmology cases, yielding 94% raw agreement and an ICC of 0.89 for 'measurement standardisation'. It also contains a chess demonstration contrasting SHAP with Logia, and a discussion of challenges such as expert entrenchment and commercial bootstrap. The abstract claims that, under bounded conditions, LLMs can produce reliable standardized assessments of expert–AI interactions, and describes a statistical protocol involving paired bootstrap, DeLong's test, a non-inferiority margin of 0.05, and Holm–Bonferroni correction. The core of the paper is conceptual; the empirical contribution is a small pilot.","tokens_in":20371,"tokens_out":4494,"duration_ms":43265,"significance":"If the framework could be shown to produce reliable, standardized measurements of AI-output risk from passively captured expert interactions, it would offer a genuinely model-agnostic governance tool with applications in regulated domains like healthcare, finance, and law. The paper's strength is its explicit staged research programme and its candid acknowledgement of limitations: it distinguishes feasibility from population-level validation, lists what cannot yet be tested, and specifies Phase 2/3 requirements. The proposed grammar is simple and concrete, and the use of published clinical guidelines as RAG documents is a reproducible starting point. However, the current evidence is far too thin to support the abstract's reliability claim, and several conceptual issues — particularly the circular use of expert overrides in both calibration and outcome definition — remain unresolved. As a concept proposal with an honest pilot, it is worth pursuing; as a demonstration of measurement reliability, it is not yet convincing.","major_comments":[{"comment":"The central claim that 'this study demonstrates that standardised measurement of AI outputs is feasible with 89% inter-rater reliability (ICC = 0.89), achieving good reliability for epidemiological analysis' is not supported by the reported design. Three cases and one expert cannot estimate population-level inter-rater reliability: there is no sampling frame for cases or experts, no confidence intervals around ICC, and with only three items per field a single disagreement changes ICC from 1.0 to 0.67. The study should be described as a pilot feasibility check, not as evidence of measurement reliability. Any claim of 'good reliability' should be deferred to Phase 2 with appropriate uncertainty quantification.","section":"Section 4, 'Study design' and 'Discussion'"},{"comment":"The 94% agreement and ICC = 0.89 pool semantic-capture fields (mission, conclusion, justification) with measurement-standardisation fields (risk, alignment, accuracy). Semantic capture is an information-preservation check, not an inter-rater reliability estimate. The 'lossless semantic compression' claim from 3/3 agreement on three cases is also an overreach; absence of observed information loss in three hand-picked cases does not establish losslessness. The reported ICC should be computed and reported separately for the three assessment fields, with model form (e.g., two-way random, absolute agreement), confidence intervals, and per-field ICCs.","section":"Table 1 and 'Measurement standardisation' subsection"},{"comment":"There is a circularity in the proposed reliability score. The same expert overrides are used both to recalibrate the alignment and accuracy exposure variables (e.g., the mortgage example where overrides lead to recalibration) and to define output failure in the reliability score. If scores are calibrated to predict expert overrides and then used to predict expert overrides, the apparent predictive power is partly tautological. The paper's proposed mitigation — outcome tracking — is not implemented in the feasibility study and remains future work. The manuscript should specify how an independent outcome signal would break this circularity before the reliability score is used prospectively.","section":"Section 3, 'Reliability score: predicting failure probability'; Section 6, 'Expert judgment as outcome variable'"},{"comment":"The abstract promises a statistical protocol consisting of paired bootstrap inference, DeLong's test for paired AUCs, a pre-specified one-sided non-inferiority margin of 0.05, and Holm–Bonferroni correction. None of these appears in Section 4's Methods or Results. Either the protocol should be added and the corresponding analyses reported for the feasibility data, or the abstract must be revised to reflect what was actually done. As written, the abstract claims a level of statistical rigour that the body does not deliver.","section":"Abstract vs. Section 4"}],"minor_comments":[{"comment":"The table reports ICC = 1.0 and ICC = 0.67 for fields with n = 3 cases. No confidence intervals or ICC model details are given. Please report the ICC variant (e.g., two-way random-effects, absolute agreement) and bootstrap/CI estimates.","section":"Table 1"},{"comment":"The Tracelayer statistics ('85% consensus, 165 similar cases') are explicitly simulated in the footnote, but this is easy to miss. The demonstration would be clearer if the simulated nature were stated in the main text and the table were marked as illustrative.","section":"Section 5, 'Logia structured analysis'"},{"comment":"The reliability score combination rule ('taking the lower value') is introduced without justification or citation. If this is a placeholder default, say so; if it is a substantive modeling choice, it needs a rationale and a sensitivity analysis.","section":"Footnote 3"},{"comment":"Many reference strings contain garbled tokens (e.g., 'NeurIPS.8686', 'arXiv¿8❶6❶¡79❸9❸', 'Proc.0th.Workshop'). Please clean the reference list and verify DOIs/arXiv identifiers.","section":"References"},{"comment":"The arXiv metadata title ('Towards AI epidemiology: a measurement standardisation framework for prospective risk detection') differs from the full-text title ('AI Epidemiology: achieving explainable AI through expert oversight patterns'). Please harmonize.","section":"Title and metadata"},{"comment":"The historical epidemiological passages (Bradford Hill, Goldberger, Framingham) are long relative to their technical contribution. Condensing them would improve readability without affecting the argument.","section":"Section 2"}],"recommendation":"major_revision","confidential_remarks":"The paper is best read as a concept proposal with a pilot illustration. The author is transparent about the small scale, but the abstract and Section 4 Discussion overstate the strength of the evidence. The most important fix is to reframe the feasibility study as a pilot and to remove or substantially qualify the reliability claim. The circularity of using overrides both as calibration targets and as outcome definitions is a substantive conceptual concern that needs to be addressed before the framework is tested at scale. I recommend major revision rather than rejection because the framework itself is coherent and the staged validation plan is reasonable; however, the current manuscript does not support the abstract's 'demonstrates' language."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague — quick take: read this as a concept paper with an honest staged plan, not as a demonstrated measurement standard. The abstract says the judge 'demonstrates' reliable assessment; the body's only empirical support is three ophthalmology cases, one expert, and an agreement table. That gap is the paper's main problem.\n\nWhat's actually new: the Logia grammar (mission, conclusion, justification; risk, alignment, accuracy; override, corrective option) as a standardised output-level schema, Tracelayer as a population-pattern layer, and the explicit 'AI epidemiology' framing. The distinction between output-level surveillance, correspondence-based interpretability, and aggregate statistical monitoring is clean and useful. The staged protocol — Phase 2 at 500+ cases with outcome tracking — is sensible. The limitations section is genuinely honest, and the discussion of expert entrenchment and outcome validation is a real strength. I also think the citations are broadly sound; nothing looks like citation-padding.\n\nNow the soft spots, in order of size. The reliability claim is load-bearing and unsupported. The ICC of 0.89 is computed from eight field-level items (three each for risk/accuracy, two for alignment) comparing the RAG process to a single expert. There are no confidence intervals, no second expert, no independent runs. The 94% 'inter-rater reliability' lumps semantic-capture checks (which are representation checks, not reliability) together with measurement-standardisation items. And the abstract promises a statistical protocol — paired bootstrap, DeLong, non-inferiority margin, Holm-Bonferroni — that does not appear in the body; these look like two versions of the paper stitched together.\n\nThe reliability-score design also has a circular flavour the paper itself acknowledges: scores recalibrate from expert overrides, then are used to predict expert overrides, with outcome tracking only promised as future work. The chess demonstration is illustrative, not evidence. No code or data are shipped; 'available upon request' is weak for a measurement standard.\n\nNone of this kills the framework as a proposal. The core idea — output-level risk stratification for expert-AI interactions — is concrete, testable, and worth taking seriously. It just needs proper validation before anyone cites it as established.\n\nWho this is for: AI governance teams, interpretability researchers, and people running deployed systems who want a model-agnostic oversight layer. I'd send it to peer review rather than desk reject, but the empirical section needs a major overhaul: label it an illustrative pilot, release the RAG stack and scoring outputs, and run a real multi-rater reliability study with CIs. Until then, I wouldn't cite the reliability result.","headline":"Useful framework proposal, but the 'reliability demonstrated' claim rests on three cases and one expert; treat it as a pilot, not a result.","tokens_in":20735,"tokens_out":5051,"would_cite":false,"duration_ms":49062,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper proposes governing AI systems the way epidemiology governs disease: standardise how expert-AI interactions are measured, treat expert overrides as outcome events, and flag risky outputs before they cause harm, with a feasibility s","keywords":["AI epidemiology","measurement standardisation","AI governance","retrieval-augmented generation","expert oversight","risk stratification","inter-rater reliability","model-agnostic monitoring"],"falsifier":"Have a second independent expert score the same three published ophthalmology cases using the same RAG-generated fields; if the alignment-score disagreement in Case 2 is not resolved or new disagreements appear, the ICC = 0.89 estimate does not replicate. More decisively, run the protocol with multiple experts across dozens of diverse cases and compute the intraclass correlation; if agreement on alignment score falls below the moderate threshold, the standardisation claim fails.","tokens_in":20243,"feed_emoji":"🩺","tokens_out":4584,"duration_ms":46760,"temperature":0.7,"pith_summary":"The paper argues that the most practical way to govern opaque AI systems is not to open the model but to watch what experts do with its outputs. It proposes a standardised grammar that compresses every expert-AI interaction into eight fields, and uses retrieval-augmented generation to fill the assessment fields against institutional guidelines. Treating those assessments as exposure variables and expert overrides plus real-world outcomes as outcome variables, it claims institutions can detect unreliable AI outputs prospectively, before harm occurs. A small feasibility test with three ophthalmology cases and one expert reports perfect semantic capture and an intraclass correlation of 0.89 between the automated judge and the expert. The paper frames the population-level claims as a staged research programme rather than as results already established.","feed_headline":"Standardised AI-risk scoring hits 89% inter-rater reliability","feed_subtitle":"A proposed eight-field grammar turns expert overrides into population-level signals for spotting dangerous AI outputs early.","key_machinery":"The Logia Grammar: an eight-field schema (mission, conclusion, justification, risk level, alignment score, accuracy score, override, corrective option) that compresses expert-AI interactions into comparable records. The assessment fields are populated by retrieval-augmented generation (RAG) against institutional documents, then dynamically recalibrated through a triple-signal system: RAG gives the initial assessment, expert overrides give medium-term validation, and tracked outcomes give the most reliable long-term signal. This turns alignment and accuracy into exposure variables, with override and adverse outcomes as outcome variables, and a composite reliability score that predicts output","core_discovery":"The central claim is that, under bounded conditions, large language models can produce reliable, standardised assessments of the evidential and policy alignment of expert-AI interactions. The study reports 89% inter-rater reliability (ICC = 0.89) across risk level, alignment score, and accuracy score, with 100% agreement on semantic capture of mission, conclusion, and justification. The larger claim is that once such standardised measurement exists, alignment and accuracy scores become exposure variables that predict output failure, enabling an 'AI epidemiology' that acts on statistical patterns before mechanistic understanding is available.","pith_inferences":["As an editorial extension: the reported ICC of 0.89 is computed from only three cases against a single expert, so the true reliability of the judge across the diversity of deployment is likely to be substantially lower; a multi-expert, many-case study would give a more honest estimate.","As an editorial extension: the judge is itself a large language model, so the framework has a potential circularity problem — an unreliable AI system is being used to score the reliability of other AI systems; comparing RAG-generated scores against a blinded human panel on adversarial cases would test how much this matters.","As an editorial extension: the chess demonstration uses simulated population statistics (85% consensus, 165 similar cases) rather than real accumulated data, so it illustrates the grammar's explanatory format but does not yet show that pattern recognition works in practice."],"forward_implications":["If the reliability result holds at scale, institutions can flag low-reliability AI outputs for mandatory review before they are acted on, shifting oversight from post-hoc correction to pre-hoc triage.","Automatic audit trails can be built from passive monitoring of expert-AI interactions, with zero data-entry burden on experts and no need for model-internal access.","Governance can survive model updates and vendor switches because the standardised assessments attach to observable outputs, not to proprietary internals.","The corrective-option field enables semantic explanations, such as 'similar outputs were overridden 71% of the time because they violated triage protocols,' and can generate targeted retraining datasets.","Outcome tracking can expose and correct systematic expert bias when expert overrides diverge from real-world consequences, preventing dogmatic reliance on expert consensus."],"fun_headline_variants":["AI epidemiology framework standardises expert-AI risk checks","89% reliability in standardising expert-AI alignment scoring","Proposed framework turns expert-AI logs into risk signals","AI risk detection via standardised overrides: 89% agreement","From expert-AI interactions to AI epidemiology: new framework"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that one expert's agreement with the automated judge on three ophthalmology cases is an informative estimate of how reliably the judge would assess the full diversity of expert-AI interactions in deployment.","fun_headline_variants_meta":{"raw":{"variants":["AI epidemiology framework standardises expert-AI risk checks","89% reliability in standardising expert-AI alignment scoring","Proposed framework turns expert-AI logs into risk signals","AI risk detection via standardised overrides: 89% agreement","From expert-AI interactions to AI epidemiology: new framework"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000196,"raw_usage":{"total_tokens":1225,"prompt_tokens":801,"completion_tokens":424,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":352}},"tokens_in":545,"tokens_out":424,"duration_ms":5433,"temperature":1.0,"reasoning_tokens":352,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T16:28:08.295681+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have a second independent expert score the same three published ophthalmology cases using the same RAG-generated fields; if the alignment-score disagreement in Case 2 is not resolved or new disagreements appear, the ICC = 0.89 estimate does not replicate. More decisively, run the protocol with multiple experts across dozens of diverse cases and compute the intraclass correlation; if agreement on alignment score falls below the moderate threshold, the standardisation claim fails.","supporting_citations":[],"review_version":1}