{"id":"f4474b31-9b63-4120-aea0-906e7af26d24","arxiv_id":"2606.20820","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"CELEUS applies e-processes to LLM evaluation to produce anytime-valid CIs that reach target precision with 54-62% fewer samples than baselines while preserving coverage guarantees.","lead":"CELEUS uses e-processes to create anytime-valid confidence intervals for LLM performance scores, combining uncertainty-guided sampling and surrogate approximations to cut the number of evaluated samples. A smart generalist might read it to see a statistically grounded way to make LLM benchmarks more efficient and trustworthy even when stopping early based on results.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest_assumption matches the load-bearing mathematical step. Because the query provides only the abstract plus a placeholder for the full text, no further technical flaw can be isolated; the non-finding is therefore honest rather than manufactured.","tokens_in":1816,"tokens_out":237,"duration_ms":35695,"concrete_test":"Locate the section containing the unbiasedness proof (likely the main theorem on signals); re-derive the conditional expectation step-by-step, confirming that the sampling probabilities (which depend on past uncertainty estimates) and any surrogate parameters (fitted on past data) do not introduce bias in the expectation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on proving that the uncertainty-guided sampling plus surrogate signals remain conditionally unbiased for the evaluation score given the past. The abstract states this property explicitly and claims it enables the e-process construction. No internal inconsistency, hidden assumption in the sampling rule, or flaw in the variance-reduction argument is visible from the given material; the reader's note already isolates exactly this point as requiring full-text verification of the proof.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper proposes CELEUS, a framework for certifiable LLM evaluation that uses e-processes to construct anytime-valid confidence intervals. It introduces composite signals that combine uncertainty-guided sampling of informative examples with surrogate-assisted approximations for unevaluated items, proves that these signals are conditionally unbiased for the evaluation score given the past, establishes near-parametric shrinkage rates (up to log factors) for the resulting CIs, and reports that the method reaches target precision with 54-62% fewer evaluated samples than baselines while preserving coverage guarantees.","tokens_in":1911,"tokens_out":436,"duration_ms":22535,"significance":"If the conditional-unbiasedness claim holds, the work supplies a statistically rigorous, anytime-valid alternative to existing sequential LLM evaluation procedures. The combination of e-process theory with practical variance-reduction techniques (uncertainty sampling plus surrogates) and the reported sample savings would be a meaningful contribution to efficient, certifiable evaluation, especially given the high cost of human or model-based scoring of large test sets.","major_comments":[{"comment":"The central load-bearing claim is the proof that the uncertainty-guided sampling plus surrogate signals remain conditionally unbiased for the evaluation score given the past (abstract and §3). Because this property is what licenses the e-process construction and the anytime-valid coverage, the derivation must be checked in full for any dependence introduced by the sampling rule or the surrogate training that could violate the martingale property.","section":"§3 (proof of unbiasedness)"}],"minor_comments":[{"comment":"The experimental section should report the precise definition of the target precision (e.g., CI width) and the stopping rule used in the anytime-valid setting so that the 54-62% reduction claim can be reproduced.","section":"Experiments"},{"comment":"Notation for the surrogate model and the uncertainty measure should be introduced once and used consistently; several symbols appear to be overloaded between the sampling rule and the e-process construction.","section":"§2-3"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the positive assessment and recommendation for minor revision. We address the sole major comment below.","responses":[{"response":"We appreciate the referee underscoring the centrality of this property. In §3 the proof proceeds by verifying that the composite signal S_t satisfies E[S_t | ℱ_{t-1}] = \theta (the target score) for the filtration ℱ_{t-1} generated by all prior evaluations and the surrogate parameters fitted on them. The uncertainty-guided sampling probabilities are ℱ_{t-1}-measurable by construction, and the surrogate approximation is likewise a deterministic function of ℱ_{t-1}. Consequently the product of the sampling indicator and the (surrogate or true) score remains a martingale difference; no extra dependence is introduced that would break the conditional unbiasedness. We have re-examined every step of the derivation and confirm it holds. In a minor revision we will insert a short paragraph after the main proof explicitly stating the measurability of each component with respect to ℱ_{t-1}.","revision_made":"partial","referee_comment":"[§3 (proof of unbiasedness)] The central load-bearing claim is the proof that the uncertainty-guided sampling plus surrogate signals remain conditionally unbiased for the evaluation score given the past (abstract and §3). Because this property is what licenses the e-process construction and the anytime-valid coverage, the derivation must be checked in full for any dependence introduced by the sampling rule or the surrogate training that could violate the martingale property."}],"tokens_in":1390,"tokens_out":347,"duration_ms":16727,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that this paper shows how to get confidence intervals for LLM performance scores that stay valid no matter when you decide to stop evaluating. It does so by feeding e-processes with signals built from uncertainty-guided sampling and surrogate approximations for the rest of the items, and the authors state they prove those signals remain unbiased conditional on the past.\n\nWhat stands out is the direct application to the sequential evaluation setting. Existing certifiable methods update CIs but lose the coverage guarantee under repeated peeking; tying e-process theory to this domain with the two added ingredients is the concrete step forward. The paper also supplies an oracle analysis for the variance-optimal rule and shows the empirical version gets close, plus the near-parametric shrinkage rate. The reported efficiency numbers are large enough to matter for anyone running repeated benchmarks.\n\nThe soft spot is exactly the unbiasedness claim under the adaptive rule. The abstract presents it as proven, but any dependence introduced by the surrogate or the way uncertainty is estimated could affect the conditional property, and that needs the full derivation to confirm. The experiments preserve coverage while cutting samples, yet the strength of that result depends on how the baselines were implemented and how sensitive the surrogates are to their training data.\n\nThis is for researchers who work on statistical guarantees for model evaluation and want practical anytime-valid tools. It is worth sending to peer review because the core technical move is non-routine, the efficiency claim is substantial, and the theoretical framing is clear enough for referees to check the proofs directly.","headline":"Celeus applies e-processes to LLM eval for anytime-valid CIs by proving conditional unbiasedness under uncertainty sampling plus surrogates, with reported 54-62% sample cuts.","tokens_in":2439,"tokens_out":386,"would_cite":true,"duration_ms":30869,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Celeus constructs anytime-valid confidence intervals for LLM evaluation by proving that uncertainty-guided sampling signals combined with surrogate approximations remain unbiased conditional on past observations.","keywords":["LLM evaluation","e-processes","anytime-valid CIs","uncertainty sampling","surrogate approximation","certifiable evaluation","sequential sampling"],"falsifier":"Observing that the empirical coverage of the constructed confidence intervals falls below the nominal level, such as 95%, in repeated experiments where stopping decisions are made adaptively based on the intervals.","tokens_in":2712,"feed_emoji":"📊","tokens_out":656,"duration_ms":30530,"temperature":0.7,"pith_summary":"The paper aims to close the gap between theoretical guarantees and practical use in LLM evaluation by developing methods that provide valid confidence intervals even when the evaluation process adapts based on intermediate results. Existing approaches can lose their coverage guarantees if intervals are updated repeatedly to decide when to stop sampling more examples. Celeus addresses this by introducing signals that merge uncertainty-guided selection of samples with approximations for unevaluated ones, proving these signals are unbiased given the history. This unbiasedness supports the use of e-processes to form intervals that are valid at any stopping time. If correct, evaluators can confidently halt testing once a desired precision is achieved, knowing the reported interval truly covers the model's performance with the stated probability.","feed_headline":"Anytime-valid CIs for LLM evaluation achieved with 54-62% fewer samples","feed_subtitle":"Unbiasedness of uncertainty-guided and surrogate signals conditional on past data supports early stopping without coverage loss.","key_machinery":"Unbiased signals from uncertainty-guided sampling and surrogate-assisted approximations that enable e-process based anytime-valid confidence intervals.","core_discovery":"The central discovery is that signals combining uncertainty-guided sampling to select informative samples for evaluation and surrogate-assisted approximations for unevaluated samples remain unbiased for the evaluation score conditional on the past. This property enables the construction of statistically-grounded and anytime-valid e-process confidence intervals. The approach also reduces estimation variance, allowing the target precision to be reached with fewer evaluated samples, and the intervals shrink at a near-parametric rate up to logarithmic factors.","pith_inferences":["Similar unbiased signal constructions might apply to sequential testing in other machine learning tasks like model selection or hyperparameter tuning.","The efficiency gains could make large-scale LLM benchmarking more feasible under resource constraints.","If extended, this might provide a template for certifiable evaluation in non-LLM settings where adaptive sampling is used."],"forward_implications":["The CIs remain valid regardless of when the evaluation stops based on the intervals themselves.","Target precision is reached with 54-62% fewer evaluated samples than baselines.","Confidence intervals shrink at near-parametric rates up to logarithmic factors.","An oracle variance-optimal sampling rule exists that motivates the empirical uncertainty-guided approach."],"fun_headline_variants":["Celeus delivers 54-62% sample savings in anytime-valid LLM evaluation","Anytime-valid CIs from E-processes need 54-62% fewer LLM samples","Unbiased signals enable anytime-valid LLM CIs with reduced evaluations","E-processes support early stopping in certifiable LLM evaluation"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The signals combining uncertainty-guided sampling to select informative samples and surrogate-assisted approximations for unevaluated samples remain unbiased for the evaluation score conditional on the past.","fun_headline_variants_meta":{"raw":{"variants":["Celeus delivers 54-62% sample savings in anytime-valid LLM evaluation","Anytime-valid CIs from E-processes need 54-62% fewer LLM samples","Unbiased signals enable anytime-valid LLM CIs with reduced evaluations","E-processes support early stopping in certifiable LLM evaluation"]},"model":"grok-4.3","cost_usd":0.006543,"raw_usage":{"total_tokens":3016,"prompt_tokens":744,"num_sources_used":0,"completion_tokens":77,"cost_in_usd_ticks":65428000,"prompt_tokens_details":{"text_tokens":744,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2195,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":744,"tokens_out":77,"duration_ms":17395,"temperature":1.0,"reasoning_tokens":2195,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T04:46:22.693366+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Observing that the empirical coverage of the constructed confidence intervals falls below the nominal level, such as 95%, in repeated experiments where stopping decisions are made adaptively based on the intervals.","supporting_citations":[],"review_version":2}