{"id":"0ce7b508-b70f-49b9-92d3-f0fcb9bca733","arxiv_id":"2604.00660","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.5,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"SUPG-IT and GAMCAL route streaming semantic-SQL rows through cheap proxies with joint precision/recall guarantees or a single cost-error tradeoff, cutting oracle calls while keeping high F1.","lead":"The paper proposes two streaming model-cascade algorithms that cut expensive LLM calls in semantic SQL while targeting precision and recall. They matter because warehouses now run per-row LLM operators that dominate query cost, and prior cascades block pipelines or cannot guarantee both metrics.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Provided full text is an unrelated paper; cascade claims remain unsupported.","rationale":"The reader correctly flagged the manuscript mismatch and issued UNVERDICTED / LOW confidence. The supplied full text does not contain any of the cascade algorithms, proofs, or production-engine results asserted in the abstract and strongest_claim; it is a different paper on spatial contact processes. No internal analysis of threshold refinement, monotone GAM calibration, or transfer of batch-local statistics to the full stream is available to stress-test. The load-bearing issue is therefore still the complete absence of supporting material, not a subtle flaw inside a present argument. Verdict remains UNVERDICTED until the correct manuscript is supplied; no adjustment is warranted.","tokens_in":11550,"tokens_out":458,"duration_ms":11827,"concrete_test":"Retrieve the actual arXiv:2604.00660 PDF/source; confirm presence of (i) a theorem stating joint (precision, recall) guarantees for SUPG-IT under streaming parallel workers and (ii) the six-benchmark tables reporting F1 and oracle counts. If either is absent or the body is the pedestrian paper, the claims stay unverified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The strongest claim (SUPG-IT as first streaming cascade with joint probabilistic precision/recall guarantees at failure probability δ, plus GAMCAL/SUPG-IT F1 ≥ 0.95 and up to 58% oracle savings on six semantic-SQL benchmarks) has no supporting content in the supplied manuscript. The CACHEABLE full text is Gambaudo & Génois (arXiv:2604.00652) on 2D random-walk pedestrian models and contact-number distributions; it contains no formalization of streaming cascade routing, no SUPG-IT threshold iteration, no GAM calibration, no δ-guarantees, and no semantic-SQL experiments. Consequently the joint-guarantee and F1/oracle claims rest solely on the abstract and cannot be checked for validity under independent parallel workers or batch-local statistics.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The submission’s abstract claims two streaming cascade algorithms for semantic SQL (SUPG-IT and GAMCAL) that avoid a global proxy-score pass under independent parallel workers: SUPG-IT iteratively refines dual thresholds from batch oracle labels with joint probabilistic precision/recall guarantees at failure probability δ, and GAMCAL uses a monotone GAM to calibrate proxy scores to true-positive probabilities with pointwise uncertainty for α-tradeoff stochastic routing. It further claims F1 ≥ 0.95 on six production benchmarks, leadership at a 20% delegation budget, up to 58% fewer oracle calls than LOTUS SUPG, and mean best-case F1 of 0.989 for SUPG-IT. The body supplied with this review, however, is an unrelated manuscript (Gambaudo & Génois) on 2D random-walk pedestrian models and contact-number distributions; it contains no cascade formalization, no SUPG-IT/GAMCAL algorithms, no δ-guarantees, and no semantic-SQL experiments.","tokens_in":11810,"tokens_out":818,"duration_ms":13279,"significance":"If the abstract’s claims were substantiated, the work would be significant for production semantic SQL: joint precision–recall streaming guarantees without a blocking global proxy pass, plus large oracle savings, would address a real cost bottleneck in LLM-augmented warehouses. Those contributions cannot be assessed from the material provided, because the manuscript body does not develop or evaluate them.","major_comments":[{"comment":"Manuscript–abstract mismatch: the full text is a physics/soc-ph paper on targeting rules for pedestrian contact networks (Figs. 1–4, RW/TW rates r,s, three target-choice mechanisms). It does not formalize streaming cascade routing, define SUPG-IT or GAMCAL, state joint (precision, recall, δ) guarantees, or report the six semantic-SQL benchmarks. Load-bearing claims in the abstract are therefore unsupported by the supplied manuscript and cannot be checked.","section":null},{"comment":"Central technical premise (abstract): that batch-local oracle accumulation and monotone GAM calibration remain valid under independent parallel streaming workers without a global proxy-score pass. No model, algorithm, proof sketch, or experiment for this premise appears in the provided body, so the joint-guarantee and F1/oracle-savings claims are unevaluable.","section":null},{"comment":"Empirical claims (F1 ≥ 0.95, 58% fewer oracle calls vs LOTUS SUPG, mean best-case F1 0.989, leadership at 20% budget on six datasets) have no corresponding tables, figures, datasets, or engine description in the supplied text. Without those results, the performance narrative cannot be verified or stress-tested.","section":null}],"minor_comments":[{"comment":"Title/abstract metadata (paper_id 2604.00660, cs.DB) do not match the body (arXiv:2604.00652-style pedestrian contact study). This should be corrected before any scientific review of the cascade work.","section":null},{"comment":"If the intended cascade manuscript is restored, the abstract’s free parameters (δ, α, dual thresholds, delegation budget) and the ‘first streaming cascade with joint guarantees’ priority claim will need explicit theorem statements and comparison baselines in the body.","section":null}],"recommendation":"uncertain","confidential_remarks":"The review package appears corrupted or mis-bound: abstract and paper_id describe a cs.DB cascade paper, while the full text is an unrelated physics manuscript on face-to-face contact distributions. I cannot produce a scientific accept/reject decision on 2604.00660 until the correct PDF is supplied. Please re-queue with the matching manuscript; no author-side scientific fault can be inferred from this packaging error alone."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: the full text in the packet is not this paper. It is Gambaudo & Génois on 2D random walks and face-to-face contact counts (arXiv:2604.00652). Everything load-bearing about SUPG-IT, GAMCAL, δ-guarantees, and the six semantic-SQL benchmarks lives only in the abstract of 2604.00660. I cannot verify the math, the streaming model with parallel workers, or the reported F1/oracle numbers.\n\nOn the abstract alone, the problem is real and well posed. Semantic operators make per-row LLM cost dominate SQL; prior SUPG-style cascades need a global proxy pass that blocks pipelines and usually optimize one of precision or recall. Framing cascade routing for streaming workers, then offering (i) iterative dual-threshold refinement with joint precision/recall guarantees and (ii) a monotone GAM calibration plus stochastic routing under a single cost–error tradeoff α, is a coherent systems contribution if the proofs and experiments hold. The production-engine evaluation and the head-to-head with LOTUS’s SUPG cascade are the right kind of evidence for this subfield.\n\nSoft spots, in proportion: without the manuscript we have zero check on whether batch-local oracle labels actually support the joint (precision, recall, δ) claim under independent workers, whether the monotone GAM calibration transfers across the stream, or how free parameters (δ, α, thresholds, 20% budget) were set. Those are not invented flaws; they are the claims the abstract hangs on and that the wrong PDF does not address. Circularity risk looks ordinary from the abstract (external benchmarks, not tautological metrics), but that is also uncheckable here.\n\nWho this is for: people building LLM-in-the-loop warehouses and selective-prediction cascades. A serious editor should send the real paper to referees if the full text matches the abstract’s formal and empirical claims. I would not cite or put this in reading group until we have the correct manuscript, code, and the guarantee proofs. Get the right PDF; then re-read.","headline":"Wrong manuscript was attached: we only have the abstract for the cascades paper, so the joint-guarantee and oracle-savings claims cannot be checked.","tokens_in":12373,"tokens_out":530,"would_cite":false,"duration_ms":9820,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Two streaming cascade algorithms cut expensive LLM calls in semantic SQL while hitting high precision and recall without a blocking global score pass.","keywords":["semantic SQL","model cascades","streaming query execution","proxy-oracle routing","precision-recall guarantees","generalized additive model calibration","LLM inference cost"],"falsifier":"Re-run the six production benchmarks with independent parallel workers and no global proxy pass; if either algorithm fails to reach F1 of 0.95 at its claimed operating points, or if GAMCAL does not cut oracle calls by roughly the reported margin versus the prior SUPG cascade at comparable quality, the central performance and guarantee claims fail.","tokens_in":12416,"feed_emoji":"⚡","tokens_out":1046,"duration_ms":15958,"temperature":0.7,"pith_summary":"Semantic SQL operators call large language models on every qualifying row, which is far costlier than ordinary SQL. Model cascades try to fix this by sending most rows through a cheap proxy and only the uncertain ones to a costly oracle, but earlier SUPG-style designs need a full global proxy-score pass that itself is an LLM workload and stalls pipelined engines; they also optimize only precision or only recall. This paper formalizes cascade routing for streaming engines with independent parallel workers and gives two complementary algorithms. SUPG-IT iteratively refines two thresholds from accumulating batch oracle labels and is the first streaming cascade with joint probabilistic guarantees on user-specified precision and recall at a chosen failure probability. GAMCAL replaces fixed targets with a single error-versus-cost tradeoff, fits a monotone generalized additive model that turns proxy scores into true-positive probabilities plus pointwise uncertainty, and routes stochastically. On six classification, filtering and join benchmarks inside a production semantic SQL engine both methods reach F1 at least 0.95; GAMCAL leads at a 20 percent delegation budget and needs up to 58 percent fewer oracle calls than the prior SUPG cascade, while SUPG-IT posts the highest best-case F1 (mean 0.989).","feed_headline":"Streaming cascades cut LLM calls in semantic SQL","feed_subtitle":"Two algorithms hit F1 above 0.95 with joint guarantees and up to 58 percent fewer oracle calls","key_machinery":"Cascade routing under independent parallel streaming workers: SUPG-IT's iterative dual-threshold refinement from accumulating batch oracle labels, and GAMCAL's monotone Generalized Additive Model that maps proxy scores to calibrated true-positive probabilities plus pointwise uncertainty for stochastic routing.","core_discovery":"Streaming semantic SQL can run model cascades without a blocking global proxy-score pass. SUPG-IT extends single-pass SUPG to iterative dual-threshold refinement across batches and supplies the first joint probabilistic guarantees on user-chosen precision and recall at failure probability delta. GAMCAL learns a monotone GAM that calibrates proxy scores to true-positive probabilities with pointwise uncertainty for stochastic routing under a single tradeoff parameter alpha. Both attain F1 greater than or equal to 0.95 on six production benchmarks, with GAMCAL using substantially fewer oracle calls.","pith_inferences":["If batch-local statistics remain representative under highly skewed or non-stationary streams, the same iterative and GAM machinery could extend to continuous query settings beyond the six static benchmarks.","The joint precision-recall guarantee of SUPG-IT suggests a natural path to multi-objective cascades that also bound latency or monetary cost under the same delta.","Because GAMCAL supplies pointwise uncertainty, it could be combined with active-learning style oracle selection to further reduce labels when the proxy is already well-calibrated.","The formalization for independent workers implies that cascade logic can be pushed into the worker itself rather than a central coordinator, which would matter for massively parallel warehouse deployments."],"forward_implications":["Pipelined semantic SQL engines can emit results without waiting for a full proxy-score materialization.","Workloads that need both high precision and high recall can be served by one cascade rather than separate precision-only or recall-only designs.","A single alpha knob lets operators trade classification error against oracle cost without hand-tuning two thresholds.","Up to roughly half the oracle LLM calls can be avoided while still reaching F1 of 0.95 on the evaluated classification, filtering and join tasks.","The same streaming cascade model can be reused for any semantic operator that exposes a cheap proxy score and an expensive oracle label."],"fun_headline_variants":["Streaming cascades cut oracle calls without global proxy pass","SUPG-IT gives joint precision-recall guarantees in streaming SQL","GAMCAL calibrates proxy scores for lower-cost cascade routing","Both cascades reach F1 of 0.95 with far fewer LLM oracle calls","Streaming dual-threshold cascades serve precision and recall together"],"cache_read_input_tokens":128,"weakest_assumption_plain":"The claim rests on the premise that thresholds refined from successive batch oracle labels, and a monotone GAM fitted to those labels, transfer well enough across independent parallel workers and the whole stream that the joint precision-recall guarantees and reported F1 and oracle savings still hold without any global score pass.","fun_headline_variants_meta":{"raw":{"variants":["Streaming cascades cut oracle calls without global proxy pass","SUPG-IT gives joint precision-recall guarantees in streaming SQL","GAMCAL calibrates proxy scores for lower-cost cascade routing","Both cascades reach F1 of 0.95 with far fewer LLM oracle calls","Streaming dual-threshold cascades serve precision and recall together"]},"model":"grok-4.5","effort":"low","cost_usd":0.004484,"raw_usage":{"total_tokens":1408,"prompt_tokens":895,"num_sources_used":0,"completion_tokens":73,"cost_in_usd_ticks":44840000,"prompt_tokens_details":{"text_tokens":895,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":440,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":895,"tokens_out":73,"duration_ms":4118,"temperature":1.0,"reasoning_tokens":440,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T14:55:24.429518+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Re-run the six production benchmarks with independent parallel workers and no global proxy pass; if either algorithm fails to reach F1 of 0.95 at its claimed operating points, or if GAMCAL does not cut oracle calls by roughly the reported margin versus the prior SUPG cascade at comparable quality, the central performance and guarantee claims fail.","supporting_citations":[],"review_version":1}