{"id":"733fe751-dd49-4ab7-a585-a79704090a42","arxiv_id":"2508.16318","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":2,"one_line_summary":"SATORI statically infers REST API test oracles from OpenAPI specs via LLMs, reporting F1 74.3%, above AGORA+'s 69.3%, with 18 confirmed bugs; the supplied full text, however, is a different paper.","lead":"This paper describes SATORI, a tool that uses large language models to infer test oracles for REST APIs from OpenAPI specifications without executing the API, reporting higher F1 than a state-of-the-art dynamic approach on 17 operations. A generalist should care because cheaper oracle generation could make automated API testing more practical; note the review basis is the abstract only, as the supplied full text is a different paper.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground-truth oracles may have been derived from the same OpenAPI text SATORI reads, so F1 could measure spec-text agreement, not behavioral validity; the mismatched full text leaves the annotation protocol uncheckable.","rationale":"The central claim is that SATORI can infer correct REST API oracles statically from OpenAPI field names and descriptions, beating the dynamic AGORA+ and complementing it (90% joint coverage). For this to be true, two things must hold: (1) the spec text carries enough semantic signal, and (2) the 'ground-truth' oracles used to measure F1 and coverage are correct expected behaviors, not restatements of the same text the LLM reads. I focus on (2) because it is the more direct correctness threat. If (2) fails, even a perfect LLM would appear to do well: it would simply mimic the annotation source, and the comparison to AGORA+ would not measure real-world usefulness. The abstract gives no annotation protocol, and the supplied full text is a different manuscript (ConceptGuard), so no protocol can be inspected. The 18 bugs provide some external validation, but they are few and the abstract does not say they were found blindly or that they cover the oracle counts; documentation updates can also result from ambiguity in the maintainer's own specs rather than an API bug. A concrete check — blind re-annotation from executed behavior for a few operations — would settle whether F1 is inflated. This concern does not change the reader's verdict: the paper is already UNVERDICTED because the actual text is unavailable. If the full text reveals independent annotation, the central claim would deserve a proper correctness review; if not, the empirical headline is likely circular. The reader's weakest_assumption identified spec-text sufficiency as the primary issue and ground-truth independence as a second structural premise; I agree with the latter but make it primary, so agreement is partial.","tokens_in":941,"tokens_out":2391,"duration_ms":67865,"concrete_test":"In the full manuscript, locate the ground-truth annotation protocol and answer: were annotators shown the OpenAPI response-field names/descriptions, or independent behavioral sources (API execution traces, external documentation)? Then have a fresh annotator, blind to the OpenAPI text, label valid oracles for a random sample of 3-5 of the 17 operations using only executed responses and independent documentation; recompute SATORI's F1 against those labels. If F1 drops by more than ~5 absolute points, or if original-versus-fresh label agreement is below ~80%, circular inflation is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"SATORI's abstract claims it infers valid test oracles from OpenAPI response-field names and descriptions, achieving F1=74.3% on 17 operations versus AGORA+'s 69.3%. The metric is computed against an 'annotated ground-truth dataset.' The load-bearing premise is that those ground-truth oracles were created independently of the spec signal SATORI consumes. If annotators wrote oracle labels after reading the same field names and descriptions, the F1 is an agreement-with-source-text measure, not a measure of correct expected API behavior. OpenAPI descriptions rarely specify valid value sets, ranges, or cross-field invariants; they say 'status' or 'total cost.' An annotator who has only that text can only re-encode the shallow semantics the LLM already sees. The 90% joint coverage with AGORA+ and the 18-bug claim are downstream of this: bugs are only reported for a few APIs and the abstract does not say they were found blindly or that they cover the other oracle counts. The supplied manuscript text is actually ConceptGuard, a different paper, so the annotation protocol, oracle definition, and per-operation results are absent. Thus the core claim is unverified, not refuted.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The submission consists of an abstract for a paper titled 'SATORI: Static Test Oracle Generation for REST APIs' followed by the full text of an unrelated paper titled 'ConceptGuard: Neuro-Symbolic Safety Guardrails via Sparse Interpretable Jailbreak Concepts' (arXiv:2508.16325). The abstract claims that SATORI, a black-box static approach, uses LLMs to infer test oracles from OpenAPI specifications, achieving F1=74.3% on 17 operations from 12 industrial APIs (outperforming AGORA+'s 69.3%), with complementary joint coverage of 90% and 18 bugs found in popular APIs. The full text, however, describes a mechanistic-interpretability framework for LLM safety guardrails and contains no methodology, evaluation, or results for SATORI. The claimed SATORI approach is therefore entirely absent from the manuscript body.","tokens_in":16735,"tokens_out":2496,"duration_ms":27934,"significance":"If SATORI's claims were supported, the work could be significant for black-box REST API testing: generating valid test oracles from static specifications without executing the API would be a useful advance over dynamic approaches like AGORA+. The potential complementary coverage of static and dynamic inference is also an interesting hypothesis. However, as submitted, the manuscript provides no evidence for these claims: the full text is a different paper on a different topic. The abstract's quantitative results cannot be evaluated or reproduced. The significance of the claimed contribution is real but entirely unbacked in this document.","major_comments":[{"comment":"The manuscript body is not the paper described in the title/abstract. Sections 1–6 and the Appendix present ConceptGuard, an LLM safety-guardrail framework using sparse autoencoders, with no mention of REST APIs, OpenAPI specifications, test oracles, or SATORI. The central claim of the abstract—that SATORI generates valid oracles with F1=74.3%—is therefore unsupported by any methodological description, experimental setup, or results. This is a load-bearing omission that cannot be addressed by minor revision.","section":"Full text (all sections)"},{"comment":"The abstract reports F1=74.3% versus AGORA+'s 69.3%, 90% joint coverage, and 18 bugs, but the manuscript provides no evaluation section, no definition of oracle validity, no ground-truth annotation protocol, no per-operation breakdown, and no statistical significance or variance information. Even if the correct full text were supplied, these abstract-level numbers alone would be insufficient to support the generalization claims made.","section":"Abstract (results claims)"},{"comment":"The provided full text itself contains an explicit limitation statement (Section 6: 'our analysis presents a proof-of-concept, limited to a single hook point') and acknowledges heuristic pruning and potential over-aggressiveness. These are appropriate caveats for ConceptGuard, but they are irrelevant to SATORI. Their presence underscores that the submitted document is not the claimed paper and that no SATORI-specific limitations or validation are discussed.","section":"Full text (ConceptGuard limitations)"}],"minor_comments":[{"comment":"The title and abstract identify the paper as SATORI, while the body is ConceptGuard. This mismatch should be resolved before any resubmission; it is a fundamental presentation issue.","section":"Title/Abstract"}],"recommendation":"reject","confidential_remarks":"This appears to be a submission assembly error: the abstract is for SATORI but the full text is ConceptGuard. Under the reviewing rule to treat the provided text as the manuscript, the submission cannot be reviewed as a scientific paper on SATORI. If the authors intended to submit SATORI, a resubmission with the correct full text would be necessary. As it stands, the central claims are unsupported by any body text, so rejection is the only defensible recommendation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The abstract describes SATORI, a static approach to test-oracle generation for REST APIs using LLMs on OpenAPI specs, claiming F1 of 74.3% versus AGORA+’s 69.3% and 18 maintainer-confirmed bugs. That is a genuinely useful idea: oracle generation is often the bottleneck, and doing it without deploying the API is attractive. The comparison to dynamic SOTA and the complementarity claim are the new parts, and the bug reports are real evidence if they hold up.\n\nNow the trouble. The supplied full text is a different paper entirely — ConceptGuard, about jailbreak guardrails. I cannot inspect the annotation protocol, the definition of a “valid oracle,” the per-API breakdown, or the LLM configuration. My judgment rests on the abstract alone, so everything here is provisional.\n\nThe main scientific soft spot is the ground-truth annotation. Oracles are inferred from response-field names and descriptions, and F1 is computed against an annotated dataset. If the annotators saw the same OpenAPI text, the metric partly measures agreement with the spec, not behavioral correctness. OpenAPI descriptions are often too shallow to distinguish valid from invalid oracles. The paper needs to show the annotation was independent of the signal the LLM consumes — for example, deriving ground truth from existing test suites, bug reports, or maintainer input. Without that, the 74.3% could be inflated. The 18 bugs mitigate this for a subset of APIs, but the abstract doesn’t say whether they were found blindly or how many oracles they cover.\n\nThe sample size (17 operations, 12 APIs) is small for a generalization claim, and the F1 gap is reported without variance or significance testing. That’s fixable, not fatal; I’d want per-API results and at least a paired comparison.\n\nIf the full paper delivers on the abstract and the annotation is independent, this is a solid SE contribution worth citing. If not, it’s another LLM benchmark result. My recommendation: take it seriously, send it to peer review, and make the dataset and annotation protocol a condition. The mismatch means I cannot endorse it as-is, but the direction is worth a careful look.","headline":"Plausible static-oracle approach, but the full-text mismatch and unverified annotation protocol mean the paper needs referee scrutiny, not blind acceptance.","tokens_in":17352,"tokens_out":1618,"would_cite":true,"duration_ms":18591,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that REST API test oracles can be generated statically from an OpenAPI specification alone, using an LLM to infer expected response behavior from field names and descriptions, and that this beats dynamic oracle generation","keywords":["REST API testing","test oracle generation","OpenAPI Specification","large language models","static analysis","black-box testing","API testing","behavioral assertions"],"falsifier":"Strip property descriptions from a set of OpenAPI specs and rerun SATORI: if F1 stays near the reported level, descriptive semantics are not the active ingredient. Separately, build ground truth by observing actual API responses rather than reading the spec, and compare; if agreement is much lower than 74.3%, the score measured spec-consistency rather than behavioral validity.","tokens_in":16366,"feed_emoji":"🧪","tokens_out":5549,"duration_ms":61418,"temperature":0.7,"pith_summary":"REST API test generators are good at producing input data but weak at knowing what a correct response looks like: their oracles usually stop at crashes, status-code errors, or schema noncompliance. SATORI tries to close that gap without running the API. It treats the OpenAPI specification as the source of truth and asks a large language model to infer, from each response field's name and description, what the API should return. On 17 operations from 12 industrial APIs, the paper reports an F1 of 74.3%, above the 69.3% of a dynamic state-of-the-art approach on the same oracle types, with the two methods together covering 90% of an annotated ground-truth set. The paper also reports 18 real bugs found in widely used APIs, which maintainers fixed by updating documentation.","feed_headline":"OpenAPI specs alone can yield behavioral test oracles","feed_subtitle":"SATORI’s LLM-based static inference beats dynamic oracle generation and found 18 real API bugs.","key_machinery":"The OpenAPI Specification is the central object. SATORI reads one operation at a time, takes the names and descriptions of its response fields, and asks an LLM to infer expected properties—field presence, types, ranges, formats, and semantic invariants. The resulting oracles are converted into executable assertions by an extension of an existing Postman assertion tool, which is what lets the static inferences be run as real tests.","core_discovery":"On the paper's terms, the discovery is that the semantic metadata already present in an OpenAPI specification—response field names and their natural-language descriptions—carries enough signal for an LLM to infer what a correct response should look like, so behavioral oracles can be produced without executing the API. The paper supports this with 17 operations from 12 industrial APIs: SATORI reached 74.3% F1, above the 69.3% of a dynamic state-of-the-art approach on the same oracle types, and the two approaches together recovered 90% of an annotated ground-truth set. It also reports that the generated oracles uncovered 18 bugs in widely used public APIs, which maintainers addressed with docu","pith_inferences":["The reported F1 may partly reward agreement with the specification text itself, since the ground-truth oracles and the LLM read the same descriptions; a validation set built from observed API responses would separate behavioral correctness from spec-fidelity.","Specifications with sparse or generic descriptions are likely the failure mode; supplementing the prompt with endpoint context, examples, or request schemas is a plausible extension the paper does not explore.","The same static-inference chain should transfer to other interface contracts such as GraphQL schemas or gRPC protos, and to request-side constraints where no execution is needed either.","Because the generated oracles are executable and static, they could be rerun continuously to detect drift between a published specification and a live API."],"forward_implications":["Behavioral oracles become available before an API is deployed, because generation needs only the specification file, not a running system.","Static and dynamic oracle inference are complementary: combining them covered 90% of the annotated ground-truth oracles, pointing toward hybrid test pipelines.","REST API test suites can be enriched from status-code checks to hundreds of field-level behavioral checks per operation.","The 18 documentation bugs found in widely used APIs show that specification-driven oracles can surface real discrepancies between documented and actual behavior."],"supporting_citations":[],"fun_headline_variants":["Static oracle generation outperforms dynamic with LLMs","LLMs turn API specs into test oracles without execution","SATORI: Behavioral oracles from static API specs alone","API specs enough for test oracles, 18 bugs found"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The whole approach rests on OpenAPI response descriptions being informative enough that an LLM can infer correct expected behavior; if real-world specs are vague, missing, or misleading, the oracle quality drops.","fun_headline_variants_meta":{"raw":{"variants":["Static oracle generation outperforms dynamic with LLMs","LLMs turn API specs into test oracles without execution","SATORI: Behavioral oracles from static API specs alone","API specs enough for test oracles, 18 bugs found"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000298,"raw_usage":{"total_tokens":1602,"prompt_tokens":825,"completion_tokens":777,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":569,"completion_tokens_details":{"reasoning_tokens":719}},"tokens_in":569,"tokens_out":777,"duration_ms":8011,"temperature":1.0,"reasoning_tokens":719,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T17:22:46.441527+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Strip property descriptions from a set of OpenAPI specs and rerun SATORI: if F1 stays near the reported level, descriptive semantics are not the active ingredient. Separately, build ground truth by observing actual API responses rather than reading the spec, and compare; if agreement is much lower than 74.3%, the score measured spec-consistency rather than behavioral validity.","supporting_citations":[],"review_version":1}