{"id":"2337e19a-cac6-45fd-90d8-6d2f52c1214c","arxiv_id":"2607.27453","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Eleven production voice agents complete 43–71% of 100 stateful card-support phone calls when graded jointly on speech and tool-execution traces.","lead":"VAmoS Bench tests complete voice agents on 100 simulated bank card-support phone calls, jointly grading what the agent says and what it does in a live SQL backend. Teams picking among cascaded, hosted, and speech-to-speech stacks get a containment-style score instead of only WER or latency.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"Unvalidated single-LLM judge on synthetic traces is still the load-bearing risk for the reported completion rates and ordering.","rationale":"The paper’s strongest contribution is the multi-transport, SQL-backed, jointly graded phone-call protocol plus an 11-stack snapshot on a fixed 100-scenario set. The authors already frame results as a descriptive snapshot rather than a resolved ranking and flag judge/synthetic-user limits in §6. The reader correctly identified the unvalidated LLM judge (no gold set, synthetic caller, same platform generating and grading) as the weakest assumption underwriting cross-system completion numbers. No stronger internal inconsistency or hidden mathematical flaw appears; proprietary harness and three-run variance are secondary reproducibility/precision issues, not replacements for the judge-validity concern. A human-agreement audit is the single check that would settle whether that concern lands. Until then the CONDITIONAL verdict (publishable/useful if materials are released or audited and claims stay scoped) remains appropriate; no upgrade or downgrade is warranted from this pass.","tokens_in":13013,"tokens_out":558,"duration_ms":22802,"concrete_test":"Stratified sample of ~200 graded traces (balanced across the 11 agents and the simple/complex/adversarial partitions). Blind human annotators apply the identical per-scenario assertions to the same joint records; compute per-assertion and per-call agreement (Cohen’s κ) with the LLM judge and test for architecture-class bias. If κ < 0.8 or completion deltas reverse for any top-tier pair, the snapshot rates/ordering are not reliable enough for the strongest claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim is that joint assertion grading of conversation + tool trace yields comparable task-completion (containment) rates across eleven stacks (43–71%, complex group weakest at 52.2%). That claim rests on Section 3.5’s grader: one LLM applying fixed natural-language binary assertions to the combined record, with success only if all assertions pass. Limitations §6 states explicitly there is no human-annotated gold set, the caller is fully synthetic, and the authors’ Veris platform both generates scenarios and runs the harness/grader. Without calibration, systematic judge bias (favoring certain transcript styles, tool-log verbosity, or refusal phrasings) or synthetic-user artifacts could shift absolute rates and relative order, especially given only three runs and run-level SEs up to 5.5 pp. Joint-trace design is methodologically stronger than final-state-only rewards, but the unvalidated judge is the least secure condition for treating the numbers as true cross-system containment.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces VAmoS Bench, an end-to-end benchmark for complete voice-agent systems on a stateful card-operations task. Riley, a fixed-policy credit-card agent with five real-SQL tools against isolated per-call PostgreSQL backends, is exercised over live audio on 100 scenarios (30 simple / 38 complex / 32 adversarial after re-partition), each with a simulated caller holding a private goal and fixed binary assertions. A single LLM judge scores the joint record of transcript plus tool invocations, arguments, and returned rows; success requires all assertions to pass. Eleven stacks (self-hosted frameworks, hosted platforms, bundled APIs, and native speech-to-speech endpoints) are each run three times (3,300 calls). Observed task completion ranges from 43.0% (Nemotron) to 71.0% (Pipecat); the complex group is weakest for every agent (pooled 52.2%). The authors emphasize joint say/do grading, transport-flexible adapters, and an extensible leaderboard protocol rather than a definitive product ranking.","tokens_in":13308,"tokens_out":1616,"duration_ms":46762,"significance":"If the protocol and measurements hold, this is a useful systems contribution: public voice benchmarks still under-measure containment-style end-to-end correctness on stateful tool-using phone calls. Strengths that should be credited explicitly include (i) real SQL-backed tools and per-call backend isolation rather than declarative mocks, (ii) joint assertions over conversation and execution trace, which can catch say/do mismatches and unauthorized disclosure that final-state-only rewards miss, (iii) controlled near-matched cascade comparisons (Pipecat vs LiveKit; baseline vs all-NVIDIA under the same harness), (iv) counting unconnected/ungraded calls as misses in the denominator, and (v) transparent reporting of run-level SE, connect failures, and operational metrics alongside completion. The design is closer to deployment value than WER/naturalness-only suites and is a natural audio extension of the τ-bench pattern. The main significance risk is that absolute rates and fine ordering rest on an unvalidated judge and a single synthetic domain.","major_comments":[{"comment":"Section 3.5 and Limitations §6: the central empirical claim (Table 1 completion 43–71%; Table 2 complex group weakest at 52.2%) is produced by a single LLM applying natural-language binary assertions to synthetic-caller traces, with no human-annotated gold set and no reported human–LLM or inter-annotator agreement. The joint-trace design is methodologically stronger than final-state-only rewards, but without a calibrated subsample (e.g., double-annotated stratified traces covering say/do, confidentiality, and sequence assertions) the absolute rates and near-top ordering cannot be treated as reliable cross-system containment. This is load-bearing for the results section; a modest human validation study or release of graded traces with agreement statistics would address it.","section":"Section 3.5; §6"},{"comment":"Section 4.1 and Table 1: one-SE bars across three run-level rates reach 5.5 pp, and §6 reports per-run completion spans up to 17 pp (Nemotron), 12 (Cartesia), and 8 (Pipecat)—larger than several gaps near the top. The text correctly calls the row order descriptive, yet the table is still ordered by observed completion with bolded “best” cells, and the abstract/conclusion lead with the 43–71% range. Either add enough runs to resolve neighboring differences, report pairwise uncertainty more formally, or de-emphasize ordered ranking (e.g., unsorted or banded presentation) so the snapshot claim matches the statistics.","section":"Section 4.1; Table 1; §6"},{"comment":"Sections 3.1–3.3 and Acknowledgments: scenario generation, simulation harness, and grading all run on the authors’ Veris platform, and the same team reviewed the 100 scenarios. Isolation and fixed assertions reduce some risks, but platform-specific transport adapters, logging verbosity, and assertion wording could still interact with stack behavior. The manuscript should state more sharply what is held fixed vs. platform-mediated, and ideally provide an external reproducibility path (scenario pack, assertion text, trace schema, and a non-Veris grading reference) so third parties can re-score without the authors’ stack.","section":"§3.1–3.3; Acknowledgments"},{"comment":"Section 4.3 and Failure analysis §5: confidentiality failures (≈33–39% pass on repeated failed verification and field-probe scenarios) and policy-order collapse under freeze-before-replace pressure (3.0% and 39.4% on the two cited branches) are among the most actionable findings, but they are pooled end-to-end rates that mix connect misses, judge decisions, and agent policy errors. Separating (a) connect/setup failures, (b) assertion failures with tool evidence, and (c) pure conversational disclosure failures—and releasing per-scenario pass counts—would make the “refusal vs confidentiality” claim falsifiable and comparable across future leaderboard entries.","section":"Section 4.3; §5"}],"minor_comments":[{"comment":"Figure 3: Pareto ovals and hollow/filled dots are useful but dense; a short caption note defining “Pareto-optimal” on this two-metric slice (and that Nemotron is latency-only) would help.","section":"Figure 3"},{"comment":"Appendix A / Table 3: pin exact model snapshot IDs and dates more uniformly (some rows already do); preview models (Gemini, Realtime) will drift quickly for an evolving leaderboard.","section":"Appendix A"},{"comment":"Section 3.4: the Nova-3 vs nova-3-general identifier split for Vapi vs Pipecat/LiveKit is acknowledged; state explicitly in the table footnote that this is not a pure harness isolation.","section":"Section 3.4"},{"comment":"Typos/orthography: title and headers alternate “V AmoS” / “VAmoS”; “Acascadetypically” and “Aspeech-to-speechsystem” need spaces (p.1); “eitherthe” (p.4).","section":"Throughout"},{"comment":"Cost column: §6’s exclusions (self-hosted servers, annual licenses, telephony) are appropriate; add a one-line footnote under Table 1 pointing to that paragraph so the bold “lowest cost” cells are not over-read.","section":"Table 1; §6"}],"recommendation":"major_revision","confidential_remarks":"Mild COI/fit note for the editor: all authors are affiliated with Veris AI, which builds and operates the simulation platform used to generate scenarios, run calls, and grade traces. The benchmark contribution looks real and the limitations section is candid, but the empirical leaderboard should not be read as an independent third-party bake-off until judge calibration and an external re-scoring path exist. Scope is appropriate for a systems/benchmark track; I would not reject on affiliation alone."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing worth knowing is that this is a practical containment bench, not another WER/latency bake-off. They run full phone calls against isolated Postgres, grade conversation and tool traces together with fixed binary assertions, and report 3,300 calls across eleven named stacks (self-hosted, hosted, bundled, native audio). Observed completion sits 43–71%, with complex multi-step flows weakest for everyone (~52%). That joint say/do check is the real methodological win over final-state-only rewards and over EVA-Bench’s decoupled metrics.\n\nWhat they do well: methods are explicit (isolation, connect failures as misses, three full repeats, SE on run-level rates, mutually exclusive simple/complex/adversarial split). Controlled comparisons are honest—Pipecat vs LiveKit under a pinned cascade is nearly a wash; swapping the whole stack for Nemotron drops hard and they separate connect failures from answered-call gaps. Failure analysis is concrete: policy-order collapse under pressure, confidentiality leaks after correct refusals, and at least 27 explicit say/do mismatches. They mostly frame results as a fixed-set snapshot rather than a product ranking, which matches the error bars.\n\nSoft spots in proportion: the load-bearing risk is the single unvalidated LLM judge on synthetic-caller traces with no human gold set, on a harness the authors operate. That can bias absolute rates and maybe order; they admit it. Single banking vertical, proprietary platform (scenario defs and grader materials matter for trust), and cost figures are rough snapshots. None of that sinks the design—joint trace grading is still stronger than transcript-only or DB-hash-only—but it caps how hard you should lean on the leaderboard numbers until assertions and calibration are public or audited.\n\nCitations sit cleanly on τ-bench, EVA-Bench, τ-Voice; novelty is breadth of deployment products plus path-sensitive joint assertions, not a new theory. Math is not the point; the empirical protocol is.\n\nWho it’s for: people building or buying voice agents, and anyone writing the next agent bench. I’d bring it to a systems/agents reading group. It deserves serious referee time—publishable with pressure to release scenarios/assertions and tighten judge validation—not a desk reject. Engage if you care about end-to-end voice agents; skim if you only want component speech metrics.","headline":"Solid systems benchmark for voice-agent containment: joint speech+tool grading on 11 real stacks is useful, even if the unvalidated LLM judge and single-domain setup keep the numbers as a snapshot, not a ranking.","tokens_in":13968,"tokens_out":596,"would_cite":true,"duration_ms":16099,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Voice agents must be scored on what they say and do together, not on components or final database state alone.","keywords":["voice agents","end-to-end evaluation","containment","tool-using agents","joint speech-execution grading","customer support simulation","stateful benchmarks","adversarial callers"],"falsifier":"Build a human-annotated gold set on the same 100 traces and show that the LLM judge’s pass/fail labels disagree with humans enough to reorder the top stacks or erase the complex-versus-adversarial gap; or replace the synthetic caller with real customers and watch completion and relative ranking move outside the reported run-to-run error bars.","tokens_in":13865,"feed_emoji":"📞","tokens_out":938,"duration_ms":22956,"temperature":0.7,"pith_summary":"Most voice-agent benchmarks score pieces of the stack—transcription error, latency, naturalness, turn-taking—or final database state. Contact centers care about containment: whether the automated system resolved the call correctly on its own, including when the right answer is refusal or redirect. VAmoS Bench puts complete production stacks through the same 100 live audio phone calls on a stateful bank card-operations task. Each call runs against an isolated seeded database with real SQL tools, a simulated caller holding a private goal, and fixed binary assertions graded on the joint trace of speech plus tool invocations, arguments, and returned rows. That joint record catches agents that claim success without acting, act while leaking protected data, or skip required policy steps. On eleven current stacks, observed completion ranges from 43% to 71%, with multi-step complex flows the weakest group for every agent.","feed_headline":"Voice agents top out at 71% on real card-support calls","feed_subtitle":"Joint speech-and-tool grading shows multi-step flows fail most, across eleven stacks","key_machinery":"Joint grading of speech and execution: fixed per-scenario binary assertions evaluated by a judge against one evidence record that includes the transcript, tool invocations, arguments, returned rows, and their order. A single assertion can require that spoken claims match system state and that no state change precede verification.","core_discovery":"End-to-end voice-agent quality is not captured by component scores or final state alone. On a fixed 100-scenario card-operations task with real audio, real SQL tools, and per-scenario assertions applied to conversation and execution together, eleven production stacks complete between 43.0% and 71.0% of calls. Complex multi-step flows are the hardest group for every agent (pooled 52.2%), while direct adversarial bypasses are handled better than confidentiality under failed verification. Joint grading surfaces say/do mismatches and policy-order failures that either channel alone would miss.","pith_inferences":["If say/do mismatches remain common even on strong stacks, production monitoring will need continuous joint-trace audits, not transcript review or DB snapshots alone.","The large drop when swapping an entire model stack inside one harness suggests component choice still dominates orchestration for this task class.","Adversarial callers who fish for which verification field failed may be a more realistic production failure mode than classic prompt injection.","Extending the same joint-assertion design to healthcare intake or insurance claims would test whether the complex-flow weakness is domain-general."],"forward_implications":["Teams choosing among self-hosted frameworks, hosted platforms, and native-audio APIs can compare them on the same phone-call protocol and tool interfaces rather than on WER or latency alone.","Containment scores must treat correct refusal and redirect as success, not only fulfilled requests.","Path-sensitive checks (verify before disclose; freeze before replace) become first-class success criteria once the judge reads tools and speech together.","Later benchmark versions can keep the same protocol while expanding domains and scenarios, supporting an evolving leaderboard on fixed versions.","Confidentiality under failed verification deserves separate scoring from direct bypass resistance."],"fun_headline_variants":["Voice agents top out at 71% on real card-support calls","Eleven stacks clear only 43–71% of end-to-end card calls","Multi-step card flows hardest: pooled 52% across agents","Joint speech-and-tool grading exposes say/do mismatches","VAmoS: real audio and SQL put voice containment at 71% max"],"cache_read_input_tokens":128,"weakest_assumption_plain":"An unvalidated single language-model judge scoring synthetic callers on natural-language assertions is reliable enough to compare live voice-agent stacks.","fun_headline_variants_meta":{"raw":{"variants":["Voice agents top out at 71% on real card-support calls","Eleven stacks clear only 43–71% of end-to-end card calls","Multi-step card flows hardest: pooled 52% across agents","Joint speech-and-tool grading exposes say/do mismatches","VAmoS: real audio and SQL put voice containment at 71% max"]},"model":"grok-4.5","effort":"low","cost_usd":0.004115,"raw_usage":{"total_tokens":1359,"prompt_tokens":901,"num_sources_used":0,"completion_tokens":82,"cost_in_usd_ticks":41148000,"prompt_tokens_details":{"text_tokens":901,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":376,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":901,"tokens_out":82,"duration_ms":6398,"temperature":1.0,"reasoning_tokens":376,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T00:24:53.992835+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Build a human-annotated gold set on the same 100 traces and show that the LLM judge’s pass/fail labels disagree with humans enough to reorder the top stacks or erase the complex-versus-adversarial gap; or replace the synthetic caller with real customers and watch completion and relative ranking move outside the reported run-to-run error bars.","supporting_citations":[],"review_version":1}