{"id":"9323b1ba-ce0d-4350-a806-0bc1fb1bd692","arxiv_id":"2607.03926","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":7.0,"correctness_risk":"low","formal_verification":"none","parameter_count":5,"one_line_summary":"Across 49 datasets and 11 generators, distance-based fidelity overstates synthetic tabular quality: best query-centric score is only 0.75, with systematic failures on high-cardinality support, local conditionals, and extreme tails.","lead":"TabQueryBench scores synthetic tables by whether they return the same answers as real data on grounded analytical SQL queries. Leading generators still lag real data on this measure and fail on high-cardinality columns, local filters, and rare tails.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified","rationale":"The paper's central claim is empirical and scoped: under the released single-table query-centric assessors, current models that look strong on distance metrics still under-perform on analytical query answers, with consistent failure modes and a clear cost–fidelity tradeoff. That claim is supported by open artifacts, a large suite (49 datasets, 11 models, >100 queries/dataset), family-level breakdowns, and an explicit stability study on the only non-deterministic step. The reader's weakest assumption correctly flags instrument design as the residual risk, but the paper already mitigates it with public provenance, fixed policies, validators, and ranking-stability numbers; it does not silently assume the 44 templates are exhaustive of all SQL. Multi-table joins, privacy-primary evaluation, and alternative score aggregations are out of scope by design (Sections 3.2, 6) and do not falsify the reported single-table patterns. No internal contradiction or load-bearing hidden assumption that would reverse the five patterns was found. Verdict remains ACCEPT; no adjustment needed.","tokens_in":31519,"tokens_out":570,"duration_ms":5592,"concrete_test":"Re-run the full 49\times11 evaluation after replacing the LLM realizer with a purely rule-based SQL expander that only fills the already-bound slots from the fixed template skeletons (no free generation); if RealTabFormer remains top on query overall, the local-vs-global conditional gap and high-cardinality/tail failures persist at similar magnitude, and BayesNet stays on the Pareto frontier, the instrument-dependence concern does not land.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption (template representativeness and non-cherry-picked grounding) is the right residual risk for any query-centric instrument, but the paper already bounds it tightly enough that it does not undermine the central claim. Stage 1 templates are fixed, public-source-attributed, and deduplicated (Appendix A, Table 7); Stage 2 uses deterministic profiling/binding/validation with LLM only inside constrained realization; Section 5.5 reports high ranking stability (mean Kendall W 0.927, mean pairwise Spearman 0.903, 83.7% of cells move by at most one rank) across three regenerations on a 9-dataset probe. The five patterns (distance–query mismatch, high-cardinality collapse, local vs global conditional drop, tail degradation, BayesNet cost–fidelity Pareto) are multi-dataset, multi-model, and family-level rather than single-template artifacts. Within the paper's explicit single-table mimicry scope, the instrument is sufficiently stable and transparent that the strongest claim holds.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"TabQueryBench proposes a query-centric evaluation framework for synthetic tabular data: reusable SQL-shaped analytical queries act as structural assessors of fidelity rather than relying only on distance-based, privacy, or ML-utility metrics. From 12 public analytical-query sources the authors distill 44 templates in five families (subgroup, conditional, tail/rarity, missingness, cardinality/range), ground them to each of 49 datasets via a policy-guided template-to-SQL pipeline (deterministic profiling/binding/validation; LLM only in constrained realization), and evaluate 11 generative models. The main empirical claims are that (i) strong distance-based scores can coexist with substantially lower query-centric fidelity (best model RealTabFormer at 0.75±0.15 vs REAL=1.00), (ii) high-cardinality discrete support often collapses, (iii) local conditional slices are harder than global counterparts, (iv) tail fidelity degrades under stricter rarity thresholds, and (v) BayesNet offers the best fidelity–cost tradeoff on a common-9 runtime slice. A three-run ranking-stability study and open release of code, templates, and artifacts support the instrument.","tokens_in":31837,"tokens_out":1187,"duration_ms":16808,"significance":"The paper addresses a genuine and practically important gap: synthetic tabular data are often used for analytics, system testing, and data sharing, yet existing benchmarks rarely treat analytical query answers as first-class evaluation objects. The contribution is not a single theorem but a carefully constructed, extensible instrument with public provenance for templates, large multi-dataset/multi-model coverage, family-level diagnostics, a cost Pareto view, and an explicit stability audit of LLM-assisted grounding. If the reported patterns hold under the stated single-table mimicry scope—and the evidence is multi-faceted enough that they appear to—the work should influence both model selection practice and future generative-model design (e.g., high-cardinality support, rare-region, and local-slice objectives). Open code and artifacts further raise the work’s value as a community foundation.","major_comments":[{"comment":"The central numerical claims (e.g., RealTabFormer 0.75±0.15 query overall; local-slice drop of 0.11; ~40.7% rare-value recovery) depend on a precise definition of how a synthetic query answer is scored against the real answer. The main text and Appendix describe families, templates, and aggregate tables (Table 8, Figures 4–9) but do not give an explicit, auditable scoring map from (real result, synthetic result) to [0,1] per template type (counts, rates, rankings, support sets, range envelopes, missing rates). Please add a short formal definition (or algorithm box) covering result alignment, empty-support cases, and aggregation from template → family → overall, so that the headline numbers are independently checkable.","section":null},{"comment":"Section 5.1 states that models were tuned within bounded search ranges, and free parameters include split ratio, synthetic row count, and family aggregation. For load-bearing ranking claims (RTF best; BayesNet best cost–fidelity), please report the final selected hyperparameters per model (or a compact appendix table) and state whether the overall query score is an unweighted mean of activated templates/families or a fixed weighted scheme. Without this, residual sensitivity of the reported orderings cannot be fully assessed even though the three-run stability study (Section 5.5) already bounds query-regeneration variance well.","section":null}],"minor_comments":[{"comment":"Figure 2 caption and surrounding text use abbreviated model names (T-DDPM, TPF, T-Syn) that are defined later in Table 2; define them at first use or move the abbreviation note earlier.","section":null},{"comment":"Table 1 uses u for user-specified scale; a one-line legend note would help readers scanning the comparison table.","section":null},{"comment":"In Finding 2 / Table 4, the prose sometimes cites slightly different distinct counts than the table (e.g., title 96,777 vs 96,779 elsewhere). Align the numbers for consistency.","section":null},{"comment":"Section 3.2 scopes out long join chains and multi-table settings; a single sentence in the abstract or introduction stating the single-table focus would set expectations earlier for database readers.","section":null},{"comment":"Appendix Table 9 heatmap uses TF for technical failure; ensure the main-text cost discussion (Figure 9, common-9) explicitly notes which models failed on which datasets so readers do not over-interpret missing cells.","section":null},{"comment":"Minor typography: “ocurring” → “occurring” (Introduction); “pilcrow” artifacts in the author line of the source text should be cleaned in the camera-ready version.","section":null}],"recommendation":"minor_revision","confidential_remarks":"I agree with the reader’s high-confidence accept lean: the residual risk is template representativeness, which the paper already bounds with public attribution, fixed Stage-1 library, constrained grounding, and ranking stability. My minor_revision is driven by missing explicit query-answer scoring formalization—important for a methods/benchmark paper at a serious venue—not by doubt about the five empirical patterns. Scope fit for a DB-oriented journal is good; the work is closer to systems/benchmark methodology than to pure generative-model theory."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"This is a real systems/DB contribution, not a metric rebrand. The new piece is treating recurring OLAP-style SQL as structural assessors for synthetic tabular data: 44 templates from 12 public sources, schema-aware grounding, 49 datasets × 11 generators, and family-level diagnostics that distance and ML-utility suites do not give you.\n\nWhat they do well is the instrument and the campaign. Templates are attributed and deduplicated; Stage 2 keeps profiling/binding/validation deterministic and puts the LLM only inside constrained realization. The five patterns are multi-dataset and multi-model: distance can look fine while query scores sit well below REAL (RTF best at ~0.75), high-cardinality support collapses, local conditional slices drop relative to global counterparts, tail recovery worsens as rarity tightens, and BayesNet sits on the practical cost–fidelity frontier. The three-run ranking stability study (Kendall W ~0.93, most cells move by at most one rank) is the right check for an LLM-assisted grounder, and the open code/data package makes the work usable.\n\nSoft spots are real but proportional. Representativeness of the 44 templates is the residual instrument risk for any query-centric bench; they bound it with public provenance, fixed policies, and the stability probe, and they scope out multi-table and DP. Score aggregation and the common-9 cost slice are free parameters, not hidden circularity. Citations and baselines look honest relative to Synthcity, SDGym, TabArena, and the TPC/ClickBench lineage.\n\nThis is for people who ship or select synthetic tables for sharing, system testing, or analytics prototyping, and for anyone building the next generator who needs failure modes beyond JSD/Wasserstein. I would bring it to reading group, cite the benchmark when I need query-answer fidelity, and send it to peer review. Engage.","headline":"Solid open benchmark that makes analytical SQL answers the evaluation object for synthetic tables; the five empirical patterns hold up under a large 49×11 campaign and a real stability check.","tokens_in":32421,"tokens_out":485,"would_cite":true,"duration_ms":5651,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Synthetic tables that look statistically close still fail the SQL queries analysts actually run.","keywords":["synthetic tabular data","query-centric fidelity","SQL benchmarks","tabular generative models","high-cardinality support","tail fidelity","fidelity-cost tradeoff"],"falsifier":"If regenerated or independently authored query sets that still target the same five analytical families reverse the model ranking or erase the reported gaps on high-cardinality support, local slices, or extreme-tail recovery, the central claim that current generators systematically fail query-centric fidelity would not hold.","tokens_in":32423,"feed_emoji":"📊","tokens_out":581,"duration_ms":5878,"temperature":0.7,"pith_summary":"Synthetic tabular data is usually scored on how close its columns look to the real ones, or how well a machine-learning model trained on it performs. This paper argues that those scores miss the structure that matters for everyday analytics: the answers to the kinds of SQL questions people run on tables. The authors build TabQueryBench by distilling recurring analytical logic from public query collections into 44 reusable templates, then grounding those templates to each dataset so the same query families can be run fairly across many generators. On 49 datasets and 11 generators they show that even the strongest model reaches only about three-quarters of real-data query fidelity, with systematic collapse on high-cardinality discrete columns, local filtered slices versus global counterparts, and rare-event tails. The practical upshot is a clearer map of where current generators break and a cost-quality frontier in which a simple Bayesian network often wins for users who care about both answer quality and generation speed.","feed_headline":"Synthetic tables score high on distance, fail on SQL","feed_subtitle":"Best model hits only 0.75 query fidelity; rare values and local slices break first","key_machinery":"TabQueryBench: 44 reusable SQL-shaped query templates, taxonomized into five families (subgroup, conditional, tail/rarity, missingness, cardinality/range) from public analytical sources and grounded to each dataset by a policy-guided template-to-SQL pipeline that keeps queries schema-aware and comparable.","core_discovery":"Distance-based and ML-utility scores can make synthetic tables look faithful while the same tables still give wrong answers to analytical SQL. Across 49 datasets, RealTabFormer is the best query-centric model yet only reaches 0.75 ± 0.15 (real data = 1.00), and failures concentrate on high-cardinality discrete support, local conditional slices, and extreme tails.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Synthetic tables ace distance scores but flunk SQL queries","Best model reaches only 0.75 query fidelity vs real data","Distance wins hide SQL failures in synthetic tables","Query fidelity lags: RealTabFormer tops at 0.75","High-cardinality and local slices break synthetic tables"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The 44 templates drawn from the chosen public sources, once grounded by the fixed pipeline, form a fair and representative set of structural tests for the analytical uses the paper claims to cover.","fun_headline_variants_meta":{"raw":{"variants":["Synthetic tables ace distance scores but flunk SQL queries","Best model reaches only 0.75 query fidelity vs real data","Distance wins hide SQL failures in synthetic tables","Query fidelity lags: RealTabFormer tops at 0.75","High-cardinality and local slices break synthetic tables"]},"model":"grok-4.5","effort":"low","cost_usd":0.003684,"raw_usage":{"total_tokens":1272,"prompt_tokens":891,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":36840000,"prompt_tokens_details":{"text_tokens":891,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":319,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":891,"tokens_out":62,"duration_ms":7966,"temperature":1.0,"reasoning_tokens":319,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T22:58:16.353808+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If regenerated or independently authored query sets that still target the same five analytical families reverse the model ranking or erase the reported gaps on high-cardinality support, local slices, or extreme-tail recovery, the central claim that current generators systematically fail query-centric fidelity would not hold.","supporting_citations":[],"review_version":1}