{"id":"fa56458b-ed21-4197-a752-6c9a3afc87b3","arxiv_id":"2608.09296","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":10,"one_line_summary":"CADEngBench, a layered benchmark of 600 parametric tasks and 120 assembly pairs, shows current AI models edit supplied CAD far more easily than they generate it, and rarely match reference physics or recover exact assembly joints.","lead":"CADEngBench is a benchmark that tests AI-generated CAD models on engineering behavior, not just appearance: valid solids, parameter response, functional edits, matched physics simulation, and assembly joints. Across eight AI models, editing an existing design proved much easier than creating one from scratch, while matching reference physics and exact joint grounding remained rare.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"A-track metrics may penalize mechanically valid but unrecorded joint alternatives, so the assembly-failure headline could overstate model error. Human expansion of the accepted joint set would settle it.","rationale":"The reader's weakest_assumption identifies exactly the load-bearing risk I find: A1/A2 scoring assumes the source-recorded joint alternatives are the complete set of correct answers. My stress test agrees and adds two pieces of internal evidence that make the concern concrete rather than hypothetical. First, Supplement A.2 explicitly disclaims exhaustiveness while the scoring code credits only source-recorded pairs, creating a mismatch between the stated semantics and the metric. Second, the paper's own discussion of rigid joints admits that several mechanically plausible flat mating faces can exist, so a model can be mechanically right and still fail Typed@k. This makes the concern load-bearing: the assembly track is one of two pillars of the central claim, and the abstract's assertion that plausible assembly predictions 'fail to recover the recorded mating relation' loses much of its force if the recorded relation is one valid option among several. I do not see a comparably strong threat to the P-track conclusions: L1 checks are tied to explicitly stated requirements, L2-Z tests a declared parameter against measured geometry, and L3 compares matched engineering quantities under a disclosed analysis card, so those numbers would survive even if the A-track ground truth were expanded. The lack of released code/data is a genuine reproducibility limitation, but it is secondary to the ground-truth completeness question for the correctness of the empirical claims; the reader's conditional verdict already accounts for both. A human-annotation study that expands the accepted joint set would either confirm the benchmark's assembly conclusions or show them to be overstated, and the recommended verdict would change only in the latter case. Since the reader already conditioned acceptance on this assumption plus artifact availability, my assessment leaves the verdict unchanged.","tokens_in":26322,"tokens_out":5650,"duration_ms":64098,"concrete_test":"Select a stratified sample of 30-40 CADEngBench-A pairs, oversampling rigid and pin-slot items where ambiguity is acknowledged. Have two independent mechanical engineers, blind to source records, enumerate every mechanically valid joint family and exact face/edge pair using the same candidate atlas and body geometry; add the union to the accepted-alternative set. Recompute Entity@1, Typed@1, Typed@3, and A2 E2E for all eight models with this expanded set. If Typed@1 remains below about 25% and A2 E2E below about 20%, the assembly-failure conclusion is robust; if these rates rise substantially (for example, Typed@1 above 40%), the current metrics materially overstate assembly failure and the abstract's assembly claim needs qualification. Run the proposed alternative relations through PyBullet as a secondary check that they actually satisfy the allowed-motion contract.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is that the Fusion 360 Assembly-Joint source records provide a complete set of correct joint alternatives for each body pair. The paper states in Supplement A.2 that 'the source records provide positive joint alternatives rather than an exhaustive catalogue of every mechanically feasible relation' and that 'the absence of a relation is not treated as a verified negative example,' but A1 scoring (Supplement F.2) still credits only source-recorded role-aware entity pairs. The paper's own rigid-joint discussion concedes the concrete failure mode: 'two rigidly connected bodies often admit several plausible flat mating faces, so a model can correctly infer no relative motion, yet still select a different face pair than ground truth, failing the stricter Typed metric.' If unrecorded valid pairs are frequent, then the headline 'only 14.8% of rank-1 hypotheses recover the recorded joint' and the conclusion that assembly grounding fails beyond region localization conflate dataset ambiguity with model error. A model that proposes a mechanically valid cylindrical or planar mating not present in the source is scored a miss even if it has found a correct engineering relation. This directly threatens the assembly half of the central claim: if the 'correct' answer set is incomplete, the evaluation is partly measuring source-record completeness rather than model competence. The P-track claims (L1, L2, L3) rest on explicitly stated requirements and are not similarly affected.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CADEngBench, a two-track benchmark for evaluating CAD generation and assembly behavior beyond appearance. CADEngBench-P comprises 300 parametric parts used for 600 zero-to-CAD and functional-editing tasks, evaluated through B-Rep validity (L0), engineering and DFM requirements (L1), parametric perturbation and functional editing (L2-Z and L2-E), and matched linear-static FEA in CalculiX (L3). CADEngBench-A comprises 150 body pairs (120 evaluation pairs after development exclusion) scored on ranked joint retrieval, exact face/edge grounding, joint-frame prediction, and kinematic verification in PyBullet. Eight multimodal models are evaluated. The reported results show that executable code frequently violates engineering requirements, matched FEA agreement is rare, editing is easier than generation, and assembly predictions often locate the correct region but rarely recover the source-recorded joint and mating entities.","tokens_in":26741,"tokens_out":6266,"duration_ms":63563,"significance":"If the evaluation is valid, the benchmark addresses a genuine gap: existing CAD benchmarks often stop at shape similarity or executable-code checks, whereas CADEngBench attempts to measure behavioral properties such as parametric control, functional editing, physical response, and assembly grounding. The paper has notable methodological strengths: scoring is based on deterministic B-Rep measurement and replay; thresholds and analysis cards are disclosed; confidence intervals use item-clustered bootstraps; pairwise comparisons use exact McNemar tests with Holm correction; and limitations are stated explicitly. The two-track design and the L0-L3 gated hierarchy are useful contributions. The principal risk is the completeness of the assembly ground-truth set: because source-recorded joint alternatives are treated as the accepted answers, mechanically valid but unrecorded mating relations are scored as misses, which could overstate the assembly-failure conclusions.","major_comments":[{"comment":"The L2-Z results are reported with model-specific denominators (N_s = 287-300 in Table 2), while the experimental protocol states that malformed responses, timeouts, and build failures are retained as failures. If an item is excluded from the denominator because a model's output does not satisfy the required interface, the reported pass rates are not directly comparable across models and are not strict item pass rates. The paper should either report L2-Z over the fixed 300-item denominator for every model, with excluded cases counted as failures, or justify why model-specific evaluability is appropriate and show the sensitivity of the conclusions to the choice of denominator.","section":"Supplement A.2, F.1, F.2; main text 'Complex Assembly Behaviour'"}],"minor_comments":[{"comment":"The main text states that '554 solve but violate engineering limits or disagree with the reference stress or deformation,' but Figure 7 shows 270 engineering-limit failures and 593 FEA-quantity mismatches, which sum to 863. Please correct the number or clarify which subset of outcome categories is included in the 554 figure.","section":"Results, Fig. 7"},{"comment":"The sentence 'Functional-edit L2 also falls from these slices show that the remaining difficulty is not merely parsing an edit request' is grammatically incomplete. It should be rewritten to state the intended observation, for example that the functional-edit gap is concentrated in histories involving joins, cuts, or multiple bodies.","section":"Results, 'Generation and Editing'"},{"comment":"When describing the source relations in the A1 prompt, the manuscript says the dataset 'consolidates alternative joints observed for identical pairs of parts across different assemblies'; this is helpful and should also be stated in the main text near the assembly results, since it directly qualifies how the 14.8% result should be read.","section":"Supplement B.5 and F.1"},{"comment":"The log-ratio tolerance is described in the main text as a multiplicative comparison, and the supplement gives equivalent ratio intervals in Table 8. The connection between them is clear, but a one-sentence note that log(1.25) corresponds to the [0.8, 1.25] interval would prevent reader confusion, especially for the asymmetric body-acceleration interval [2/3, 1.5].","section":"Supplement D.3 and main Eq. (3)"},{"comment":"The caption for Table 2 says 'P uses 300 tasks except L2-Z,' but the L3 reach and pair-pass columns also have model-dependent denominators. Please clarify that L3 pair pass is conditional on comparable generated/reference state pairs, as is done in the text.","section":"Table 2 caption"}],"recommendation":"major_revision","confidential_remarks":"The benchmark is valuable and the general design is sound, but I am not comfortable accepting until the assembly ground-truth completeness issue is addressed. The paper's own admission that the source records are not exhaustive makes the strict A1/A2 scoring a correctness risk for the assembly half of the central claim. A human-validated alternative set on a subset, or a clear reframing of the claims as recorded-joint retrieval, would resolve the concern. The L2-Z denominator issue is smaller but should also be fixed. I did not find circularity: ground-truth labels come from source artifacts and deterministic replay, and LLM/VLM annotations are not used as scoring labels."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth your time. This is a benchmark paper that actually does what it claims. The novelty is the combination: the same generated program is checked for executability, engineering/DFM requirements, parameter-family behavior, functional editing, and matched FEA; assembly hypotheses are scored on exact B-Rep entity grounding, joint type, frame, and executed motion. Prior benchmarks cover one or two of these. The paper also ships unusually careful methodology: deterministic B-Rep measurement, disclosed thresholds, bootstrap confidence intervals clustered by part/assembly graph, Holm-corrected McNemar tests, and a supplement that documents scoring paths and failure categories. The empirical claims that editing outperforms generation, that matched FEA is rarely passed, and that assembly models find the region but not the recorded joint are all backed by the reported numbers.\n\nThe soft spots are real but mostly disclosed. The most important is the A-track ground truth. The source Fusion data provides positive joint alternatives, not an exhaustive set of mechanically valid relations. The paper says this in Supplement A.2 and even gives the concrete example: two rigidly connected bodies have several plausible flat mating faces, so a model can correctly infer no relative motion but pick a different face pair and be scored wrong. Yet the A1/A2 metrics only credit source-recorded pairs. So the 14.8% rank-1 joint recovery headline likely overstates model error, conflating dataset incompleteness with failure. This doesn't touch the P-track results, which rest on explicit requirements and measured geometry. But it does mean the assembly half of the central claim needs a caveat. The fix is straightforward: human expansion of the accepted joint set for a sample of pairs, or a sensitivity analysis.\n\nThe other practical problem is that no code, data, or evaluation artifacts are released with the preprint. That makes the benchmark un-auditable today. For a benchmark paper, that matters. I'd also note that L3 load cases are benchmark-assigned rather than source conditions, which is a sensible controlled design and is disclosed, not a flaw.\n\nOverall: this deserves a serious referee. I'd conditionally accept it, asking for artifact release and either an expanded A-track ground-truth set or a sensitivity analysis. If the authors deliver that, it will be a useful reference. I'd bring it to reading group.","headline":"A serious, unusually honest CAD benchmark that moves evaluation from appearance to engineering behavior; the assembly track has a disclosed ground-truth completeness issue that should temper the headline, not sink the paper.","tokens_in":27163,"tokens_out":3483,"would_cite":true,"duration_ms":31723,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"CADEngBench argues that CAD evaluation must test engineering behavior instead of appearance, and shows that current models fail most such checks.","keywords":["parametric CAD generation","functional editing","boundary representation","finite element analysis","assembly joint grounding","kinematic verification","design for manufacturability","LLM benchmark"],"falsifier":"Take a sample of assembly pairs and ask a human engineer to list every mechanically plausible joint between the two bodies without seeing the source records; if a large share of engineer-approved mates are not in the recorded set, then CADEngBench-A's A1/A2 pass rates understate true assembly competence. In the opposite direction, exhibit one scored-miss rank-1 hypothesis that is kinematically valid in simulation, meaning the predicted joint type, entity pair, and frame allow exactly the permitted motions and block the forbidden ones; a single such case shows the metric can reject a working hypothesis.","tokens_in":26092,"feed_emoji":"⚙️","tokens_out":7057,"duration_ms":65866,"temperature":0.7,"pith_summary":"The paper's central claim is that a CAD artifact is not engineering-grade just because it looks correct or executes: it must satisfy stated requirements, respond predictably to parameter changes, survive functional edits without collateral damage, match a reference structural response under the same analysis contract, and connect to other parts through valid joints. To test that claim it builds CADEngBench, with a parametric track of 300 parts used for 600 generation and editing tasks and an assembly track of 150 body pairs. Across eight code-capable models, only 432 of 1,030 executable programs pass the engineering and DFM checks, matched finite-element agreement is rare (best model 46.3% of comparable parameter states), and only 14.8% of rank-1 assembly hypotheses recover the recorded joint family and mating entities. The point is that appearance-based or executability-based evaluation overstates progress, while these layered checks reveal where a plausible artifact stops behaving like an engineering model.","feed_headline":"Only 42% of executable CAD passes engineering checks","feed_subtitle":"New benchmark tests parametric parts and assembly joints beyond looks; best model matches reference FEA in 46% of cases.","key_machinery":"The load-bearing object is a layered evaluation hierarchy in which a candidate passes a stage only after satisfying all earlier stages. The parametric track maps a program $C(\\theta)$ to a solid $B = C(\\theta)$ and tests: L0 executes and re-imports a STEP solid; L1 checks engineering requirements and DFM screens (minimum wall 1.0 mm, hole diameter at least 2.0 mm, depth-to-diameter ratio at most 8); L2-Z rebuilds the same $\\text{build}(\\text{params})$ function over a fixed set of parameter values and requires every state to keep the declared parameter-to-geometry relation while preserving protected invariants; L2-E applies a source-native edit and checks the requested target plus hidden preservation checks; L3 compares generated and reference solids under identical materials, supports, loads, and mesh profiles, using log-ratio disagreement $\\delta = |\\log((q(C)+\\epsilon)/(q(R)+\\epsilon))|$ on 95th-percentile von Mises stress, normalized displacement, and compliance or stress concentration. The assembly track takes a hypothesis $h = (j, e_A, e_B)$ with a joint family $j$ and exact face or edge entities from each body, scores role-aware retrieval (Entity@k and Typed@k), then predicts the joint frame and executes allowed and blocked motion in a rigid-body simulator. This machinery is what separates “looks right” from “behaves right”.","core_discovery":"On the paper's own terms, the discovery is that each evaluation layer removes a large share of apparent successes, and the layers behave like nearly independent capabilities. Executable code can violate design intent; parameter families that rebuild successfully can still diverge from the reference stress and deformation response; and assembly predictions that localize the correct region often fail to recover the recorded joint type and exact B-Rep entities. The evidence includes 1,030 of 2,400 generated programs passing L0 execution but only 432 also passing L1 engineering and DFM requirements; editing supplied CAD passing far more often than zero-to-CAD generation (1,071 item–model pairs pass editing only versus 231 generation only); L3 matched FEA pair pass reaching 46.3% for the best model and being uncorrelated with how many simulations a model can run (ρs = 0.024); and in assembly, 542 of 960 requests putting a correct B-Rep entity pair at rank 1, but only 143 of those also specifying the correct joint family and body ordering, with end-to-end kinematic pass at 15.8% for the best model. The conclusion is that CAD evaluation must measure engineering behavior, not plausible shape.","pith_inferences":["If the source joint records are incomplete, A1/A2 pass rates likely underestimate assembly competence; a human-annotation study that counts mechanically valid but unrecorded mates as correct would quantify that ceiling.","The pin-slot 0% result rests on six items, which the paper itself flags as too few for a broad claim; a larger pin-slot subset could change family-level conclusions.","The L2-Z perturbation protocol could transfer to other generative design domains, such as layouts, structural frames, or circuit boards, where a plausible default hides wrong parameter wiring.","The “editing is easier than generation” result may be partly a function of construction-history coupling, so future benchmark design should stratify edit tasks by how deeply the edited feature is joined or cut into existing geometry."],"forward_implications":["A benchmark that scores only executability or shape similarity overstates capability: 42% of executable generated programs fail stated engineering or DFM requirements, so requirement checks are needed to detect failures of design intent.","Parameter-family testing is necessary because a parameter can appear correct at its default value while being inert or coupled to the wrong feature elsewhere; L2-Z requires every evaluated state to behave.","A successful FEA solve does not establish correct physics: L3 reach is uncorrelated with agreement to the reference (ρs = 0.024), so benchmarks should compare generated and reference structural response under the same analysis contract.","Assembly evaluation should require exact B-Rep entity grounding and typed joint recovery rather than region location; the gap between Entity@1 and Typed@1 shows that most correct localizations still miss the recorded joint.","Capability ranks are weakly associated across generation, editing, parametric, and assembly scores, so no single model ranking describes CAD ability and multi-stage evaluation is required."],"supporting_citations":[{"why":"Supplies 159 executable CadQuery programs and part families that form the BenchCAD subset of CADEngBench-P, providing the generation and editing source artifacts.","marker":"Zhang et al. 2026"},{"why":"Supplies the Fusion 360 Gallery replayable construction histories used for 141 reconstruction-based parts and their functional-edit tasks.","marker":"Willis et al. 2021"},{"why":"Introduces the assembly-joint dataset from which CADEngBench-A draws its 150 body pairs, recorded joint alternatives, frames, and motion semantics.","marker":"Jones et al. 2021"},{"why":"Provides the finite-element mesh generator that produces matched quadratic tetrahedral meshes for the reference and candidate solids in L3.","marker":"Geuzaine and Remacle 2009"},{"why":"Provides the CalculiX solver that runs the matched linear-static analyses and produces the stress and displacement fields used in L3 comparisons.","marker":"Dhondt 2004"}],"fun_headline_variants":["CAD that looks right often fails engineering checks","New benchmark: appearance alone isn't enough for CAD","Only 42% of executable CAD passes engineering validation","Benchmark shows CAD needs engineering tests, not just looks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The assembly scores treat the source-recorded joint alternatives as the complete set of correct answers, so a mechanically valid but unrecorded mating is counted as a miss; the paper notes that absence of a relation is not treated as a verified negative example.","fun_headline_variants_meta":{"raw":{"variants":["CAD that looks right often fails engineering checks","New benchmark: appearance alone isn't enough for CAD","Only 42% of executable CAD passes engineering validation","Benchmark shows CAD needs engineering tests, not just looks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000523,"raw_usage":{"total_tokens":2552,"prompt_tokens":994,"completion_tokens":1558,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":610,"completion_tokens_details":{"reasoning_tokens":1496}},"tokens_in":610,"tokens_out":1558,"duration_ms":15763,"temperature":1.0,"reasoning_tokens":1496,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:55:59.237952+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a sample of assembly pairs and ask a human engineer to list every mechanically plausible joint between the two bodies without seeing the source records; if a large share of engineer-approved mates are not in the recorded set, then CADEngBench-A's A1/A2 pass rates understate true assembly competence. In the opposite direction, exhibit one scored-miss rank-1 hypothesis that is kinematically valid in simulation, meaning the predicted joint type, entity pair, and frame allow exactly the permitted motions and block the forbidden ones; a single such case shows the metric can reject a working hypothesis.","supporting_citations":[],"review_version":1}