{"id":"10f97410-8122-4879-b122-a7acf93b1042","arxiv_id":"2606.12687","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"DICE-MMM is a two-stage diagnostic framework that separates forecasting accuracy from graph-aligned attribution in neural MMMs and localizes decoder bypass via controlled graph tests.","lead":"The paper shows that graph-based neural marketing mix models can achieve low forecasting error without routing attribution through the graph structure, a failure mode called attribution bypass. A smart generalist might read it to see why prediction accuracy alone does not validate causal claims in business analytics models.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"CIG/AR-CIG metrics may not isolate graph-routed perturbation influence even when the claim requires it","rationale":"The reader's weakest_assumption directly identifies the same point: the unverified mapping from the three tests to 'routes perturbation-induced influence through the supplied graph.' Because the full text was not supplied in the initial query and the abstract only invokes the tests without derivation or sanity checks, the concern remains load-bearing and keeps the verdict at UNVERDICTED.","tokens_in":1868,"tokens_out":364,"duration_ms":16282,"concrete_test":"Construct a minimal linear dynamical system on a known sparse graph G where each node's next state is a linear function of its graph neighbors plus noise; train a DICE-style decoder on it; compute AR-CIG on held-out perturbations. If AR-CIG remains near zero despite the decoder being forced to use G, the metric does not detect graph routing.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that low MSE hides attribution bypass and that graph-support selection is the bottleneck—rests on CIG, AR-CIG, and graph-swap tests correctly quantifying whether decoder sensitivity to perturbations travels through the supplied graph. The abstract states these tests evaluate decoder use after DICE training, and the frozen graph-swap result (nAUPRC rising from -0.044 to 0.894) is used to localize the failure. If the metrics can register low values for reasons unrelated to bypass (e.g., normalization artifacts, autoregressive leakage, or sensitivity to node ordering) or high values without actual graph mediation, the separation of forecasting accuracy from attribution alignment does not hold. No independent validation against a known graph-aligned linear system is described in the provided abstract.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper claims that graph-based neural marketing mix models (MMM) can achieve low forecasting error (e.g., MSE@7 ~0.004) via decoder autoregression or other mechanisms while failing to route perturbation-induced influence through the supplied graph (attribution bypass). It introduces the DICE-MMM two-stage framework (graph encoder training followed by frozen-encoder graph-safe decoder training) and uses CIG, AR-CIG, and graph-swap tests to separate graph recovery, forecasting accuracy, and graph-aligned decoder use. Experiments with controlled R/d/T swaps and an external multi-graph stress test show non-oracle graphs yield near-zero nAUPRC while an oracle graph reaches 0.807, localizing the bottleneck to graph-support selection rather than decoder capacity.","tokens_in":2040,"tokens_out":403,"duration_ms":20885,"significance":"If the CIG/AR-CIG metrics validly isolate graph-routed sensitivity, the work supplies a concrete diagnostic framework that prevents conflating forecasting performance with attribution reliability in neural MMMs. The frozen graph-swap result (nAUPRC rising from -0.044 to 0.894) and the oracle vs. learned contrast at matched MSE provide a falsifiable stress test that credits the separation of concerns and the use of external benchmarks.","major_comments":[{"comment":"Abstract and evaluation description: the claim that low MSE hides attribution bypass and that graph-support selection is the bottleneck rests on CIG, AR-CIG, and graph-swap tests correctly quantifying whether decoder sensitivity to perturbations travels through the supplied graph. No independent validation against a known graph-aligned linear system is described, leaving open the possibility that low nAUPRC arises from normalization artifacts, autoregressive leakage, or ordering sensitivity rather than bypass; this is load-bearing for the central separation of forecasting from attribution alignment.","section":"Abstract"}],"minor_comments":[],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the careful reading and for identifying a load-bearing assumption in our evaluation. We address the concern directly below and agree to strengthen the description of validation.","responses":[{"response":"The oracle graph-swap experiment provides the requested independent validation against a known correct graph: the identical DICE-trained decoder and data yield nAUPRC = -0.044 under the learned graph but nAUPRC = 0.894 under the oracle graph at matched MSE@7. This isolates the metric's sensitivity to graph alignment rather than decoder capacity or autoregressive leakage. The controlled R/d/T swap benchmarks further supply known ground-truth graphs for recovery and alignment tests. We did not include a separate linear-system benchmark, which is a fair observation. We will revise the abstract and evaluation section to explicitly frame the oracle swap and synthetic controls as validation against known graph-aligned systems and to discuss potential normalization/ordering sensitivities.","revision_made":"partial","referee_comment":"[Abstract] Abstract and evaluation description: the claim that low MSE hides attribution bypass and that graph-support selection is the bottleneck rests on CIG, AR-CIG, and graph-swap tests correctly quantifying whether decoder sensitivity to perturbations travels through the supplied graph. No independent validation against a known graph-aligned linear system is described, leaving open the possibility that low nAUPRC arises from normalization artifacts, autoregressive leakage, or ordering sensitivity rather than bypass; this is load-bearing for the central separation of forecasting from attribution alignment."}],"tokens_in":1537,"tokens_out":327,"duration_ms":16047,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The new piece here is the DICE-MMM setup that splits graph recovery, forecasting, and whether the decoder's perturbation sensitivity actually travels through the supplied graph. The abstract gives concrete numbers: no-graph and full-graph decoders hit MSE@7 around 0.004 with nAUPRC near zero, while the oracle graph reaches 0.807 at similar error, and swapping to oracle inputs lifts nAUPRC to 0.894. That separation is useful and not routine in the MMM literature.\n\nThe experiments use controlled swaps and an external multi-graph test, which helps localize the bottleneck. The claim that forecasting accuracy is not an attribution certificate holds up on the reported results.\n\nThe soft spot is whether CIG and AR-CIG actually isolate graph-routed influence. The stress-test note flags that these metrics could pick up artifacts from normalization or autoregression instead of true mediation through the graph. The abstract does not describe an independent check against a known linear system where the routing is guaranteed, so the interpretation rests on the tests behaving as intended. If that assumption slips, the localization to graph-support selection weakens.\n\nThis is for people working on neural marketing mix models who already care about attribution beyond raw forecast error. It is narrow but addresses a real practical gap. The quantitative demonstration is sharp enough that a serious editor should send it to referees rather than desk reject.","headline":"The paper's core contribution is a diagnostic that shows low forecasting MSE in graph MMMs can mask decoder bypass of the graph, with experiments localizing the issue to graph selection rather than model capacity.","tokens_in":2527,"tokens_out":364,"would_cite":false,"duration_ms":11159,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Low forecasting error can conceal attribution bypass in graph-based neural marketing mix models","keywords":["marketing mix models","graph-based neural networks","attribution bypass","forecasting accuracy","decoder alignment","graph recovery","stress testing"],"falsifier":"A controlled experiment in which swapping from a learned graph to an oracle graph produces no improvement in AR-CIG or nAUPRC scores for a DICE-trained decoder at fixed MSE.","tokens_in":2753,"feed_emoji":"","tokens_out":766,"duration_ms":36030,"temperature":0.7,"pith_summary":"The paper shows that forecasting accuracy and attribution are distinct tasks in graph-based neural marketing mix models. A high-capacity decoder can reach low error through shortcuts such as target autoregression without routing perturbation effects through the supplied graph. DICE-MMM is presented as a two-stage framework that trains a restricted graph-mediated decoder then a graph-safe latent decoder, and evaluates decoder use with CIG, AR-CIG, and graph-swap tests. Experiments on controlled swaps and an external stress test demonstrate that MSE remains low for both no-graph and full-graph decoders while attribution metrics stay near zero, yet an oracle graph yields high scores at comparable error. The work concludes that the bottleneck lies in graph-support selection rather than forecasting or decoder capacity.","feed_headline":"Low MSE conceals attribution bypass in neural MMM","feed_subtitle":"Graph selection is the bottleneck; only oracle graphs raise attribution scores while MSE stays low.","key_machinery":"DICE-MMM, a bounded two-stage diagnostic and training framework that enforces graph mediation in the decoder and evaluates alignment via CIG, AR-CIG, and graph-swap tests.","core_discovery":"Attribution bypass occurs when a decoder obtains low forecasting error while failing to route counterfactual sensitivity through the supplied graph. DICE-MMM separates graph recovery, forecasting accuracy, and graph-aligned decoder influence through stage-wise training with a restricted graph-mediated decoder followed by a graph-safe latent decoder. Decoder alignment is measured by CIG, AR-CIG, and graph-swap tests. Across R/d/T swaps and a multi-graph rawlog test, the framework shows that no-graph and full-graph decoders achieve MSE@7 around 0.004 with AR-CIG near or below zero, while an oracle graph reaches 0.807 nAUPRC at similar MSE. Frozen graph-swap tests localize the failure to the in","pith_inferences":["Graph selection methods for MMM applications could be ranked by how well they pass alignment tests such as graph-swap.","The separation of forecasting from graph-aligned influence may apply to other neural time-series models that use graphs for interpretability.","Practitioners could add these alignment checks before deploying neural MMM for channel attribution decisions."],"forward_implications":["Forecasting accuracy measured by MSE does not certify that the decoder routes influence through the graph for attribution.","The unresolved bottleneck is graph-support selection, not forecasting performance or decoder capacity.","In a sparse-target benchmark, no-graph and full-graph decoders reach MSE@7 around 0.004 while AR-CIG remains near or below zero.","An oracle graph input raises nAUPRC to 0.807 +/- 0.129 at comparable MSE, and graph-swap raises it from -0.044 to 0.894 for the same decoder."],"fun_headline_variants":["Low MSE does not certify attribution in neural MMM","Decoder bypass undetected at low forecast error","Bottleneck is graph not capacity in MMM attribution","Oracle graphs yield higher attribution at fixed MSE"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That the CIG, AR-CIG, and graph-swap tests correctly measure whether the decoder routes perturbation-induced influence through the supplied graph in a manner relevant to attribution.","fun_headline_variants_meta":{"raw":{"variants":["Low MSE does not certify attribution in neural MMM","Decoder bypass undetected at low forecast error","Bottleneck is graph not capacity in MMM attribution","Oracle graphs yield higher attribution at fixed MSE"]},"model":"grok-4.3","cost_usd":0.005451,"raw_usage":{"total_tokens":2719,"prompt_tokens":862,"num_sources_used":0,"completion_tokens":55,"cost_in_usd_ticks":54512000,"prompt_tokens_details":{"text_tokens":862,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1802,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":862,"tokens_out":55,"duration_ms":15082,"temperature":1.0,"reasoning_tokens":1802,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-27T10:02:00.417272+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled experiment in which swapping from a learned graph to an oracle graph produces no improvement in AR-CIG or nAUPRC scores for a DICE-trained decoder at fixed MSE.","supporting_citations":[],"review_version":1}