{"id":"e28c69b2-d1f0-4bb4-b308-85fa9ea4ec07","arxiv_id":"2608.04038","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A source-blind protocol with structural estimators recovers the shape of seed-level uncertainty in text-to-video generation of modern art, but adding that structure does not improve prediction of semantic disagreement over a single scalar.","lead":"A new protocol analyzes how a text-to-video model's random seeds interpret modern art, classifying the seed-to-seed uncertainty into shapes such as compact, one outlier, two modes, or diffuse, and separating useful diversity from a failure to cover the original artwork. It matters because current uncertainty scores return one number, which cannot distinguish these different situations, guide curation, or tell when a generator misses the reference.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline diagnostic numbers come from synthetic seed sets; real-world transfer is unverified, and the structural components that drive them are the least stable on real data (Table 5).","rationale":"The paper is internally coherent, reports negative results, and avoids overclaiming: RQ6 admits structure does not beat a scalar, and Table 5 admits structural components are fragile. My concern is not internal inconsistency but the gap between the controlled experiment and real deployment. The reader's weakest assumption was synthetic-to-real transfer; I agree and sharpen it with Table 5: the exact components that achieve AUROC 1.00 in the controlled setting have the worst cross-encoder stability on real data. This does not refute the protocol; it means the headline numbers should be read as proof-of-concept on idealized geometry, and the conditional verdict, with release of artifacts and a direct transfer check, is appropriate. I therefore recommend UNCHANGED.","tokens_in":7159,"tokens_out":8994,"duration_ms":86528,"concrete_test":"Re-run Experiment 2 with synthetic sets calibrated to the real corpus: for each of the 250 real seed sets, estimate the within-set covariance and the empirical distribution of pairwise clip distances; generate compact, 3+1, 2+2, and diffuse sets by sampling from these empirical distributions (or by perturbing real clip embeddings with noise at the observed scale), then recompute balanced topology accuracy and outlier AUROC. If accuracy falls materially below 0.98, the original controlled geometry was too easy and the headline numbers do not transfer.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central diagnostic claim (Sec. 5.2, Table 1) is established only on 'controlled seed sets of known geometry' (compact, 3+1, 2+2, diffuse), yet the paper never states how those sets were generated or how cluster separation was chosen. If the synthetic sets are well-separated draws from the same parametric families (vMF, Kent, mixtures) that the protocol fits, then the 0.98 balanced accuracy and 1.00 AUROC measure the internal consistency of the clustering objective on idealized inputs, not recovery on real Wan2.1 seed distributions. On the real corpus there is no topology ground truth; RQ4/RQ5 validate a median split of uncertainty and coverage, not the predicted topology. The paper's own Table 5 shows the structural components behind the headline, outlier gap and mode ratio, have the lowest cross-encoder rank agreement on real data (mean Spearman 0.23 and 0.35, min 0.17 and 0.22), while only simple dispersion estimators are stable (0.80). So the transfer from controlled geometry to real, four-seed, cross-encoder settings is the load-bearing assumption, and it is unsupported. The paper's candor about the 2+2 failure (AUROC 0.49) and about structural components needing more seeds narrows but does not close this gap.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces a protocol for quantifying the structure, rather than only the magnitude, of seed-to-seed variation in text-to-video generation when rendering modern art. The authors construct a corpus of 250 artworks, each animated under four seeds by Wan2.1 14B and embedded with four encoders, and they propose seven source-blind and six reference-aware estimators, a distributional profile with an explicit topology selector, and an artwork-level resampling protocol. Eight research questions are investigated; the headline results are that the topology selector recovers controlled synthetic geometries with balanced accuracy 0.98, that an outlier-gap component detects the 3+1 configuration at AUROC 1.00 where a scalar reaches 0.35, and that high-uncertainty artworks split into reference-covering (n=97) and reference-missing (n=56) groups. The paper also reports two negative results: the full distributional profile does not outperform pairwise dispersion for predicting semantic disagreement, and source-blind uncertainty does not track reference fidelity.","tokens_in":7438,"tokens_out":6687,"duration_ms":61040,"significance":"If the central claims hold, the paper would be a useful contribution to generative-media uncertainty quantification: it provides a reusable protocol, a purpose-built corpus, and honest negative results. The statistical hygiene is a clear strength—artwork-level bootstrap, permutation p-values, Benjamini-Hochberg correction, a calibration/validation/test split, and explicit reporting of failures such as the 2+2 multimodality AUROC 0.49 are all present. However, the headline diagnostic numbers are produced on controlled synthetic seed sets, and the paper's own cross-encoder analysis (Table 5) shows that the structural components driving those headline results are the least stable on real data. The significance is therefore prospective: the transfer from synthetic geometry to real Wan2.1 seed distributions is the load-bearing assumption, and it is not yet supported.","major_comments":[{"comment":"The controlled recovery experiment is the sole evidence for the headline topology balanced accuracy (0.98) and outlier AUROC (1.00), but the manuscript never describes how the 'controlled seed sets of known geometry' were generated: the parametric families, separation parameters, sample sizes, and whether the same four-encoder embedding pipeline was used are all omitted. Without these details, the reported numbers may largely reflect the topology selector and outlier-gap component re-identifying the very families they were designed to detect, rather than recovery on real Wan2.1 seed distributions. Please provide the full generation procedure and add a real-data validation, for example with human-annotated or independently derived topology labels on a sample of the 250 artworks.","section":"§5.2, Table 1"},{"comment":"The abstract's claim that the protocol is reliable 'from three seeds and across encoders' is not supported for the structural components. Table 4 shows 3-of-4 seed reliability only for simple dispersion estimators, while Table 5 reports mean cross-encoder Spearman values of 0.23 for outlier gap and 0.35 for mode ratio, with minima of 0.17 and 0.22. Because the real corpus has no topology ground truth and RQ4/RQ5 validate median splits rather than predicted topology, the transfer from controlled synthetic geometry to real four-seed, cross-encoder settings is the central unsupported step. Please either provide such a transfer test or restrict the reliability claims in the abstract and conclusions to the simple estimators.","section":"§5.6, Table 5; Abstract"},{"comment":"The protocol describes a six-model ablation (vMF, Kent, ACG, Student t, kernel, mixture) and states it gauges how much parametric machinery four seeds can support, but no results from this ablation are reported anywhere in Section 5. This is a missing component of a stated protocol element. Please report the ablation or remove it from the protocol description.","section":"§3, 'Distribution model ablation'"},{"comment":"The two-axis taxonomy separating reference-covering from reference-missing diversity is based on median splits of uncertainty and coverage, but the manuscript does not report a statistical test establishing that the two axes are independent or that the quadrant proportions differ from chance beyond the descriptive counts. Since the separation of high-uncertainty artworks into n=97 and n=56 is one of the abstract's claims, please add a permutation or bootstrap test for the association between the uncertainty and coverage axes.","section":"§5.4, Table 3"}],"minor_comments":[{"comment":"The header 'CVR 2' should be written as 'CV R²' or 'CV R2' to avoid confusion with a metric named CVR.","section":"Table 2"},{"comment":"The anisotropy response values (ρ = −0.14 and 0.41 under directional and isotropic perturbation) are reported without stating the unit of correlation; please clarify whether these are computed over perturbation levels, seed sets, or artworks, and report confidence intervals.","section":"§5.2"},{"comment":"The temporal decomposition values (=0.016, 0.289, 0.144) are introduced without an equation or explicit definitions of 'within video instability,' 'across seed same timestep disagreement,' and 'trajectory disagreement'; please define these formally or point to a supplement.","section":"§5.7"},{"comment":"The term 'source blind' is used in the abstract before it is defined; a one-sentence definition in the abstract would improve readability.","section":"Abstract and §3"},{"comment":"The novelty claims 'first corpus' and 'first systematic study' should be checked against concurrent work, given that the related-work section already cites several 2025–2026 preprints in closely related areas.","section":"§4"}],"recommendation":"major_revision","confidential_remarks":"The paper is honest, well structured, and statistically careful, but the gap between synthetic controlled recovery and the real-data structural claims is the central weakness. The omitted synthetic-generation details and the low cross-encoder stability of the structural components (Table 5) make the headline abstract claims premature. This is addressable in a major revision by adding a real-data transfer test or by substantially tempering the claims; I do not see a load-bearing error that would force rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this paper builds a reusable protocol and a real corpus for a genuinely new question — not just how much seeded video generations disagree, but how they disagree (compact vs. outlier vs. bimodal vs. diffuse), and whether the disagreement still covers the reference artwork. The corpus is real: 250 modern-art captions, Wan2.1, four seeds, four encoders, references held out. The statistical hygiene is better than most of what crosses my desk: artwork-level bootstrap, permutation p-values, BH correction, calibration/validation/test split, and explicit negative results. The authors openly report that their richer profile does not beat a scalar for predicting semantic disagreement, and that source-blind uncertainty is not reference fidelity. That candor is real and earns credit.\n\nThe genuinely new pieces: the distributional profile itself, the eight-question identification protocol, the reference-covering vs. reference-missing taxonomy (Table 3), and the 1000-video corpus. The individual estimators are mostly prior work, but the combination and the measurement framework are new, and the two-axis taxonomy is a clean, interpretable result.\n\nThe soft spot, and it is the main one: the headline numbers (topology balanced accuracy 0.98, outlier AUROC 1.00) come from \"controlled seed sets of known geometry,\" and the paper never says how those synthetic sets were generated or how separation was chosen. If they are well-separated draws from the same parametric families (vMF, Kent, mixtures) that the protocol fits, then the recovery exercise is partly a self-consistency check. On the real corpus there is no topology ground truth, so RQ4/RQ5 validate a median split, not the predicted structure. Table 5 makes the concern concrete: the structural components behind the headline — outlier gap and mode ratio — have the lowest cross-encoder rank agreement (mean Spearman 0.23 and 0.35), while simple dispersion is stable at 0.80. So the paper's own evidence says the richer structure is fragile at four seeds and across encoders. The authors are upfront about some of this, including the 2+2 failure (AUROC 0.49), but the transfer from synthetic to real remains load-bearing and unsupported. That is not a fatal flaw; it is a missing experiment. I would want to see synthetic sets matched to real seed distributions, or a small hand-labeled real set to check topology.\n\nMinor: the release of artifacts is claimed but not verifiable from the text, and the thresholds are given but not the generation details for the synthetic sets.\n\nWho this is for: people building uncertainty and quality metrics for generative video, and creative-AI folks looking for a measurement language. It deserves a serious referee — the corpus and the honest negative results give the field a real baseline. I'd send it to peer review with a request to address the synthetic-to-real gap directly. I'd cite it, probably for the corpus and the coverage/uncertainty separation.","headline":"A careful, honest first protocol for characterizing the structure of seed-to-seed uncertainty in text-to-video generation, but its headline recovery numbers come from synthetic seed sets and the transfer to real data is not yet shown.","tokens_in":7959,"tokens_out":2321,"would_cite":true,"duration_ms":21672,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that the seed-to-seed disagreement of a text-to-video model on modern art has a recoverable distributional shape—compact, outlier, bimodal, or diffuse—identifiable from a handful of seeds without seeing the original…","keywords":["uncertainty quantification","generative video model","modern art animation","seed-to-seed diversity","distributional profile","source blind evaluation","Wan2.1","semantic entropy"],"falsifier":"Have human raters label each real artwork's four-seed set by the same five topology classes (compact, $3{+}1$, $2{+}2$, $2{+}1{+}1$, all-distinct) and compare the protocol's predictions to those labels; agreement near chance, or an outlier AUROC that fails to beat a scalar on real sets, would show the synthetic shapes are easier to separate than real seed distributions.","tokens_in":6963,"feed_emoji":"🎨","tokens_out":6585,"duration_ms":55909,"temperature":0.7,"pith_summary":"The paper argues that when a text-to-video model animates the same modern artwork under different random seeds, the disagreement between the resulting films is structured signal, not noise: it can look compact, like a dominant reading with an outlier, like two competing modes, or like diffuse instability. It introduces a source-blind protocol that decomposes a set of seeded generations into these named structures and shows the structure is recoverable from as few as three or four seeds. On controlled seed sets, the protocol classifies seed-set topology at balanced accuracy 0.98 and isolates the outlier configuration at AUROC 1.00, where a standard dispersion scalar reaches only 0.35. The paper also separates high seed uncertainty into 97 artworks whose seeds still cover the reference and 56 whose seeds miss it, showing high diversity is not one phenomenon. Two negative results bound the claim: a full distributional profile does not beat plain pairwise dispersion at predicting semantic disagreement (cross-validated $R^2$ 0.240 vs 0.243), and source-blind uncertainty tracks interpretation, not fidelity to the hidden artwork.","feed_headline":"Four seeds expose the shape of AI art uncertainty","feed_subtitle":"New protocol tells compact, outlier, bimodal and diffuse readings apart—and flags when no seed matches the artwork.","key_machinery":"The carrying object is the distributional profile, a suite of seven source-blind and six reference-aware estimators combined with a controlled topology-recovery experiment. For each seed set, the profile computes robust spread, outlier influence, an explicit discrete topology via the partition $\\hat{C}_i = \\arg\\min_C S_{\\text{within}}^i(C)+\\lambda|C|$ over candidate shapes, multimodality, anisotropy, leave-one-seed influence, and reference coverage, with calibration constants $\\tau=0.76$, $\\lambda=0.50$, $h=1.07$, and a 50-dimensional PCA space fit on a held-out split. Synthetic seed sets of known geometry (compact, $3{+}1$, $2{+}2$, diffuse) provide the ground truth that shows the topology selector and outlier gap recover structure a scalar cannot; a distribution model ablation (vMF, Kent, ACG, Student $t$, kernel, mixture) then gauges how much parametric machinery four seeds can support.","core_discovery":"The central claim is that the seed-to-seed variance of a text-to-video model on ambiguous modern art has a recoverable distributional shape, and that shape—not just its magnitude—can be identified from a handful of generations without ever consulting the original artwork. The paper's distributional profile names the structure of a seed set (compact, dominant-plus-outlier, bimodal, diffuse, directional) and recovers it reliably: topology balanced accuracy 0.98, outlier detection AUROC 1.00 versus 0.35 for the directional-concentration scalar. On the 250-artwork corpus, the protocol divides the high-uncertainty cases into reference-covering diversity (97 artworks) and reference-missing diversity (56), meaning the same scalar uncertainty can correspond to meaningful variation or to failure to capture the artwork. The paper is equally explicit about what structure does not buy: adding outlier, mode, and anisotropy components does not improve prediction of cross-seed semantic disagreement over a single dispersion scalar, and source-blind uncertainty is interpretive rather than reconstructive.","pith_inferences":["Inference: If the controlled synthetic shapes transfer to real generation noise, the same protocol could serve as a cheap runtime quality gate for any text-to-video model animating stylistically ambiguous input, with no access to the original artwork.","Inference: The two-axis split suggests a concrete failure-mode test: a model whose high-uncertainty outputs are mostly reference-missing should be treated as unreliable for art animation, whereas a model whose high-uncertainty outputs are mostly reference-covering is producing legitimate alternative readings.","Inference: Running the identical protocol on a second text-to-video model or on a natural-image corpus would require re-estimating the calibration constants ($\\lambda$, $h$, $\\tau$) and would reveal whether the 0.98 topology accuracy is tied to Wan2.1's particular seed behavior or is a general property of seeded video generation."],"forward_implications":["A curator can route a high-uncertainty artwork to different actions—accept, discard a seed, present alternatives, or abstain—based on whether the protocol names the disagreement compact, outlier, bimodal, or diffuse.","Because three-of-four and even two-of-four seed subsets reproduce the full four-seed rankings at Spearman 1.00 for the simple estimators, future studies can run the protocol with fewer generations and save compute.","The negative RQ6 result implies that for predicting semantic disagreement, adding structural components to a scalar does not help; a practitioner should not expect a richer scalar to predict meaning better than pairwise dispersion.","Source-blind uncertainty measures interpretation, not reference fidelity, so any deployment that needs to know whether seeds still captured the original artwork must use the reference-aware evaluation branch rather than the deployable source-blind scores.","The simple estimators stay stable across four encoders (pairwise and DCU agree at mean Spearman 0.80), while the richer structural components are fragile, so the protocol's portable parts are the dispersion and coverage estimators."],"supporting_citations":[{"why":"Supplies the Wan2.1 14B text-to-video model on which the entire 250-artwork, 1000-video corpus is generated.","marker":"[Team Wan et al., 2025]"},{"why":"Supplies the directional concentration uncertainty (DCU) scalar baseline that the protocol's outlier detection outperforms (AUROC 1.00 vs 0.35).","marker":"[Chattopadhyay et al., 2026]"},{"why":"Supplies the semantic entropy baseline and the meaning-equivalence partition approach that the protocol includes as one of its seven source-blind estimators.","marker":"[Kuhn et al., 2023]"},{"why":"Extends semantic entropy and serves as a second baseline for uncertainty quantification that the paper contrasts with its structural profile.","marker":"[Farquhar et al., 2024]"}],"fun_headline_variants":["AI art uncertainty has a shape, not just a size","Four seeds expose the shape of AI art uncertainty","Beyond spread: decoding AI art's uncertainty structure","Source-blind protocol maps AI art's interpretation modes","Seed variance in AI art: from scalar to structure"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline diagnostic numbers come from synthetic seed sets crafted to look like compact, dominant-plus-outlier, two-plus-two, and diffuse configurations, and the paper assumes real Wan2.1 seed sets on modern art vary in the same way; the 250 real artworks have no ground-truth topology label, so nothing in the paper directly verifies that transfer.","fun_headline_variants_meta":{"raw":{"variants":["AI art uncertainty has a shape, not just a size","Four seeds expose the shape of AI art uncertainty","Beyond spread: decoding AI art's uncertainty structure","Source-blind protocol maps AI art's interpretation modes","Seed variance in AI art: from scalar to structure"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000712,"raw_usage":{"total_tokens":3256,"prompt_tokens":1048,"completion_tokens":2208,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":664,"completion_tokens_details":{"reasoning_tokens":2134}},"tokens_in":664,"tokens_out":2208,"duration_ms":15148,"temperature":1.0,"reasoning_tokens":2134,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:57:15.552649+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have human raters label each real artwork's four-seed set by the same five topology classes (compact, $3{+}1$, $2{+}2$, $2{+}1{+}1$, all-distinct) and compare the protocol's predictions to those labels; agreement near chance, or an outlier AUROC that fails to beat a scalar on real sets, would show the synthetic shapes are easier to separate than real seed distributions.","supporting_citations":[],"review_version":2}