{"id":"ffeb1f64-7bad-4bc4-aeca-8e2924ba04d7","arxiv_id":"2412.09625","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Illusion3D generates 3D objects with multicolor textures that reveal different pictures from different viewpoints, using a 2D text-to-image diffusion model and score-distillation optimization.","lead":"Researchers built a system that draws a 3D object, such as a cube or a reflective cylinder, whose surface reveals one picture from one angle and a different picture from another angle. It uses a text-to-image AI model to steer the texture, so a single object can show a campfire from the front and a butterfly from the side.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Per-view instance counts are not measured: the reported metrics can be high even when a view contains duplicate or blended content, so the central claim of distinct interpretations is not established by the evidence.","rationale":"The reader identified the local-minimum problem as the weakest assumption; I agree but sharpen it. The paper's own failure cases and Section 5 statement show the problem is real. The missing piece is that the evaluation metrics do not measure the defining property of the illusion—per-view uniqueness—so the headline result could be an artifact of aggregate scoring. My proposed test directly counts instances and would settle whether the method produces distinct interpretations or cheats via duplication/blending. I also note the prompt distribution is style-constrained, narrowing the tested domain relative to the abstract's claim. These are reasons to keep the CONDITIONAL verdict and add a reproducibility/validation requirement, not to reject the paper: the qualitative results and ablations are plausible evidence that the method works for the demonstrated cases.","tokens_in":14354,"tokens_out":7610,"duration_ms":83716,"concrete_test":"Run the released or author-provided implementation on the 86 evaluation prompt pairs plus 50 unconstrained prompt pairs (no style prefix). For each generated object, render all target views and use GroundingDINO or human annotation to count instances of each target concept per view. Pre-register a threshold: at least 90% of views must contain exactly one instance of their intended concept, and no view may contain another view's concept. If the method falls below this, the 'different interpretations' claim is not supported for general prompts.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that each target viewpoint yields its intended concept and that concepts do not duplicate or blend across views. The optimization has no such guarantee: Section 5 concedes 'the model can still cheat the optimization criteria,' and Fig. 10 shows duplicate monkey faces and blended content. The quantitative support (Tables 1-3) uses CLIP score, aesthetic scores, and aggregate alignment/concealment; none of these count instances within a view. A view containing two copies of the target object still receives a high CLIP score, so the reported numbers are compatible with the failure modes the paper acknowledges. Furthermore, the 43 prompt pairs in Section 4 are all constructed as 'random painting style + primary object'; Supp. Section D states that prompts without style constraints have lower success rates, so the abstract's 'user-provided text prompts' claim is extrapolated beyond the tested distribution. The load-bearing premise—that scheduled camera jitter, patch denoising, and progressive resolution scaling reliably avoid duplicate/blended local minima—is therefore unverified for the general claim. This does not disprove the method, but it means the central claim is conditional on prompt distribution and on a per-view uniqueness check that the paper does not report.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Illusion3D proposes an optimization-based method for generating 3D multiview illusions from text prompts or reference images. A neural texture field (an InstantNGP hash-encoding MLP) on a fixed 3D shape (cube, sphere, beanbag, Lego, or reflective cylinders/mirrors) is optimized with Variational Score Distillation (VSD) using a pre-trained Stable Diffusion model, with each target viewpoint associated with a distinct prompt or, in the 'Waldo' case, an L2-supervised image view. Three techniques are introduced to stabilize the under-constrained optimization: scheduled camera jittering (Sec. 3.2), random-patch denoising, and progressive render-resolution scaling (Sec. 3.3). The evaluation uses 43 randomly constructed style+object prompt pairs (86 examples on cubes and spheres), comparisons against inverse projection, latent blending, and Burgert et al. [3], systematic ablations, a 40-participant user study, and demonstrations including reflective surfaces, 8-view cubes, and 2048x2048 texture maps.","tokens_in":14609,"tokens_out":13731,"duration_ms":123129,"significance":"The problem is timely and well motivated, and the paper is a credible step from 2D diffusion illusions toward full-color 3D multiview illusions, going beyond shadow/wire art in expressiveness. The method is clearly specified and reproducible: a single hyperparameter set is used for all experiments (Supp. A), prompt pairs are randomly constructed rather than cherry-picked, the three proposed techniques are each ablated, and the reported metrics move consistently with the qualitative story. The authors are also commendably candid about failure cases (Sec. 5, Fig. 10) and about the role of style constraints in prompt selection (Supp. D). The main deficiency is evidentiary rather than conceptual: the metrics do not measure the defining property of an illusion (each view containing its intended concept exactly once, with no duplication or blending), and the tested prompt distribution is narrower than the abstract claims. Both gaps are addressable with additional measurement and a modestly scoped claim, so the contribution is, in my assessment, defensible at the core.","major_comments":[{"comment":"The central claim — that each viewpoint yields its intended concept without duplication or blending — is not measured by the reported quantitative evidence. CLIP score, aesthetic score, and the alignment/concealment scores are all insensitive to the failure modes the paper itself documents: a view containing two monkeys (Fig. 10, row 1) still produces a high CLIP score for the 'monkey' prompt, and a blended view still partially matches both prompt texts. The paper even concedes in Sec. 5 that 'the model can still cheat the optimization criteria,' and Supp. D notes the prompt distribution matters. The user study (Supp. E) asks only about visual appeal and prompt alignment, not about whether a view contains duplicate or mixed primary content. The stress-test concern therefore lands: the reported numbers are compatible with the acknowledged failures. I request (i) a per-view instance-count evaluation, e.g., open-vocabulary detection counting detections of the prompt's object noun per view, and (ii) a reported success rate over the 43 prompt pairs defined as 'no duplicate or blended content in any view.' These are measurement additions, not changes to the method.","section":"Sec. 4 (Tables 1-2), Sec. 5, Fig. 10"},{"comment":"The abstract and introduction claim the method works from 'user-provided text prompts,' but the evaluation distribution is much narrower. All 43 prompt pairs in Sec. 4 are constructed as 'random painting style + primary object,' and Supp. Sec. D explicitly states that prompts without style constraints converge with lower success rates and that 'illusion on real prompts is hard to succeed.' As written, the evidence supports only style-constrained noun+style prompts, not arbitrary prompts. The authors should either restrict the claim to the tested distribution or evaluate on unconstrained or user-provided prompts and report the per-prompt success rate. This matters because the distinction determines whether the paper delivers what its title and abstract promise.","section":"Abstract, Sec. 1, Sec. 4, Supp. Sec. D"},{"comment":"All reported quantitative results are point estimates without error bars or significance tests. The 86 examples come from only 43 prompt pairs (Sec. 4), so the prompt pair is the natural sampling unit and per-prompt variance is computable; I request bootstrap confidence intervals or paired tests for the headline comparisons in Tables 1-2 and for the user-study preference percentages in Table 3. The importance is visible in the Table 2 margins (e.g., Ablation-C vs. Ablation-D: CLIP 0.155 vs 0.165) and in Table 3 (Ablation-C 29.50% vs Ablation-D 39.80%), which are small relative to the 40-participant, 24-comparison study described in Fig. 12. Without variance information, the claim of 'best among all the metrics' (Sec. 4.1) is not fully substantiated.","section":"Sec. 4 (Tables 1-3), Fig. 12"}],"minor_comments":[{"comment":"The sentence 'we leverage 2D diffusion priors from a pre-trained Stable Diffusion model [42] via a Score Distillation Sampling [37] is a common method for utilizing 2D diffusion priors, it often produces over-smoothing and color over-saturation artifacts' is garbled and should be split into two correct sentences; the intended meaning (SDS is a common method, VSD is adopted instead) is clear but the current text reads as two sentences fused.","section":"Sec. 3.1"},{"comment":"Subfigure (a) is cited as '2D flip illusion [42],' but reference [42] is the Stable Diffusion/LDM paper; the cited work for 2D diffusion illusions should be Burgert et al. [3] and/or Geng et al. [8] (which the text itself cites in Sec. 2 for this figure).","section":"Fig. 2 caption"},{"comment":"The abstract claims the method optimizes 'textures and geometry,' but Sec. 3 describes only texture-field optimization; the main experiments use fixed geometry, and the 3D-shape case (Fig. 8) is described only by a caption with no training procedure. Also, the 1024/2048 resolution claims: Eq. (6) specifies a final resolution of b=1024, while the abstract promises 2048x2048 outputs; please document the 2048 setting (schedule endpoints, patch size, budget) or align the claims with the described procedure.","section":"Abstract, Sec. 3, Supp. Sec. C"},{"comment":"In the pairwise preference data, percentages for a pair do not sum to 100 (e.g., Ours 60.61% vs. Burgert et al. 10.14% in Table 3), so it is unclear whether ties were allowed or whether the normalization denominator varies per pairing; please state the exact question format, whether choice was forced, and the denominator used in Eq. (7).","section":"Table 3 / Fig. 12"},{"comment":"Typos and formatting artifacts to clean up: Table 2 caption 'CLIP sore' should be 'CLIP score'; Sec. 4.1 'best among all the matrices' should be 'metrics'; Supp. Sec. A 'we set the training time step for 2000' should be 'to 2000'; Fig. 3 caption contains 'duplicate pattern issue issue'; and the reference list contains stray trailing numbers (e.g., refs [3], [14], [25], [43]) that appear to be leftover build artifacts.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"For the editor: the paper is within scope for a computer-vision venue and is, in my view, a competent contribution whose core optimization story is credible. The two risks most likely to draw criticism are (1) the mismatch between the abstract's 'user-provided text prompts'/'geometry' phrasing and the actual tested distribution (style-constrained noun prompts, fixed geometry in the main experiments), and (2) the absence of any direct measurement of the defining illusion property (per-view uniqueness). Both are fixable, which is why I recommend major revision rather than rejection. The citation oddities in the bibliography suggest a need for a careful author proofread."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's the short version: Illusion3D delivers something new—automatic generation of full-color 3D multiview illusions from text or image prompts on arbitrary shapes, beyond the shadow/wire/2D-plane illusions in prior work. The three techniques (scheduled camera jitter, patch denoising, progressive resolution scaling) are sensible design choices, and the ablations show they each help. The paper is also honest: it explicitly says VSD is not a novelty and concedes in the Limitations that the model can 'cheat the optimization criteria.' I believe the method genuinely works for the demonstrated cases.\n\nNow the soft spots. The evaluation never directly measures the central property: that each target view contains its intended concept exactly once, with no duplicates or blending. CLIP, aesthetic, and alignment/concealment scores can all stay high when a view contains two copies of the object. The paper shows duplicate and blending failure cases but doesn't report how often they occur across the 43 prompt pairs. So the headline claim about distinct per-view interpretations is supported by qualitative examples, not by a quantitative count.\n\nSecond, the abstract says 'user-provided text prompts,' but all experiments use 'random painting style + primary object' pairs. The supplement admits that without style constraints, success rates drop. That's a genuine gap between the claim and the tested distribution.\n\nNo code is released, and the tables lack error bars. The user study has 40 participants, which is small but reasonable for this type of work.\n\nNone of this is fatal. The core idea is novel, the method is clearly specified, and the limitations are openly stated. This is exactly the kind of paper that deserves a serious referee. In revision I'd ask for code, a per-view duplicate/blend count (or a metric that penalizes multiple instances), and either broader prompt testing or a more careful claim in the abstract.","headline":"Solid, honest methods paper that makes 3D multiview illusions from text prompts work in practice, but the central 'one distinct concept per view' claim is not directly measured and the prompt distribution is narrower than the abstract suggests.","tokens_in":15156,"tokens_out":3687,"would_cite":false,"duration_ms":36533,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single texture field can hide several pictures in one 3D object.","keywords":["multiview illusion","3D illusion generation","diffusion priors","score distillation","neural texture field","differentiable rendering","high-resolution texture"],"falsifier":"Render a generated cube illusion and ask a judge (human or CLIP-based) to match each of its three views to its intended prompt; if a substantial fraction of views is not best matched to its own prompt, or if a single face shows duplicated content visible from two corners, the central claim of per-view fidelity fails. Concretely, run the released pipeline on 50 random prompt pairs and count views where the intended concept is not the top match.","tokens_in":14143,"feed_emoji":"🎭","tokens_out":6822,"duration_ms":58230,"temperature":0.7,"pith_summary":"This paper claims that one 3D object can carry several different, detailed pictures on its surface and reveal them one at a time as the viewer changes angle. The method assigns each camera viewpoint its own text prompt or reference image, then optimizes a shared neural texture field so that every rendered view matches its own prompt. The optimization uses a pre-trained text-to-image diffusion prior through Variational Score Distillation, and the paper's contribution is showing that three stabilizers—scheduled camera jitter, patch-wise denoising, and progressive resolution scaling—make this under-constrained process converge to clean, non-duplicated illusions. If it works, it turns illusion-making into a prompt-driven process that applies to cubes, spheres, soft shapes, and reflective mirrors at texture resolutions up to 2048x2048.","feed_headline":"One object, many views: diffusion-painted 3D illusions","feed_subtitle":"Text prompts become textures on cubes, spheres, and mirrors, with each viewpoint showing its own image.","key_machinery":"The core machinery is Variational Score Distillation (VSD) applied through a multi-resolution hash-encoding MLP texture field (Instant-NGP style), rendered with differentiable rasterization and optimized against a pre-trained Stable Diffusion model. Three stabilization techniques carry the argument: scheduled camera jitter (adding growing Gaussian noise to rotation, translation, and field of view), patch-wise denoising (optimizing random 512x512 patches of a larger render so the low-resolution diffusion prior can guide high-resolution output), and progressive resolution scaling (ramping the render size from 512 up to 1024 or 2048 so the main object is centered and duplication is suppressed). Together they prevent the VSD loss from settling into local optima that duplicate or blend the per-view concepts.","core_discovery":"The central claim is that 3D multiview illusions can be generated by optimizing a single texture field on a given 3D shape with 2D diffusion priors, rather than by hand-crafting shadow, wire, or reflective art. Using the Variational Score Distillation gradient from ProlificDreamer, the method distills a separate Stable Diffusion prompt into each assigned viewpoint of a multi-resolution hash-encoding MLP texture field. The paper argues that naive VSD optimization gets stuck in local minima—duplicated concepts on one face, blended content across views, and VAE blind-spot artifacts—and that its three techniques fix these failures: linearly scheduled camera jitter that grows over training, random 512x512 patch denoising of a progressively enlarged render, and a sigmoid schedule for render resolution. Results are shown for cube, sphere, beanbag, Lego, reflective cylinder, and curved-mirror setups, with three to eight views per object, and the method also accepts an image input for one view via L2 supervision.","pith_inferences":["The VAE blind-spot explanation for artifacts suggests that any latent-diffusion 3D optimization, not just illusions, could benefit from small amounts of scheduled camera jitter; this is a testable transfer to text-to-3D without illusion requirements.","The duplication and blending failures are likely inherent to optimizing a single texture field for multiple full-image objectives, so more principled fixes, such as spatially explicit view partition masks or equivariance constraints on the VAE encoder, might replace the heuristic schedules.","If the method extends beyond static textures to geometry optimization, then view-dependent shape changes could create even stronger illusions, since perspective distortion itself would contribute to hiding.","The paper's success with reflective surfaces suggests a broader recipe: any optical element that maps different rays to different image content, such as mirrors, lenses, or water surfaces, could serve as the canvas for a diffusion-guided illusion."],"forward_implications":["A user can go from two text prompts to a physical 3D object whose front shows one concept and side shows another, with no manual image editing.","High-resolution texture output (1024x1024 and 2048x2048) means the illusions survive being printed, wrapped, or displayed at human scale rather than only on screen.","Reflective surfaces can host multiple simultaneous illusions, with the paper's two-cylinder and mirror setups giving three views, which traditional hand-crafted reflective art does not achieve.","Eight-view cube results show the approach can scale a single object with a sequence of looks, so a sculpture or product could encode a story as the viewer walks around it.","Image-conditioned views (the Waldo example) indicate the same pipeline can embed a specific picture into one view, enabling personalized objects that hide private content in plain sight."],"supporting_citations":[{"why":"Supplies the Variational Score Distillation loss and LoRA update that drive texture optimization.","marker":"[50]"},{"why":"Provides the pre-trained Stable Diffusion model that serves as the text-to-image prior for each view.","marker":"[42]"},{"why":"Defines the multi-resolution hash-encoding MLP used to parameterize the texture field.","marker":"[26]"},{"why":"The 2D optimization-based illusion method this work extends and the main baseline compared with Dream Target loss.","marker":"[3]"},{"why":"Introduces score distillation sampling, the predecessor loss that motivates the choice of VSD.","marker":"[37]"},{"why":"Represents the 2D joint-denoising illusion paradigm that does not transfer to 3D, motivating the optimization-based approach.","marker":"[8]"}],"fun_headline_variants":["Diffusion-painted 3D objects reveal different images per angle","One shape, many views: diffusion-optimized textures create 3D illusions","3D multiview illusions crafted from text prompts with diffusion priors","Text prompts become different images on different sides of one 3D object","Diffusion priors paint 3D shapes to show a distinct image per viewpoint"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Each viewpoint's diffusion loss can be satisfied simultaneously on one shared surface without one view's content leaking into another; if the optimization settles into a state where the same concept appears on multiple faces or a view blends two concepts, the claimed illusion is broken.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion-painted 3D objects reveal different images per angle","One shape, many views: diffusion-optimized textures create 3D illusions","3D multiview illusions crafted from text prompts with diffusion priors","Text prompts become different images on different sides of one 3D object","Diffusion priors paint 3D shapes to show a distinct image per viewpoint"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000657,"raw_usage":{"total_tokens":2999,"prompt_tokens":926,"completion_tokens":2073,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":1975}},"tokens_in":542,"tokens_out":2073,"duration_ms":12474,"temperature":1.0,"reasoning_tokens":1975,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T16:50:55.008724+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Render a generated cube illusion and ask a judge (human or CLIP-based) to match each of its three views to its intended prompt; if a substantial fraction of views is not best matched to its own prompt, or if a single face shows duplicated content visible from two corners, the central claim of per-view fidelity fails. Concretely, run the released pipeline on 50 random prompt pairs and count views where the intended concept is not the top match.","supporting_citations":[{"cited_title":"Prolificdreamer: High-fidelity and diverse text-to-3d generation with variational score distilla- tion","cited_arxiv_id":null,"evidence_quote":"Supplies the Variational Score Distillation loss and LoRA update that drive texture optimization."},{"cited_title":"Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer","cited_arxiv_id":null,"evidence_quote":"Provides the pre-trained Stable Diffusion model that serves as the text-to-image prior for each view."},{"cited_title":"Instant neural graphics primitives with a multires- olution hash encoding","cited_arxiv_id":null,"evidence_quote":"Defines the multi-resolution hash-encoding MLP used to parameterize the texture field."},{"cited_title":"Diffusion illusions: Hiding images in plain sight","cited_arxiv_id":null,"evidence_quote":"The 2D optimization-based illusion method this work extends and the main baseline compared with Dream Target loss."}],"review_version":1}