{"id":"e86ca01a-46e6-4d8e-b3e9-4dab7293e2af","arxiv_id":"2607.05373","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":7,"one_line_summary":"A single pixel-space diffusion model jointly performs 3D scene reconstruction and generation by supervising flow matching on rendered multi-view images, matching SOTA reconstruction and outperforming latent-space generation.","lead":"PixWorld trains a single diffusion model directly in pixel space to both reconstruct and generate 3D scenes, eliminating the VAE/RAE bottleneck used by prior latent-space methods. A unified model that handles both tasks at competitive quality could simplify 3D pipelines and reduce training overhead.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"Evaluation protocol does not specify how NVS metrics are computed across methods with different output resolutions; PixWorld runs at 336×448 while video-diffusion baselines operate at higher resolution (Appendix E), potentially confounding the generation comparisons in Tables 2–3.","rationale":"The reader identified narrow margins and resolution scalability as concerns but did not connect them to a potential evaluation-protocol confound in the current results. My concern is more specific: the unspecified evaluation resolution could systematically bias the generation comparisons where the 'consistent outperformance' claim is made. However, this concern is about clarification rather than refutation. The reconstruction claim ('matches SOTA') is likely unaffected since all baselines in Table 1 are similar pixel-aligned Gaussian methods. The generation claim against Gen3R (same NFE, similar unified framework, +1.29 dB on RE10K 1-view) has enough margin that a resolution confound would need to be large to overturn it. The ablation on the geometry perception loss (Table 5, +1.13 dB PSNR) is internally controlled and unaffected by this concern. No code release is mentioned, which limits independent verification, but the experimental design is sound in structure. The verdict remains CONDITIONAL: the contribution is solid but would be strengthened by specifying the evaluation resolution protocol, reporting statistical significance, and releasing code.","tokens_in":25338,"tokens_out":7271,"duration_ms":363635,"concrete_test":"Re-evaluate all methods in Tables 2–3 at a single common resolution (e.g., 336×448) by downsampling all method outputs and ground-truth targets to that resolution before computing PSNR/SSIM/LPIPS. If PixWorld's margin over Gen3R shifts by more than 0.3 dB in either direction compared to the reported numbers, the resolution protocol is a material confound and the 'consistent outperformance' claim needs qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of 'consistently outperforming prior latent-space generation methods' rests on PSNR/SSIM/LPIPS comparisons in Tables 2–3 against video-diffusion baselines (FlashWorld, Gen3C, Gen3R). Appendix E explicitly states: 'PixWorld currently runs at a lower output resolution than the video-diffusion baselines.' The paper never specifies the evaluation resolution or how metrics are computed when methods have different native output sizes. If NVS metrics are computed by resizing all outputs to a common resolution, the interpolation direction matters: downsampling higher-resolution baseline outputs to 336×448 would tend to increase their PSNR (smoothing high-frequency mismatches), while upsampling PixWorld outputs would decrease its PSNR. The net confound direction is unknown because the protocol is unspecified. This matters most for generation comparisons where margins are modest — e.g., +0.75 dB over Gen3R on DL3DV 1-view (Table 2) and +0.19 dB over LVSM on RE10K 2-view (Table 3, where LVSM actually edges PixWorld on PSNR at 23.61 vs 23.54). The reconstruction comparison (Table 1) is less affected because all baselines there are pixel-aligned Gaussian methods likely operating at similar resolutions. The concern is not that the results are wrong, but that an unspecified resolution protocol could shift rankings in the generation setting where the 'consistent outperformance' claim is made.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"The paper introduces PixWorld, a unified pixel-space diffusion framework for 3D scene reconstruction and generation. The core idea is to perform flow matching directly on rendered images from a pixel-aligned 3D Gaussian Splatting (3DGS) representation, eliminating the intermediate VAE/RAE latent stage used by prior latent-space 3D generation methods. The model uses a two-stream DiT that processes clean (reconstruction) and noisy (generation) views jointly, supervised by a flow-matching loss on rendered images, a depth loss, and a novel geometry perception loss that aligns rendered and ground-truth views in the feature space of a frozen 3D foundation model (pi3). Experiments span four benchmarks (RealEstate10K, DL3DV-10K, WorldScore) across reconstruction (4/8-view) and generation (1/2-view) settings, with disaggregated per-configuration results in the appendix.","tokens_in":26140,"tokens_out":1450,"duration_ms":293961,"significance":"The central claim — that pixel-space diffusion with differentiable 3DGS rendering can unify reconstruction and generation while matching or exceeding latent-space methods — is well-motivated and supported by a substantial experimental effort. The geometry perception loss (Sec. 3.3) is a sensible contribution that addresses a real gap in 2D photometric supervision. The ablation (Tab. 5) shows a clear 1.13 dB PSNR drop without it. The paper provides a detailed architecture breakdown (Tab. 6), inference-speed comparison (Tab. 9), and per-configuration disaggregation (Tabs. 7–8), which strengthen reproducibility. The claim of removing the VAE/RAE bottleneck is architecturally sound and the differentiable rendering pipeline is internally consistent.","major_comments":[{"comment":"§4.2 and Appendix E: The evaluation protocol for NVS metrics (PSNR/SSIM/LPIPS) does not specify how comparisons are handled when methods have different native output resolutions. Appendix E states that PixWorld runs at a lower output resolution (336×448) than video-diffusion baselines. The paper does not state whether all outputs are resized to a common resolution before metric computation, and if so, in which direction (up vs. down). This matters for the generation comparisons in Tables 2–3, where some margins are modest (e.g., +0.75 dB over Gen3R on DL3DV 1-view in Table 2; LVSM edges PixWorld on PSNR at 23.61 vs. 23.54 on RE10K 2-view in Table 3). The reconstruction comparison (Table 1) is less affected since baselines there are pixel-aligned Gaussian methods likely operating at similar resolutions. The concern is not that results are wrong, but that an unspecified protocol could conf","section":null},{"comment":"Abstract and §4.4: The claim that PixWorld 'consistently outperforms prior latent-space generation methods' is strong but partially undercut by the resolution mismatch noted above and by the fact that LVSM (a deterministic regression model, not a latent-space diffusion method) is competitive on PSNR/SSIM in the 2-view setting (Table 3). The paper should either (a) narrow the claim to specify that it outperforms latent-space *diffusion* methods (excluding LVSM from the comparison group for this claim), or (b) clarify that 'consistently' refers to the aggregate across NVS, generation quality, and camera control metrics rather than PSNR alone. As stated, 'consistently outperforms' is not strictly supported by Table 3 where LVSM wins on PSNR and SSIM for RE10K 2-view.","section":null}],"minor_comments":[{"comment":"§3.2, Eq. (6): The LPIPS term is defined as L_lpips = 1[t > t_th] * (1/N) * sum LPIPS(bar_I_n, I_n), summing over all N views including clean ones. However, the text says 'perceptual supervision is unreliable when the noisy input is too close to pure noise,' which explains gating by t > t_th for noisy views but does not explain why clean views (which have no noise) are also gated. Please clarify whether the indicator is intended to apply to all views or only noisy ones.","section":null},{"comment":"§4.1: The loss weights (lambda_depth=1.0, lambda_lpips=lambda_geo=0.1, t_th=0.3) are stated but no sensitivity analysis is provided beyond the single ablation on the geometry loss (Tab. 5). A brief note on robustness to these choices, or at least acknowledgment that they were selected by informal search, would strengthen the paper.","section":null},{"comment":"Table 5: The ablation is conducted on a 10K-sequence subset for 30K steps, which is a reduced setting compared to the full training (67K scenes, 200K steps). While the controlled comparison is valid, the paper should note whether the relative effect of the geometry loss is expected to scale similarly at full training, or whether the ablation setting was chosen for compute reasons.","section":null},{"comment":"§3.3, Eq. (9): The geometry perception loss uses cosine distance in the feature space of a frozen 3D foundation model. The choice of pi3 is mentioned in §4.1, and VGGT is listed as an alternative in §3.3. It would help to briefly note whether results with VGGT were tried and how the choice of critic affects results.","section":null},{"comment":"Appendix E, Table 9: The inference-speed comparison notes that PixWorld runs at lower resolution than baselines, which reduces per-step compute. Stating the exact resolution difference (e.g., 'baselines operate at X×Y') would improve transparency.","section":null},{"comment":"Figure 2: The '×?' in 'Shared DiT Block × ?' should be replaced with the actual count (24, per Appendix B).","section":null},{"comment":"The paper uses '1-view' and '2-view' terminology throughout. On first use in §4.2, a brief parenthetical clarifying that '1-view' means single-image conditioning and '2-view' means two-image conditioning would improve readability for readers unfamiliar with the convention.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The stress-test concern about the unspecified evaluation resolution protocol is the most substantive issue. It is not necessarily a fatal flaw — the reconstruction results (Table 1) are less affected, and the generation margins on camera control and VBench metrics are large enough that resolution alone is unlikely to explain them. However, the NVS PSNR margins in the generation setting are small enough that the protocol matters for the 'consistently outperforms' claim. I rate this as minor revision because the fix is straightforward: specify the protocol and, if needed, add a note acknowledging the resolution mismatch as a limitation. The paper is otherwise thorough and the core architectural contribution is sound."},"author_rebuttal":{"model":"glm-5.2","summary":"We thank the referee for the careful reading and constructive feedback. We address both major comments below.","responses":[{"response":"The referee is correct that the current manuscript does not specify the resolution protocol for NVS metric computation. We will add this to §4.2. Concretely: all metrics are computed at the ground-truth resolution of each test scene. For RealEstate10K and DL3DV-10K, the ground-truth target views are at 336×448 (the native resolution of both datasets' test clips as we process them). PixWorld renders directly at this resolution. For video-diffusion baselines that produce higher-resolution outputs, we bilinearly downsample their generated frames to 336×448 before computing PSNR/SSIM/LPIPS against the ground-truth targets. This means baselines are evaluated at a resolution equal to or lower than their native output, which if anything favors them (downsampling can slightly improve PSNR by suppressing high-frequency noise). We will state this explicitly in the revised §4.2 and note in Appendix E that the resolution difference does not advantage PixWorld. We agree this was an omission that should be corrected for reproducibility.","revision_made":"yes","referee_comment":"§4.2 and Appendix E: The evaluation protocol for NVS metrics (PSNR/SSIM/LPIPS) does not specify how comparisons are handled when methods have different native output resolutions. Appendix E states that PixWorld runs at a lower output resolution (336×448) than video-diffusion baselines. The paper does not state whether all outputs are resized to a common resolution before metric computation, and if so, in which direction (up vs. down)."},{"response":"The referee raises a fair point. We note that LVSM is a deterministic regression model, not a latent-space generation method, so it already falls outside the comparison group for the claim as literally stated ('prior latent-space generation methods'). However, we agree the wording is ambiguous because LVSM appears in the same tables and a reader could reasonably interpret 'consistently outperforms' as applying to all baselines shown. We will adopt the referee's option (b) and revise the abstract and §4.4 to clarify that 'consistently' refers to the aggregate across NVS, generation quality, and camera control metrics — not PSNR alone. We will also explicitly note in §4.4 that LVSM, as a deterministic regression model rather than a generative method, is competitive on raw PSNR/SSIM in the 2-view setting (gap <0.07 dB on RealEstate10K) but trails PixWorld on LPIPS, all generation-quality metrics, and camera control AUC. This makes the scope of the claim precise without overstating it.","revision_made":"yes","referee_comment":"Abstract and §4.4: The claim that PixWorld 'consistently outperforms prior latent-space generation methods' is strong but partially undercut by the resolution mismatch noted above and by the fact that LVSM (a deterministic regression model, not a latent-space diffusion method) is competitive on PSNR/SSIM in the 2-view setting (Table 3). The paper should either (a) narrow the claim to specify that it outperforms latent-space diffusion methods (excluding LVSM from the comparison group for this claim), or (b) clarify that 'consistently' refers to the aggregate across NVS, generation quality, and camera control metrics rather than PSNR alone."}],"tokens_in":25028,"tokens_out":1226,"duration_ms":113752,"standing_objections":[]},"desk_editor":{"model":"glm-5.2","letter":"Here's my read on PixWorld (arXiv:2607.05373). The core idea is genuinely new and well-executed: a single pixel-space flow-matching model that handles both 3D reconstruction and generation by partitioning views into clean/noisy subsets, supervising a pixel-aligned 3DGS representation directly through differentiable rendering. No VAE/RAE bottleneck. The geometry perception loss — using a frozen π3 model as a structural critic — is a nice touch and the ablation shows a real 1.13 dB PSNR drop without it. The reconstruction results (Table 1) are strong: PixWorld beats YoNoSplat and DepthSplat across both datasets at 4 and 8 views, and these baselines operate at comparable resolutions, so that comparison is clean. The architecture is well-specified (24-layer MMDiT, 1.04B params, full per-component budget in Table 6), and the disaggregated per-configuration results in Appendix C are thorough and show the ranking holds across harder/easier splits. Credit where it's due: this is a clean, well-motivated framework with real engineering behind it. The stress-test concern about evaluation resolution is the main soft spot, and I think it partially lands. Appendix E explicitly states PixWorld runs at lower resolution than video-diffusion baselines, but the paper never specifies how NVS metrics (PSNR/SSIM/LPIPS in Tables 2–3) are computed when methods have different native output sizes. If baselines are downsampled to 336×448 before metric computation, that would inflate their PSNR (smoothing high-frequency error), which would actually favor PixWorld's claim. If PixWorld outputs are upsampled, that hurts PixWorld. The net direction is genuinely unknown because the protocol is unspecified. This matters most for generation comparisons where margins are modest — +0.75 dB over Gen3R on DL3DV 1-view, and LVSM actually edges PixWorld on 2-view RE10K PSNR (23.61 vs 23.54). The 'consistently outperforms' language is slightly oversold given these narrow margins and the protocol gap. The reader's other concerns are fair but minor: only one ablation (geometry loss), no code release mentioned, no confidence intervals, and seven hand-tuned hyperparameters with no sensitivity analysis. These are standard revision items, not structural problems. The central argument — that pixel-space diffusion can unify these tasks and match or beat latent-space alternatives — holds up on the reconstruction side and is plausible on the generation side pending protocol clarification. This paper is for researchers working on feed-forward 3D scene modeling, novel view synthesis, and unified generation-reconstruction frameworks. It deserves a serious referee who can push for the evaluation protocol details and broader ablations. I'd recommend conditional accept pending: (1) explicit specification of how metrics are computed across different output resolutions, (2) at least one additional ablation on loss weight sensitivity, and (3) a code/model release commitment.","headline":"Pixel-space diffusion for unified 3D scene reconstruction and generation: solid contribution, but evaluation protocol needs clarification before the 'consistent outperformance' claim fully lands.","tokens_in":26380,"tokens_out":705,"would_cite":true,"duration_ms":249808,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"Pixel-space diffusion unifies 3D scene generation and reconstruction","keywords":[],"falsifier":"If pixel-space diffusion at higher resolutions (needed for fine-grained texture) requires disproportionately more compute than latent-space alternatives, or if the geometry perception loss provides no measurable benefit beyond what multi-view photometric supervision already captures, the central claim of pixel-space superiority would weaken.","tokens_in":25547,"feed_emoji":"🖼️","tokens_out":965,"duration_ms":214716,"temperature":0.7,"pith_summary":"The paper argues that 3D scene generation and reconstruction, long handled by separate paradigms, can be unified in a single model by running diffusion directly in pixel space rather than in a compressed latent space. The central mechanism is a two-stream diffusion transformer that partitions multi-view images into clean views (driving reconstruction) and noisy views (driving generation), both decoded into a pixel-aligned 3D Gaussian representation. Because the diffusion objective is supervised on rendered images via differentiable rendering, the training signal aligns directly with 3D scene fidelity rather than with targets in an intermediate latent encoding. The authors further introduce a geometry perception loss that uses a frozen 3D foundation model as a structural critic, penalizing geometric inconsistencies that 2D photometric losses miss. The paper claims this approach consistently outperforms latent-space generation methods and matches state-of-the-art reconstruction methods, suggesting that removing the VAE/RAE bottleneck is a viable path to higher-quality 3D scene modeling.","feed_headline":"Pixel-space diffusion unifies 3D scene generation and reconstruction","feed_subtitle":"One model, no VAE bottleneck: PixWorld supervises 3D Gaussians directly through rendered images, beating latent-space generation andmatching","key_machinery":"Two mechanisms carry the argument. First, the pixel-space flow matching objective is defined on rendered multi-view images rather than latent codes, so gradients flow back through the differentiable 3D Gaussian renderer directly into the 3D representation. Second, the geometry perception loss extracts features from both rendered and ground-truth views using a frozen 3D foundation model (pi3), then minimizes the cosine distance between them, providing cross-view 3D structural supervision that 2D image-level losses cannot supply.","core_discovery":"The key finding is that a single pixel-space diffusion model, supervised through differentiable 3D Gaussian rendering and augmented with a geometry-aware feature loss from a pretrained 3D foundation model, can jointly handle 3D scene reconstruction and generation at quality matching or exceeding specialized methods. The unification works by partitioning multi-view inputs into clean and noisy subsets within one forward pass, so the same model reconstructs observed geometry and generates unobserved content without a latent autoencoder stage.","pith_inferences":["The 336x448 training resolution is a significant constraint; whether pixel-space diffusion can match latent-space efficiency at the higher resolutions needed for production-quality textures remains an open question that the paper acknowledges but does not resolve.","The geometry perception loss depends on the quality and coverage of the frozen 3D foundation model; if that model has systematic geometric blind spots (e.g., thin structures, reflective surfaces), those errors would propagate into PixWorld's learned geometry.","The reliance on pseudo-depth labels from DA3 for depth supervision introduces a dependency on an external model whose accuracy on diverse scenes directly affects PixWorld's geometric grounding.","The 15-second inference time at 100 NFE is competitive but the lower output resolution compared to video-diffusion baselines means the speed comparison is not strictly like-for-like; distilled video models may close the gap further."],"forward_implications":["If pixel-space diffusion scales, the VAE/RAE pretraining stage required by latent-space 3D methods becomes unnecessary, simplifying training pipelines and removing a source of information loss.","The clean/noisy view partitioning scheme suggests a continuum between reconstruction and generation, where the same model can interpolate between pure observation and pure synthesis based on how many views are clean.","The geometry perception loss pattern, using a frozen 3D foundation model as a differentiable structural critic, could be adopted by other 3D generation pipelines that currently rely only on photometric supervision.","If the approach generalizes beyond indoor/real-estate scenes, it could serve as a unified backbone for embodied AI applications requiring both scene understanding and scene completion."],"fun_headline_variants":["One pixel-space diffusion model handles 3D reconstruction and generation","PixWorld unifies 3D scene reconstruction and generation without latent autoencoders","Pixel-space diffusion matches specialized methods for 3D scene generation","Direct pixel supervision unifies 3D scene generation and reconstruction","PixWorld joins 3D reconstruction and generation in one diffusion model"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The approach depends on pixel-space diffusion at 336x448 resolution providing sufficient gradient signal through differentiable rendering to train a 1B-parameter model from scratch; if pixel-space diffusion does not scale to higher resolutions without prohibitive compute, the practical advantage over latent-space methods diminishes.","fun_headline_variants_meta":{"raw":{"variants":["One pixel-space diffusion model handles 3D reconstruction and generation","PixWorld unifies 3D scene reconstruction and generation without latent autoencoders","Pixel-space diffusion matches specialized methods for 3D scene generation","Direct pixel supervision unifies 3D scene generation and reconstruction","PixWorld joins 3D reconstruction and generation in one diffusion model"]},"model":"glm-5.2","effort":"high","cost_usd":0.0,"raw_usage":{"total_tokens":632,"prompt_tokens":542,"completion_tokens":90,"prompt_tokens_details":null},"tokens_in":542,"tokens_out":90,"duration_ms":50238,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-07T14:05:43.431816+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"If pixel-space diffusion at higher resolutions (needed for fine-grained texture) requires disproportionately more compute than latent-space alternatives, or if the geometry perception loss provides no measurable benefit beyond what multi-view photometric supervision already captures, the central claim of pixel-space superiority would weaken.","supporting_citations":[],"review_version":1}