{"total":14,"items":[{"citing_arxiv_id":"2607.05373","ref_index":16,"ref_count":1,"confidence":0.98,"is_internal_anchor":true,"paper_title":"PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space","primary_cat":"cs.CV","submitted_at":"2026-07-06T17:51:00+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A single pixel-space diffusion model jointly performs 3D scene reconstruction and generation by supervising flow matching on rendered multi-view images, matching SOTA reconstruction and outperforming latent-space generation.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.30045","ref_index":7,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Walking in the Implicit: Interactive World Exploration via Neural Scene Representation","primary_cat":"cs.CV","submitted_at":"2026-06-29T09:37:33+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"NeuWorld uses a transformer VAE to learn compact Neural Implicit Scenes from sparse posed frames and a diffusion transformer to evolve them conditioned on camera trajectories for consistent interactive exploration.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.27071","ref_index":19,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"PanoImager: Geometry-Guided Novel View Synthesis and Reconstruction from Sparse Panoramic Views","primary_cat":"cs.CV","submitted_at":"2026-06-25T14:15:09+00:00","verdict":"CONDITIONAL","verdict_confidence":"HIGH","novelty_score":5.0,"formal_verification":"none","one_line_summary":"An SfM-free pipeline that turns a few sparse panoramas into a more stable 3D Gaussian map by combining feed-forward pose/depth priors, diffusion view completion, and depth-constrained optimization.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.02479","ref_index":35,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Retrieve What's Missing: Coverage-Maximizing Retrieval for Consistent Long Video Generation","primary_cat":"cs.CV","submitted_at":"2026-06-01T16:49:58+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"COVRAG improves long-horizon geometric consistency in autoregressive video generation via coverage-maximizing retrieval on lightweight depth-based 3D memory evidence.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.25449","ref_index":41,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Pantheon360: Taming Digital Twin Generation via 3D-Aware 360{\\deg} Video Diffusion","primary_cat":"cs.CV","submitted_at":"2026-05-25T06:00:01+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Pantheon360 introduces a controllable 360° video diffusion framework that uses an explicit 3D cache from sparse inputs to enforce geometric consistency for digital twin generation.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.23888","ref_index":49,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"GenRecon: Bridging Generative Priors for Multi-View 3D Scene Reconstruction","primary_cat":"cs.CV","submitted_at":"2026-05-22T17:49:59+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"GenRecon lifts object-level generative priors to scene-scale reconstruction by chunking scenes and using projection-based conditioning on multi-view features, claiming 16% better results than prior methods.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.23555","ref_index":38,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Generator-Refiner-Examiner: A Tri-Module Data Augmentation Framework for 3D Human Avatar Learning from Monocular Videos","primary_cat":"cs.CV","submitted_at":"2026-05-22T12:20:46+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"TrioMan is a tri-module data augmentation framework using a Generator for pose/camera perturbations, a Refiner with one-step diffusion, and an Examiner with dual-branch attention to improve 3D avatar learning from monocular videos, claiming better results than prior methods on two benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.18052","ref_index":137,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Efficient 3D Content Reconstruction and Generation","primary_cat":"cs.CV","submitted_at":"2026-05-18T08:41:10+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"Presents Instant3D for rapid text/image-to-3D generation via multi-view diffusion plus feed-forward reconstruction, and FastMap for 10x faster structure-from-motion with comparable accuracy.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"followed by textured point clouds) guided by a 2D diffusion prior, while Magic123[186] adopts a coarse-to-fine strategy with joint 2D and 3D diffusion guidance and a trade-off parameter between the priors. ATT3D[143] shifts toward amortized inference by training a text-conditioned model to produce 3D objects without per-prompt optimization, trading per- instance fitting for a learned feed-forward generator. One-2-3-45[137] generates multi-view images with Zero123 and lifts them to a 360-degree mesh using an SDF-based generalizable surface reconstructor in a single feed-forward pass. LRM[94] predicts a triplane represen- tation from a single image using a large transformer-based reconstruction model trained on 8 large-scale multi-view data. GS-LRM[308] extends this direction to 3D Gaussian splatting"},{"citing_arxiv_id":"2604.11797","ref_index":15,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"SyncFix: Fixing 3D Reconstructions via Multi-View Synchronization","primary_cat":"cs.CV","submitted_at":"2026-04-13T17:58:06+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"SyncFix improves 3D reconstructions by synchronizing multi-view latent representations in a diffusion refinement process, generalizing from pair-wise training to arbitrary view counts at inference.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2602.15355","ref_index":47,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"DAV-GSWT: Diffusion-Active-View Sampling for Data-Efficient Gaussian Splatting Wang Tiles","primary_cat":"cs.CV","submitted_at":"2026-02-17T04:47:39+00:00","verdict":"REJECT","verdict_confidence":"MODERATE","novelty_score":4.0,"formal_verification":"none","one_line_summary":"DAV-GSWT selects views by diffusion-model uncertainty and hallucinates missing structure so Gaussian Splatting Wang Tiles can be made from sparse captures.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2512.14614","ref_index":42,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling","primary_cat":"cs.CV","submitted_at":"2025-12-16T17:22:46+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A real-time video diffusion world model that uses dual action control, reframed position encodings, and context-aligned distillation to keep generated environments consistent over hundreds of frames.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2511.00503","ref_index":50,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Diff4Splat: Controllable 4D Scene Generation with Latent Dynamic Reconstruction Models","primary_cat":"cs.CV","submitted_at":"2025-11-01T11:16:25+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A feed-forward video latent transformer that predicts time-varying 3D Gaussian primitives from one image to produce controllable 4D scenes with appearance, geometry, and motion.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2507.00990","ref_index":74,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations","primary_cat":"cs.RO","submitted_at":"2025-07-01T17:39:59+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"RIGVid shows that filtered AI-generated videos can serve as effective supervision for complex robotic manipulation tasks without any real demonstrations.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"and task description, be used as the sole source of super- vision for robotic manipulation? Recently released mod- els like SORA [16] and Kling [1] have demonstrated im- pressive capabilities in producing realistic-seeming videos from language and image inputs. At the same time, it has been shown that such videos can suffer from distorted ob- ject geometries [74, 132], physically implausible interac- tions [82, 127], and unrealistic scene dynamics [11, 39]. Consequently, while the idea of synthesizing video demon- strations is enticing, its usefulness in the robotics setting is yet to be convincingly established. Prior work incor- porating video generation into robotics typically relies on additional supervision, such as task-specific training [30]"},{"citing_arxiv_id":"2505.21996","ref_index":54,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"VRAG: Learning World Models for Interactive Video Generation","primary_cat":"cs.CV","submitted_at":"2025-05-28T05:55:44+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"VRAG improves long-horizon interactive video generation by conditioning autoregressive diffusion on retrieved historical frames and explicit global state, outperforming long-context baselines on the tested Minecraft and RealEstate10K benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}