{"total":15,"items":[{"citing_arxiv_id":"2606.29513","ref_index":29,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Scenes as Objects, Not Primitives: Instance-Structured 3D Tokenization from Unposed Views","primary_cat":"cs.CV","submitted_at":"2026-06-28T17:21:45+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"A feed-forward framework learns instance-structured 3D token groups from unposed multi-view images via differentiable rendering, enabling native object-level segmentation, editing, and retrieval without 3D supervision.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.25278","ref_index":59,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Heterogeneous and Adept Snapshot Distillation for 3D Semantic Segmentation","primary_cat":"cs.CV","submitted_at":"2026-06-24T01:31:22+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"HAS-KD combines information-oriented heterogeneous distillation from multi-modal models with adept snapshot distillation from training checkpoints to reach SOTA 3D semantic segmentation on ScanNetV2 and S3DIS without added inference burden.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.19316","ref_index":86,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"NeuMesh++: Towards Versatile and Efficient Volumetric Editing with Disentangled Neural Mesh-based Implicit Field","primary_cat":"cs.CV","submitted_at":"2026-06-17T17:39:21+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"A disentangled mesh-vertex neural radiance field enables mesh-guided geometry edits, texture swap/fill/paint, and semantic-guided edits with claimed efficiency and quality gains.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.08440","ref_index":18,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"GraspFoM: Towards Reconstruction-Driven Robotic Grasping with 3D Foundation Priors","primary_cat":"cs.RO","submitted_at":"2026-06-07T03:37:55+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"GraspFoM creates a shared 3D latent from SAM3D priors, adds an anchor-initialized diffuser for multimodal grasps, and uses reconstruction-aware scoring plus residual updates to jointly achieve SOTA reconstruction and grasping with few extra parameters.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2606.05975","ref_index":8,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"T-FunS3D: Task-Driven Hierarchical Open-Vocabulary 3D Functionality Segmentation","primary_cat":"cs.CV","submitted_at":"2026-06-04T10:16:39+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"T-FunS3D is a task-driven hierarchical method for open-vocabulary 3D functionality segmentation that constructs an open-vocabulary scene graph and applies vision-language models to achieve comparable accuracy with faster runtime and lower memory on SceneFun3D.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.29505","ref_index":38,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"ESAM++: Efficient Online 3D Perception on the Edge","primary_cat":"cs.CV","submitted_at":"2026-05-28T07:29:02+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"ESAM++ introduces a 3D Sparse Feature Pyramid Network for efficient online 3D scene perception on edge devices, claiming competitive accuracy with up to 3x faster inference and 2x smaller model size than ESAM on four benchmarks.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.25901","ref_index":17,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"AgentGrounder: Zero-Shot 3D Visual Pointcloud Grounding using Multimodal Language Models","primary_cat":"cs.CV","submitted_at":"2026-05-25T14:29:04+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"AgentGrounder performs zero-shot 3D visual grounding on colored point clouds via an offline object lookup table and an online agent that selectively retrieves, scores geometrically, and renders images on demand, reporting gains over SeeGround on ScanRefer and Nr3D.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.16901","ref_index":23,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"CAR-SAM: Cross-Attention Reconstruction for Post-Training Quantization of the Segment Anything Model","primary_cat":"cs.CV","submitted_at":"2026-05-16T09:25:23+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"CAR-SAM introduces MatMul-Aware Compensation and Joint Cross-Attention Reconstruction to enable stable 4-bit post-training quantization of SAM, outperforming prior PTQ methods by 14.6% mAP on SAM-B and 6.6% on SAM-L.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.08293","ref_index":12,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Distill, Diffuse, Segment: Unsupervised 3D Semantic Segmentation for Autonomous Driving Based on Multi-Level Distillation and Graph Diffusion","primary_cat":"cs.CV","submitted_at":"2026-05-08T09:39:59+00:00","verdict":"CONDITIONAL","verdict_confidence":"MODERATE","novelty_score":5.0,"formal_verification":"none","one_line_summary":"DDS combines DINOv2/SAM3 multi-level distillation with superpoint graph diffusion to produce label-free 3D semantic segmentation on driving LiDAR, beating several structure-oriented unsupervised baselines on nuScenes and SemanticKITTI.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2605.04506","ref_index":25,"ref_count":2,"confidence":0.9,"is_internal_anchor":false,"paper_title":"Ilov3Splat: Instance-Level Open-Vocabulary 3D Scene Understanding in Gaussian Splatting","primary_cat":"cs.CV","submitted_at":"2026-05-06T05:23:41+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"Ilov3Splat learns view-consistent CLIP and instance feature fields on 3D Gaussians to support open-vocabulary object selection and segmentation without category labels.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.08916","ref_index":61,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance Segmentation","primary_cat":"cs.CV","submitted_at":"2026-04-10T03:26:52+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":5.0,"formal_verification":"none","one_line_summary":"MV3DIS uses 3D-guided mask matching and depth consistency to produce more consistent multi-view 2D masks that refine into accurate zero-shot 3D instances.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2604.05697","ref_index":17,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"GraspSense: Physically Grounded Grasp and Grip Planning for a Dexterous Robotic Hand via Language-Guided Perception and Force Maps","primary_cat":"cs.RO","submitted_at":"2026-04-07T10:48:33+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":4.0,"formal_verification":"none","one_line_summary":"GraspSense computes force maps from object geometry to select mechanically safe grasp regions and regulate grip forces for dexterous hands.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2601.08831","ref_index":102,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"3AM: 3egment Anything with Geometric Consistency in Videos","primary_cat":"cs.CV","submitted_at":"2026-01-13T18:59:54+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":7.0,"formal_verification":"none","one_line_summary":"3AM integrates MUSt3R 3D features into SAM2 via a Feature Merger and FOV-aware sampling to deliver geometry-consistent video object segmentation from RGB alone, with large gains on wide-baseline datasets.","context_count":1,"top_context_role":"background","top_context_polarity":"background","context_text":"while DAM4SAM [77] incorporates distractor-aware updates. However, purely 2D approaches fail under large viewpoint changes as appearance features cannot establish reliable correspondence. Second, 3D instance segmentation methods: proposal-centric approaches like Mask3D [65] and OneFormer3D [38] operate on point clouds, while projection-based methods like Open3DIS [54], SAM3D [102], and SAM2Object [116] lift 2D masks to 3D. These require camera poses, depth maps, preprocessing, and super-linear computational scaling. Both paradigms have critical gaps (Fig. 2). 2D methods like SAM2 ex- cel at efficiency but lack geometric awareness, failing under viewpoint varia- tion (Fig. 2a). MOSEv2 [17] shows significant degradation under wide-baseline"},{"citing_arxiv_id":"2601.07447","ref_index":33,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"PanoSAMic: Panoramic Image Segmentation from SAM Feature Encoding and Dual View Fusion","primary_cat":"cs.CV","submitted_at":"2026-01-12T11:39:36+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"PanoSAMic modifies SAM with multi-stage feature encoding, spatio-modal fusion, spherical attention, and dual-view fusion to achieve SOTA panoramic semantic segmentation on public RGB and RGB-D datasets.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null},{"citing_arxiv_id":"2512.03370","ref_index":84,"ref_count":1,"confidence":0.9,"is_internal_anchor":false,"paper_title":"ShelfGaussian: Shelf-Supervised Open-Vocabulary Gaussian-based 3D Scene Understanding","primary_cat":"cs.CV","submitted_at":"2025-12-03T02:06:09+00:00","verdict":"UNVERDICTED","verdict_confidence":"LOW","novelty_score":6.0,"formal_verification":"none","one_line_summary":"ShelfGaussian achieves state-of-the-art zero-shot semantic occupancy prediction on Occ3D-nuScenes by jointly supervising Gaussian representations with vision foundation model features at 2D image and 3D scene levels.","context_count":0,"top_context_role":null,"top_context_polarity":null,"context_text":null}],"limit":50,"offset":0}