{"id":"e7b663bd-2e04-4d98-a9db-f5bda035c67a","arxiv_id":"2111.08897","paper_version":3,"verdict":"ACCEPT","confidence":"LOW","novelty_score":8.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"ARKitScenes is the largest real-world indoor RGB-D dataset captured with mobile LiDAR, including high-resolution depth maps and 3D furniture bounding box annotations for advancing object detection and depth upsampling.","lead":"The paper introduces ARKitScenes, the largest indoor RGB-D dataset captured using Apple's widely available mobile LiDAR sensors on iPads and iPhones, complete with laser-scanned high-resolution depth maps and manually labeled 3D bounding boxes for furniture. This resource supports development of 3D scene understanding methods that transfer to real consumer devices and everyday environments.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The reader's weakest assumption concerns label accuracy and representativeness for pushing SOTA on the two tasks. While those properties matter for downstream utility, they are not load-bearing for the headline claim of being first and largest. The paper's argument rests on the existence and scale of the released data, which can be verified directly from the provided statistics and files. No evidence of internal inconsistency or unsupported scale assertion is present.","tokens_in":1778,"tokens_out":341,"duration_ms":28242,"concrete_test":"Extract the exact counts of rooms, RGB-D frames, and 3D bounding boxes from the Dataset Statistics section; recompute the comparison table against ScanNet and Matterport3D using those numbers. If ARKitScenes is not strictly larger on at least two of the three metrics, the 'largest' claim requires qualification.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that ARKitScenes is the first RGB-D dataset captured with the now-widely-available Apple LiDAR sensor and, to the authors' knowledge, the largest indoor scene-understanding dataset released. This is a factual assertion about the release itself. The paper supplies the raw mobile captures, registered laser-scanned depth, and manually annotated 3D oriented bounding boxes, together with explicit size statistics and comparisons to prior datasets (ScanNet, Matterport3D, etc.). No internal contradiction appears in the scale argument or in the two downstream-task demonstrations; the work does not claim to have solved a new algorithmic problem, only to have supplied a resource whose characteristics better match real-world mobile capture.","agreement_with_reader":"disagree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces ARKitScenes as the first RGB-D dataset captured with Apple's widely available LiDAR sensor on mobile iPads/iPhones and, to the authors' knowledge, the largest indoor scene understanding dataset released. It supplies raw and processed mobile RGB-D captures, registered high-resolution depth maps from a stationary laser scanner, and manually annotated 3D oriented bounding boxes over a furniture taxonomy. The authors compare scale and characteristics to prior datasets (ScanNet, Matterport3D) and demonstrate utility on two downstream tasks: 3D object detection and color-guided depth upsampling, claiming the data pushes SOTA boundaries while introducing real-world challenges.","tokens_in":1896,"tokens_out":486,"duration_ms":50645,"significance":"If the scale, registration quality, and annotation accuracy hold, the release supplies a high-value resource whose mobile capture characteristics better match everyday consumer hardware than prior lab-style datasets. This can accelerate development of robust 3D scene understanding methods for mobile applications, with the laser-scanned depths and 3D boxes providing strong supervision signals for detection and upsampling benchmarks.","major_comments":[{"comment":"§4 (Dataset Statistics): the central claim that ARKitScenes is the largest indoor dataset requires an explicit side-by-side table (number of scenes, frames, annotated objects, capture conditions) against ScanNet and Matterport3D; without these numbers the size/diversity assertion is unsupported.","section":"§4"},{"comment":"§6 (Downstream Tasks): the demonstrations for 3D object detection and depth upsampling must report concrete metrics (mAP, RMSE, etc.) and baselines; the abstract states only that the data 'pushes boundaries' without evidence, which is load-bearing for the utility claim.","section":"§6"}],"minor_comments":[{"comment":"Figure captions should explicitly state what each panel shows (RGB, mobile depth, laser depth, projected boxes) and include scale bars or units.","section":null},{"comment":"The taxonomy of furniture classes and the exact annotation protocol (number of annotators, quality control) should be listed in a dedicated subsection or table.","section":"§3"}],"recommendation":"minor_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the positive recommendation of minor revision and the constructive comments. We address each point below.","responses":[{"response":"We agree that an explicit comparison table will strengthen the claim. In the revised manuscript we will insert a side-by-side table in §4 that reports number of scenes, frames, annotated objects, and capture conditions for ARKitScenes, ScanNet, and Matterport3D.","revision_made":"yes","referee_comment":"[§4] §4 (Dataset Statistics): the central claim that ARKitScenes is the largest indoor dataset requires an explicit side-by-side table (number of scenes, frames, annotated objects, capture conditions) against ScanNet and Matterport3D; without these numbers the size/diversity assertion is unsupported."},{"response":"We will revise the abstract to include the key quantitative results (mAP for detection and RMSE for upsampling) and will ensure §6 explicitly lists all metrics together with the baselines used. This will provide the concrete evidence requested.","revision_made":"yes","referee_comment":"[§6] §6 (Downstream Tasks): the demonstrations for 3D object detection and depth upsampling must report concrete metrics (mAP, RMSE, etc.) and baselines; the abstract states only that the data 'pushes boundaries' without evidence, which is load-bearing for the utility claim."}],"tokens_in":1436,"tokens_out":315,"duration_ms":42587,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The punchline is that ARKitScenes is a large new RGB-D dataset captured on consumer Apple devices with LiDAR, plus laser ground truth and 3D box labels. This is the part that matters most for the field. The work does a good job releasing data from the actual sensors people use now, instead of older Kinect-style setups. The scale is bigger than prior indoor datasets, and adding stationary laser scans gives a higher-quality reference for the mobile captures. The 3D oriented bounding boxes for furniture add another layer that supports detection tasks. They also show example uses in object detection and depth upsampling, which helps illustrate the data's relevance. The paper compares the new dataset to ScanNet and Matterport3D on size and properties, which is helpful. The claim that it introduces real-world challenges seems plausible given the mobile capture conditions. On the soft side, the description does not include much quantitative validation of the data quality or detailed collection protocols. For claims about diversity and accuracy, more evidence like error rates between mobile and laser depth or inter-annotator agreement on boxes would strengthen it. The downstream results are presented but without full numbers or baselines in the abstract, so the impact on SOTA is asserted rather than fully demonstrated here. This paper is mainly for researchers in 3D computer vision and AR who need training data that matches consumer hardware. It is worth sending to peer review because the dataset itself is a substantial new resource, even if the accompanying analysis stays at a high level.","headline":"ARKitScenes releases a large mobile RGB-D dataset from real Apple LiDAR hardware plus laser ground truth and 3D boxes, which is the useful part even if the paper stays mostly descriptive.","tokens_in":2419,"tokens_out":386,"would_cite":true,"duration_ms":47347,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":{"model":"grok-4.3","evidence":[{"relation":"unclear","rs_module":"IndisputableMonolith.Foundation.DimensionForcing","rs_theorem":null,"paper_passage":"In this paper we introduce ARKitScenes. It is not only the first RGB-D dataset that is captured with a now widely available depth sensor, but to our best knowledge, it also is the largest indoor scene understanding data released."},{"relation":"unclear","rs_module":"IndisputableMonolith.Foundation.HierarchyEmergence","rs_theorem":null,"paper_passage":"We further analyze the usefulness of the data for two downstream tasks: 3D object detection and color-guided depth upsampling."}],"headline":"Dataset for RGB-D indoor scene understanding with no connection to RS cost-based physics derivation","alignment":"orthogonal","rationale":"The paper introduces ARKitScenes, a large RGB-D dataset captured with mobile LiDAR for tasks like 3D object detection and depth upsampling. It focuses on empirical data collection, ground-truth registration, and ML benchmarks. No reference to RS concepts such as J-cost, φ-ladder, 8-tick periodicity, D=3 forcing, or recognition logic appears. The work operates in applied computer vision, a domain where RS has no opinion.","tokens_in":271587,"confidence":"high","tokens_out":300,"duration_ms":35744,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"lean_confirmation":{"model":"grok-4.3","status":"out_of_scope","citations":[],"rationale":"This is a dataset-release paper whose load-bearing content is empirical (size, diversity, hardware, task performance). Shape-of-logic contains no theorems about RGB-D sensors, indoor scene reconstruction, or ML benchmarks; the premise lies outside its scope.","tokens_in":271340,"confidence":"moderate","tokens_out":222,"duration_ms":29428,"inferential_bridge":"The paper's central claims rest on empirical assertions about dataset scale (5048 sequences, 1661 scenes), capture hardware, registration accuracy between mobile RGB-D and stationary laser scans, and quantitative improvements on downstream ML tasks. These are observational facts about real-world data collection and experimental results; they cannot be reduced to a machine-checkable mathematical identity or structural theorem in shape-of-logic.","load_bearing_premise":"ARKitScenes is the largest indoor RGB-D dataset captured with a widely available mobile depth sensor (Apple LiDAR), providing high-resolution ground truth depth maps and 3D oriented bounding boxes that enable state-of-the-art performance on 3D object detection and color-guided depth upsampling.","cache_read_input_tokens":64,"cache_creation_input_tokens":0},"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"ARKitScenes is the largest indoor RGB-D dataset captured with widely available mobile LiDAR sensors and includes laser-scanned depth plus manual 3D bounding box labels.","keywords":["RGB-D dataset","indoor scene understanding","3D object detection","depth upsampling","mobile LiDAR","3D bounding boxes","real-world dataset","ARKitScenes"],"falsifier":"A controlled test in which models trained on ARKitScenes show no improvement over models trained on prior datasets when evaluated on independent mobile RGB-D captures from varied indoor rooms would falsify the usefulness claim.","tokens_in":2680,"feed_emoji":"📱","tokens_out":736,"duration_ms":41972,"temperature":0.7,"pith_summary":"The paper presents ARKitScenes, a dataset of RGB-D captures collected from Apple iPads and iPhones that have LiDAR sensors. It augments the raw mobile data with high-resolution depth maps from a stationary laser scanner and manual 3D oriented bounding box labels for a large set of furniture categories. The authors test the data on two tasks, 3D object detection and color-guided depth upsampling, and report that it improves existing methods while exposing challenges closer to everyday conditions. A sympathetic reader would care because the captures come from devices already owned by millions of people, moving 3D scene understanding from controlled lab settings toward practical mobile use.","feed_headline":"Largest mobile RGB-D dataset released for indoor 3D scenes","feed_subtitle":"ARKitScenes pairs iPhone captures with laser depth maps and furniture labels to improve detection and upsampling","key_machinery":"The ARKitScenes dataset that pairs mobile RGB-D captures with laser-scanner depth maps and manual 3D bounding box annotations for indoor furniture.","core_discovery":"ARKitScenes is the first RGB-D dataset captured with the widely available depth sensor on iPads and iPhones and the largest indoor scene understanding dataset released. It supplies raw and processed mobile device data, high-resolution depth maps from a stationary laser scanner, and manually labeled 3D oriented bounding boxes for furniture. Evaluation on 3D object detection and color-guided depth upsampling shows the dataset pushes state-of-the-art performance and introduces new real-world challenges.","pith_inferences":["App developers could fine-tune models on this data to add room-layout awareness to consumer AR experiences without extra hardware.","The dataset could be used to study how well algorithms generalize from mobile captures to other depth sensors.","Future releases might add semantic segmentation labels or dynamic object tracks to extend the current static bounding-box focus.","Cross-validation across different device models within the captures could reveal hardware-specific biases in depth sensing."],"forward_implications":["3D object detection models achieve higher accuracy on large furniture taxonomies when trained with the labeled mobile data.","Color-guided depth upsampling produces higher-resolution outputs by using the laser scans as precise ground truth.","The dataset scale supports training larger machine-learning models for indoor scene understanding.","Methods developed on the data must handle noise and viewpoint variation typical of handheld mobile captures.","The combination of mobile and laser data creates a bridge between consumer hardware and high-precision references."],"fun_headline_variants":["ARKitScenes largest indoor RGB-D dataset using mobile sensors","ARKitScenes pairs mobile RGB-D with laser scans and 3D labels","ARKitScenes first RGB-D dataset from Apple depth sensors","ARKitScenes improves 3D detection using mobile and laser data","ARKitScenes adds furniture labels to mobile scene understanding"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The mobile RGB-D captures, laser-scanned depth maps, and manual 3D bounding box labels are sufficiently accurate and representative of real-world indoor scenes to advance state-of-the-art methods.","fun_headline_variants_meta":{"raw":{"variants":["ARKitScenes largest indoor RGB-D dataset using mobile sensors","ARKitScenes pairs mobile RGB-D with laser scans and 3D labels","ARKitScenes first RGB-D dataset from Apple depth sensors","ARKitScenes improves 3D detection using mobile and laser data","ARKitScenes adds furniture labels to mobile scene understanding"]},"model":"grok-4.3","cost_usd":0.007837,"raw_usage":{"total_tokens":3525,"prompt_tokens":727,"num_sources_used":0,"completion_tokens":84,"cost_in_usd_ticks":78365500,"prompt_tokens_details":{"text_tokens":727,"audio_tokens":0,"image_tokens":0,"cached_tokens":64},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2714,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":727,"tokens_out":84,"duration_ms":33125,"temperature":1.0,"reasoning_tokens":2714,"cache_read_input_tokens":64,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-15T10:41:47.251267+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled test in which models trained on ARKitScenes show no improvement over models trained on prior datasets when evaluated on independent mobile RGB-D captures from varied indoor rooms would falsify the usefulness claim.","supporting_citations":[],"review_version":1}