{"id":"0ef4a371-3428-4453-88c8-38a2469de5d9","arxiv_id":"2411.13149","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"New recordings of YCB household objects on a light-absorbing black screen, with automatic masking code, extending YCB-V-LUMA to transparent, reflective, dark, and deformable objects.","lead":"This paper releases a new image dataset of everyday household objects filmed against a black light-absorbing screen, along with code that turns the videos into training images for object detection. It extends an earlier dataset by adding harder objects like transparent cups, chains, and colored toys, which could test how well computer vision models cope with shiny or bendy items.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's own recording log documents masking failures on the very object classes the extension adds, and no quantitative mask-quality evaluation is provided to support the high-quality annotation claim.","rationale":"The reader's weakest-assumption analysis identifies brightness-threshold mask quality as the load-bearing point, and the paper's own Figure 2 provides internal evidence that this assumption fails for a substantial subset of the newly recorded objects. My reading reinforces that concern rather than replacing it. The paper is a dataset/extension contribution, not an algorithmic claim, so the appropriate bar is whether the released recordings and automatic annotation code actually deliver usable training data for the objects they add. Since the manuscript gives no quantitative mask-quality evaluation and its own metadata lists multiple objects with known masking failures and retake flags, the conditional verdict is correct. I would not reject the paper outright: the raw recordings may still be valuable, previous work on luminance keying suggests the method can work for many objects, and the release may be useful even if some objects need manual mask correction. But the high-quality claim needs empirical support before the dataset can be recommended for direct use. The proposed concrete test, manual IoU evaluation on the flagged and easy objects, would settle whether the concern actually lands: if the automatic masks are accurate, the objection dissolves; if they are not, the paper should be revised to scope the claim, exclude failing objects, or add a correction step. The reader's conditional verdict already captures this risk, so no verdict adjustment is needed.","tokens_in":4716,"tokens_out":5653,"duration_ms":60396,"concrete_test":"Use the released repository and dataset to regenerate masks from the raw recordings for a stratified sample of objects, including flagged cases such as YCB IDs 9, 26, 30, 36, 42, 61, 68, 74, and 79 plus at least one easy object such as ID 55. Manually annotate pixel-accurate foreground masks on 50 frames per object sampled across viewpoints, then compute the Dice/IoU between the automatic masks and the manual masks, reporting per-object means and the fraction of frames below 0.8 IoU. If any of the previously flagged objects 36, 61, or 74 has a mean IoU below 0.8, the high-quality annotation claim is not supported for that object, and the paper should either exclude such objects or document manual correction as part of the pipeline.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the new YCB-LUMA release provides high-quality recordings and code for automatic generation of training data and annotations. This claim depends on the assumption that brightness thresholding against the 99.99% light-absorbing background produces clean foreground masks for all newly recorded object types. The paper's own Figure 2 documents failures of exactly this assumption for the novel classes: object 9 has a 'dark part (moustache)' that is not detected; objects 26-28 report noisy reflective metal; object 30 (wine glass) has 'a lot of noisy results'; object 36 is marked 'Retake? Yes' because of reflective lights; objects 42-43 have darker handle parts that the script cannot read; object 61 is marked 'Retake? Yes' because the dark-shiny marble could not be handled; object 68 is a deformable chain, for which thin and self-occluded segments are prone to masking errors; objects 74-75 are marked 'Retake? Yes' due to dark, noisy results; and object 79 reports that the smallest washers could not be detected. Transparent, reflective, dark, and deformable objects have foreground pixels with luminance values close to the background, so a single global threshold will not yield complete, clean masks for them. The manuscript reports no manual verification of the generated masks, no mask-quality metric such as IoU against manual annotations, and no downstream detection or segmentation training experiment. Consequently, the 'high quality' and 'usable annotations' claims are unsupported for the new objects that constitute the contribution, and the paper's own metadata explicitly flags several of these objects for retake.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents YCB-LUMA, a dataset that extends the existing YCB-V luminance-keying recordings to the remaining objects in the YCB superset. The newly recorded objects include transparent plastics and glass, reflective metal, dark parts, and deformable objects such as a chain and a cord. The authors provide the recordings and scripts that automatically generate masks for 2D object detection and segmentation by brightness thresholding against a 99.99% light-absorbing background. The central claims are that the new release provides high-quality recordings and code for automatic generation of training data and annotations, and that the added object variety demonstrates the usefulness of luminance keying.","tokens_in":5018,"tokens_out":3248,"duration_ms":33691,"significance":"If the claims are substantiated, the contribution is practically useful: it extends a widely used benchmark object set, removes the need for manual labeling in many cases, and includes challenging object categories that are often missing from such datasets. The authors should be credited for publishing URLs to the data and code, and for including a detailed recording log in Figure 2 that transparently reports known issues. However, the usefulness claim is currently unvalidated: the paper's own log documents masking failures on many of the very objects that motivate the extension, and no quantitative mask-quality evaluation or downstream training experiment is provided. The significance is therefore conditional on the authors either fixing the problematic recordings or clearly delimiting the failure modes and measuring the quality of the generated masks.","major_comments":[{"comment":"The recording log in Figure 2 documents masking failures for many of the newly recorded objects: dark parts not detected (objects 9, 32, 42-43, 61), noisy edges from reflections (objects 26-28, 58-60, 76), and objects explicitly marked for retake (objects 36, 61, 74-75). These are precisely the transparent, reflective, dark, and deformable categories that the paper presents as the added variety of the extension. The abstract's claim of \"high quality data\" and Section 4's claim of \"high quality recordings\" are therefore not supported by the evidence in the paper. The authors should either retake or exclude these objects, or provide per-object mask-quality metrics and clearly state the failure rate in the paper.","section":"Section 3, Figure 2"},{"comment":"The central claim that the provided code generates usable training data is not validated. No experiments are reported that use the generated masks for detector or segmentation training, and no mask-quality metric is computed. A small demonstration (for example, training a standard detector on LUMA-generated masks and reporting AP on a hold-out set, or measuring IoU between generated and manually corrected masks on a sample of frames) would provide the missing evidence. Without it, the statement that the code produces usable annotations remains an assertion.","section":"Section 3, 'scripts to automatically extract training data'"},{"comment":"For the deformable objects (chain, cord), the paper states that multiple deformation states are recorded, but the automatic masking of thin, self-occluded chain links is prone to errors, as the log suggests for object 68. The release should include a per-sequence statement of whether the automatic masks were verified or hand-corrected for these objects, and the paper should indicate how much manual intervention is needed for the deformable category.","section":"Section 3, Figure 1 and object 68"}],"minor_comments":[{"comment":"The phrase \"quality insurance\" should be \"quality assurance\".","section":"Abstract"},{"comment":"The table is extremely dense and the column headers are ambiguous: \"ObjectAvailable?\", \"Different?\", and \"Difference?\" need definition in the caption, and the meaning of the \"Count\" and \"Retake?\" columns should be explicitly explained so that readers can interpret the log.","section":"Figure 2"},{"comment":"The sentence \"These new recordings complement the original ones depicting all of the objects to be found in the YCB-V subset\" is confusing, because YCB-V is itself a subset of YCB. Reword to say that the new recordings cover the remaining YCB objects not already present in YCB-V.","section":"Section 3, first paragraph"},{"comment":"The paper should state the total number of videos, frames, and per-object image counts, as well as the license for the dataset and the code, since these facts are essential for a dataset release.","section":"Section 4"},{"comment":"Reference [8] (Pöllabauer et al., \"Advanced post-processing for object detection dataset generation\") is listed without a venue or year; please provide the full publication details.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a short dataset paper with no empirical validation. The main risk is that the released masks for a substantial subset of objects are not clean enough to train from without manual correction. The authors' own Figure 2 provides evidence of this risk. A revision that retakes or excludes the flagged objects and adds a small quantitative validation would make the contribution solid. If the authors instead only add text acknowledging limitations without changing the release or measuring mask quality, I would not consider the claims adequately supported. The paper may also be better suited to a data-descriptor venue if no downstream experiments are added."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the YCB-LUMA paper. It does what it says: records the remaining YCB objects beyond YCB-V using the luminance keying setup from prior work, and ships the data and processing code. That's a real artifact. The added object types—transparent plastic/glass, reflective metal, dark plastic, deformable chain—are genuinely harder for brightness thresholding, and including them is the right stress test for the method.\n\nWhat's good: the release is complete by the authors' own accounting, the metadata table is unusually honest, and the scripts for automatic masking are provided. The paper also gives a clear path for differentiating object variations, which is useful for generalization tests. I believe the dataset will be a resource for 2D detection/segmentation training, especially for people already using YCB-V-LUMA.\n\nThe soft spot is exactly where the stress-test lands. The central claim says 'high quality recordings' and 'automatic generation of training data and annotations.' But the paper provides no quantitative evaluation of mask quality, no IoU against manual annotations, and no downstream detector/segmentation training to show the annotations are usable. Worse, the recording table documents repeated masking failures on the very classes the extension adds: object 9 misses the dark moustache; objects 26-28 have noisy reflective metal; object 30 has noisy results on the wine glass; objects 36, 61, 74-75 are explicitly marked for retake; object 79 can't detect the smallest washers; object 68 is a deformable chain where thin segments are prone to errors. A single global brightness threshold against a 99.99% absorbing background will not cleanly separate dark or transparent foreground from background. The authors acknowledge some of this in the comments but still assert high quality without validation.\n\nThis is not a fatal flaw in the dataset—recordings can still be useful, and the authors may have manually filtered or adjusted thresholds. But the paper as written overclaims. The fix is straightforward: report mask quality metrics on a held-out set of manually labeled frames, or at minimum train a YOLO/Mask R-CNN on the auto-generated data and report AP/mAP. That would turn a conditional release into a defensible one.\n\nThe citation pattern is fine; the method is from the authors' own prior work and they cite it. The paper is an incremental data extension, not a new technique, but it is a legitimate extension. I'd send it to peer review—dataset papers deserve referee time when the artifact is real—but I'd ask for the quality validation before acceptance. The reader's conditional verdict is right.","headline":"Useful dataset extension, but the paper's own recording log undercuts the 'high quality' claim for the new object classes; needs mask-quality validation.","tokens_in":5503,"tokens_out":1716,"would_cite":false,"duration_ms":16276,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"By recording the remaining YCB household objects against a 99.99% light-absorbing screen and extracting masks by brightness thresholding, this paper extends luminance-keying training data to the full YCB object set, including transparent…","keywords":["YCB dataset","luminance keying","object detection","object segmentation","synthetic data generation","training data","object localization","dataset extension"],"falsifier":"Run the provided auto-masking scripts on the recordings the paper flags for retake or noise, such as the dark-shiny large marble (object 61), the plastic bolt (object 74), and the plastic nut (object 75), and compare the resulting masks pixel-by-pixel with manually drawn masks; a large drop in intersection-over-union for these objects would show that the automatic annotation claim does not hold uniformly.","tokens_in":4552,"feed_emoji":"📷","tokens_out":8493,"duration_ms":75488,"temperature":0.7,"pith_summary":"This paper extends the YCB-LUMA training-data suite from the YCB-V subset to the complete YCB object set. Following the luminance keying method, the author records the previously missing objects — transparent plastic and glass, reflective metal, dark and multi-color variations, and deformable items such as a chain and a cord — against a 99.99% light-absorbing black screen. Brightness thresholding turns these recordings into foreground masks, and released scripts generate 2D object detection and segmentation training data with annotations. The paper's central claim is that the full YCB set can now be covered by luminance keying, increasing the variety of appearances available for testing detection and segmentation algorithms without manual labeling.","feed_headline":"Full YCB object set recorded with luminance keying","feed_subtitle":"Transparent, metallic, dark, and deformable objects join the YCB-V recordings, with code for automatic masks.","key_machinery":"Luminance keying: recording objects in front of a 99.99% light-absorbing black screen so the background has near-zero brightness, making foreground extraction a brightness threshold rather than color-based chroma keying. In this paper the mechanism both produces the automatic masks and annotations and is the method under test, because the new objects add transparency, reflection, dark surfaces, and deformation. The released processing scripts are the operational part of the machinery, converting raw recordings into 2D detector and segmentation training data without manual labeling.","core_discovery":"On its own terms, the paper's contribution is a dataset rather than a new algorithm: it records every YCB object not already present in the YCB-V subset under the same luminance keying setup, and it supplies code that automatically generates detection and segmentation training data from the recordings. The newly recorded objects intentionally include appearance classes that stress the keying approach: transparency, specular metal, dark parts, multiple color variants of the same object, and deformable shapes. Color variants are stored in separate subfolders, which enables training on one variant and testing on another or combining all variants in one set. The stated result is that the YCB-LUMA set now covers the full YCB object set with high quality recordings and automatic annotation, extending the usefulness of the previous YCB-V luminance keying data.","pith_inferences":["The paper's own metadata table marks several objects as noisy, undetected, or needing retake (for example objects 21, 26, 36, 61, 74, 75), so the pipeline likely requires per-object threshold tuning or post-processing rather than being fully automatic for every appearance class.","The paper reports no quantitative comparison between automatic masks and manual masks, leaving mask quality on the newly added objects as the main open question.","If the auto-masks are validated, the same black-screen setup could record arbitrary new objects, making luminance keying a practical alternative to 3D-model-based rendering for niche detection tasks."],"forward_implications":["Every YCB object class now has luminance-keying recordings, so 2D object detectors and segmentation models can be trained on the full YCB set without manual annotation.","The added transparent, metallic, and deformable objects provide a stress test for whether the luminance keying approach generalizes beyond the original YCB-V subset.","Separate subfolders for each color variant enable controlled experiments on generalization, such as training on one variant and evaluating on another.","The released processing scripts mean other objects recorded with the same black-screen setup can be converted into detection and segmentation training data automatically."],"supporting_citations":[{"why":"Introduces the luminance keying method and the 99.99% light-absorbing screen that the dataset depends on.","marker":"[10]"},{"why":"Supplies the dataset-generation post-processing that the released scripts build on.","marker":"[8]"},{"why":"Defines the YCB-V subset that this work complements with the remaining YCB objects.","marker":"[14]"}],"fun_headline_variants":["YCB-LUMA completes YCB set with luminance keying","Every YCB object now recorded via luminance keying","Luminance keying expands to full YCB superset","YCB-LUMA: all YCB objects with automatic masks","Dataset covers every YCB object using luminance keying"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole automatic-annotation pipeline depends on a simple brightness threshold cleanly separating each object from the black background, and the paper's own recording table shows this already fails for several dark, shiny, and reflective objects.","fun_headline_variants_meta":{"raw":{"variants":["YCB-LUMA completes YCB set with luminance keying","Every YCB object now recorded via luminance keying","Luminance keying expands to full YCB superset","YCB-LUMA: all YCB objects with automatic masks","Dataset covers every YCB object using luminance keying"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000338,"raw_usage":{"total_tokens":1834,"prompt_tokens":880,"completion_tokens":954,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":496,"completion_tokens_details":{"reasoning_tokens":870}},"tokens_in":496,"tokens_out":954,"duration_ms":7723,"temperature":1.0,"reasoning_tokens":870,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:46:45.222669+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the provided auto-masking scripts on the recordings the paper flags for retake or noise, such as the dark-shiny large marble (object 61), the plastic bolt (object 74), and the plastic nut (object 75), and compare the resulting masks pixel-by-pixel with manually drawn masks; a large drop in intersection-over-union for these objects would show that the automatic annotation claim does not hold uniformly.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the dataset-generation post-processing that the released scripts build on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the luminance keying method and the 99.99% light-absorbing screen that the dataset depends on."}],"review_version":1}