{"id":"e0311de1-1e8b-4421-a905-a6785fb5a5f2","arxiv_id":"2506.08964","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"ORIDa is a public real-world dataset of 200 objects in 30,000+ images with multiple positions per scene, designed for object compositing training and evaluation.","lead":"The authors introduce ORIDa, a new collection of over 30,000 real photographs showing 200 everyday objects across many scenes and positions, each with matching background-only shots. It is built to train AI models that insert or remove objects from photos, and a first test suggests it improves realism compared with existing methods.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The counterfactual capture assumption is not quantitatively validated; residual scene drift or misalignment between background-only and factual images would corrupt the paired supervision that the entire dataset is built on.","rationale":"The reader's weakest assumption targets the very foundation of the dataset: the factual-counterfactual pairing. This is indeed the most load-bearing concern because ORIDa's primary contribution is the paired supervision itself, and both training tasks (removal and insertion) are defined relative to it. Without a validated counterfactual, the dataset reduces to a collection of unrelated photos with unreliable targets, and the reported fine-tuning gains would be confounded. I considered two alternatives. First, the insertion model's train/inference mismatch (training input latent is a noisy ground-truth composite, while inference input is a Copy-and-Paste image) is a serious evaluation flaw that could inflate the reported user-study results; however, it threatens the secondary 'models outperform' claim, not the dataset's intrinsic value. Second, the dataset's 'publicly available' claim is not verifiable from the manuscript, but this is a practical issue and can be resolved by a link. The counterfactual validity is more fundamental and is also less checked: the paper relies on manual filtering, which cannot quantify sub-pixel alignment or slow photometric drift. A direct registration and residual-difference test on a sample of sets would settle the question. Since the reader already assigned CONDITIONAL based on this and other addressable issues, my read does not change the verdict; it strengthens the conditional status.","tokens_in":12501,"tokens_out":9676,"duration_ms":104059,"concrete_test":"Select a random subset of about 100 F-CF sets. For each set, extract SIFT/ORB keypoints in the background-only and each factual image, but only in regions outside the object mask and its visible shadow or reflection. Estimate a homography; compute the median re-projection error and the median absolute photometric difference in those regions after alignment. Also compare the background to a synthetic background obtained by inpainting the object region in a factual image; if the aligned residual RMS in non-object regions exceeds 1.0 pixel or 2/255 sRGB units, the counterfactual assumption fails and the paired target is unreliable. Report these histograms for all sampled sets.","verdict_should_be":"UNCHANGED","load_bearing_attack":"ORIDa's training signal for both object removal and insertion depends on the background-only image being a pixel-aligned, photometrically consistent counterfactual of each factual image: the only differences should be the object and its object-to-scene effects such as shadows and reflections. Section 3.2 states that shutter speed, ISO, WB, and focus were fixed and that a remote controller was used, and Section 3.3 filters obvious 'undesired background changes' manually. However, the paper provides no quantitative evidence that, after this protocol, background and factual images are registered to sub-pixel accuracy and that residual differences in non-object regions are within sensor noise. Even a tripod-mounted camera can experience micro-movement when the object is placed or removed between captures, and lighting can drift over the five-shot sequence; manual inspection cannot reliably detect sub-pixel shifts or slow illumination changes. If such drift exists, then the paired target for removal is not the true counterfactual, and the diffusion training loss will fit to incorrect pixels, undermining the dataset's central value. The same assumption underpins the four-position F-CF sets, since all four factual images share one background.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"ORIDa is a new large-scale, real-captured dataset for object compositing, containing over 30,000 images of 200 objects. Each object appears in an average of roughly 50 scenes, organized into factual-counterfactual (F-CF) sets, where one background-only image accompanies four images of the object in different positions, and factual-only (F-Only) images that provide additional scene diversity. The dataset includes raw DNG files, captions, object points, bounding boxes, and SAM2-generated segmentation masks. The authors also fine-tune Stable Diffusion Inpaint (SD-Inpaint) for object removal and insertion, reporting qualitative results, user studies, and automatic metrics that favor models trained on ORIDa over several baselines. The paper's central claim is that ORIDa is the first publicly available, large-scale, real-captured dataset for object compositing and that it enables improved realism in object removal and insertion without synthetic training data.","tokens_in":12724,"tokens_out":3188,"duration_ms":35912,"significance":"If the core capture assumption holds, ORIDa would be a substantial community resource: it is the first public dataset of this scale to combine real-captured factual-counterfactual pairs, multiple object positions per scene, multiple scenes per object, raw DNG flexibility, and rich annotations. The scale and design are clear advances over ObjectDrop, which is neither public nor multi-scene. The paper also contributes a concrete fine-tuning recipe for SD-Inpaint and reports both user and automatic evaluations. However, the empirical validation currently falls short of fully supporting the dataset's central value claim, primarily because the counterfactual-consistency assumption is asserted rather than quantitatively validated, and because the insertion experiments combine ORIDa with additional COCO data and a modified inference procedure, making the isolated contribution of ORIDa unclear. These issues are fixable through additional validation and control experiments, so the manuscript warrants major revision rather than rejection.","major_comments":[{"comment":"The dataset's paired training signal rests on the assumption that each background-only image is a pixel-aligned, photometrically consistent counterfactual of its four factual counterparts. The paper states that camera settings were fixed, tripods and remote controllers were used, and undesirable cases were filtered manually, but it provides no quantitative evidence that residual scene drift, micro-movement, or illumination change is negligible. A tripod-mounted camera can shift slightly when objects are placed or removed, and lighting can drift over the five-shot sequence; manual inspection cannot reliably detect sub-pixel shifts or slow photometric changes. I request quantitative validation: for example, registration residuals between background and factual images in non-object regions, histograms of per-pixel differences under the object mask and outside it, or a comparison against sensor-noise baselines. This is load-bearing because the entire F-CF training signal and the dataset's uniqueness depend on the counterfactual being correct.","section":"§3.2–3.3"},{"comment":"The object insertion experiments are not an isolated validation of ORIDa. Training uses an additional 60,000 COCO images with 250,000 object masks, and inference uses a skip-residual modification from DemoFusion that is not part of the pretrained SD-Inpaint pipeline. Consequently, the reported gains in identity preservation, shadow generation, and harmonization could come substantially from the COCO training data, the inference-time skip residual, or their interaction, rather than from ORIDa itself. The paper should ablate these factors, e.g., fine-tune on ORIDa F-CF data alone without COCO, or run the unmodified inference procedure, and report how each component affects the user-study and automatic results. Without such controls, the claim that ORIDa alone supports the observed insertion quality is not established.","section":"§5.1 and §B.3"},{"comment":"The user studies are reported without error bars, statistical significance tests, or inter-rater statistics. Table 2 reports mean ratings for 76 participants but no variance or pairwise significance, and Figure 11 reports preference percentages for 62 participants but no confidence intervals or significance tests. The automatic comparison in Table 3 also covers only SD-Inpaint, while the qualitative and user comparisons include LaMa and MGIE; automatic metrics should be reported for all removal baselines. Adding these statistics is necessary to support the strong claims that ORIDa-trained models 'significantly outperform' existing methods.","section":"§5.2, Table 2, Table 3, and Figure 11"},{"comment":"The evaluation for object removal is conducted on an 'out-held test set' from ORIDa, while insertion is evaluated qualitatively on COCO, internet, and MureCom images. This is reasonable as a start, but the removal numbers on a held-out subset of the same capture campaign may largely reflect the model learning dataset-specific capture conditions rather than generalizable scene understanding. I recommend adding a small cross-dataset or in-the-wild quantitative evaluation for removal (using existing paired or benchmark data where possible), or at least clearly stating and discussing this limitation in the main text.","section":"§5.2 and §5.3"}],"minor_comments":[{"comment":"The dataset name is inconsistently spelled as both 'ORIDa' and 'ORIDA' (e.g., in the Introduction and Table 1 row labels); please use a single consistent spelling.","section":"Throughout"},{"comment":"There are typographical errors in figure text: 'Factual-Counterfactal' in Figure 1 and 'Obejct' in Figure 10 should be corrected.","section":"Figure 1 and Figure 10"},{"comment":"The sentence 'adapting its colors seamlessly to the scene and and generating natural shadows' contains a duplicated 'and'.","section":"§5.3"},{"comment":"'out-held test set' should be 'held-out test set'; please also state how many F-CF sets or images are used for the automatic evaluation in Table 3.","section":"§5.2"},{"comment":"The paper should clarify whether the ISP augmentations (five Lightroom settings) are applied to all images or only to a subset, and how these augmentations interact with the raw DNG files during training; currently this is described only briefly.","section":"§3.4 and §5.1"},{"comment":"Reference [2] is cited as 'MureCom' in the text but the actual title is 'MureObjectStitch'; please align the citation name.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the unvalidated counterfactual-consistency assumption in the F-CF captures; if micro-movement or lighting drift is non-negligible, the paired supervision is degraded and the dataset's central advantage is weakened. This is not a reason to reject, because the dataset release itself is valuable, but the authors should be asked to provide direct evidence of registration and photometric stability, or an explicit quantitative analysis of the magnitude of residual drift. The insertion experiments should also be disentangled from COCO data and the skip-residual inference modification. The paper's scope fits a vision venue well, and the contribution is potentially significant if these validations are added."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The dataset is the main event, and it's real. ORIDa gives the community the first public, large-scale, real-captured factual-counterfactual collection for object compositing: 200 objects, 30k+ images, multiple positions per scene, raw DNGs, and decent annotations. That fills a concrete gap, since ObjectDrop is small and not public. The capture protocol is disciplined — tripod, remote shutter, fixed exposure/ISO/WB/focus, five consecutive shots — and the filtering from 7,000 to 5,699 F-CF sets shows they actually checked what they captured. Releasing DNGs and enabling ISP augmentations is a thoughtful extra. This is a resource the compositing community will want to train on and cite.\n\nThe soft spots are in the validation, not the collection. The stress-test concern is fair: the entire paired training signal depends on the background-only image being a pixel-aligned, photometrically consistent counterfactual, and the paper gives no quantitative evidence that this holds. Tripod and remote reduce camera motion, and manual filtering catches obvious lighting shifts, but a sub-pixel shift or slow illumination drift is exactly the kind of thing manual inspection misses. I don't think this is fatal — the protocol likely keeps drift small — but \"likely\" is not a measurement. A simple check on static background regions (registration residual or pixel-difference distribution versus sensor noise) would settle it, and the authors should run it.\n\nThe evaluation also overreaches. Removal automatic metrics compare only against SD-Inpaint, not LaMa or MGIE, which the user study does cover but without error bars or significance tests. For insertion, the model uses a modified inference (DemoFusion-style skip residual) and extra COCO training data, so the reported gains are not purely attributable to the dataset. The authors disclose this, which I respect, but it weakens the headline claim that ORIDa alone drives the improvement. The held-out test set from the same capture campaign is standard, though it does mean the numbers reflect dataset-specific conditions.\n\nNone of these are structural problems. The resource is valuable, the collection is careful, and the flaws are addressable. Send it to peer review; a serious referee should ask for the alignment validation and tighter statistics, but rejecting this would throw out a genuinely useful dataset. I'd bring it to reading group and would cite it if I worked on compositing.","headline":"A genuinely useful real-captured compositing dataset with careful collection, but the evaluation overreaches and the counterfactual alignment assumption is unquantified.","tokens_in":13239,"tokens_out":1958,"would_cite":true,"duration_ms":21524,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ORIDa introduces a 30,000-image real-captured dataset that makes object removal and insertion trainable without synthetic data, with models fine-tuned on it beating prior methods in user studies.","keywords":["object compositing","image composition dataset","factual-counterfactual pairs","object removal","object insertion","diffusion model fine-tuning","real-world image capture","RAW image data"],"falsifier":"Compare each background-only image with the object-present images in regions away from the object: if pixel differences outside the object mask are consistently nonzero and grow with capture time, the counterfactual assumption fails. A direct version of this test is to place a static fiducial marker in each scene and measure whether its appearance drifts across the five captures.","tokens_in":12331,"feed_emoji":"📷","tokens_out":6245,"duration_ms":61285,"temperature":0.7,"pith_summary":"ORIDa is a proposed dataset for object compositing: placing an object into an image so it looks genuinely present. The paper claims it is the first large-scale, real-captured, publicly available dataset of this kind, with over 30,000 images of 200 objects, each seen in roughly 50 scenes and in up to four positions per scene. The central idea is that paired factual-counterfactual captures, the same scene with and without an object, let a model learn the full visual effect of an object, including shadows and reflections, without synthetic data. The authors show that a standard inpainting diffusion model fine-tuned on ORIDa alone outperforms prior methods for object removal and insertion in user studies and automatic metrics.","feed_headline":"30,000 real photos train AI to add and remove objects","feed_subtitle":"Real with-and-without-object pairs let diffusion models learn shadows and reflections without synthetic data.","key_machinery":"The factual-counterfactual (F-CF) capture protocol is the load-bearing mechanism. Each F-CF set consists of five consecutive images of one scene from a tripod: one background-only counterfactual and four factual images with the object in different positions, with shutter speed, ISO, white balance, and focus fixed. The background-only image serves as ground truth for object removal and as the target condition for insertion, and the four positions create training signal for repositioning and for learning object-to-scene effects such as shadows and reflections.","core_discovery":"The paper's central claim is that a real-captured dataset is sufficient to train photorealistic object compositing, provided the data are structured as factual-counterfactual sets: for each scene, one background-only image and four images with the object present, captured consecutively with a tripod and fixed camera settings. These sets expose object-to-scene effects, such as shadows and reflections, that synthetic compositing data cannot supply, while multiple positions per scene expose scene-to-object effects on the object's appearance. The paper reports that fine-tuning Stable Diffusion Inpainting on ORIDa, with only real COCO images added for the insertion task, beats Copy and Paste, Paint-by-Example, AnyDoor, and ObjectStitch on insertion, and beats SD-Inpaint, LaMa, and MGIE on removal, as measured by user preference and by metrics such as PSNR, DINO, CLIP, and LPIPS.","pith_inferences":["Read within its own stated scope, the dataset was collected for rigid, portable, non-human objects, so the demonstrated gains apply to that class until tested on deformable or living subjects.","The same capture protocol could be extended to video or multi-view capture, which would add temporal and geometric consistency cues beyond what still-image pairs provide.","A direct test of the dataset's value would be to train the same model on matched-scale synthetic pairs and compare shadow and reflection fidelity, since the paper's comparison to prior methods does not isolate this factor."],"forward_implications":["A model trained on ORIDa alone can remove an object and erase its shadows and reflections without a separate harmonization or shadow-removal stage.","Object insertion can be trained without synthetic compositing pipelines, with only real images such as COCO needed to support identity preservation.","Multiple positions per scene support object repositioning as a task, not just removal or insertion.","RAW DNG files plus ISP augmentations allow training across color and lighting variations from the same captured content."],"supporting_citations":[{"why":"Supplies the factual-counterfactual capture concept that ORIDa scales from one scene and one position to many.","marker":"[43]"},{"why":"Provides the real COCO images used by the paper to train object insertion while preserving object identity.","marker":"[22]"},{"why":"Generates the segmentation masks and bounding boxes used for ORIDa's localization annotations.","marker":"[30]"},{"why":"Is the latent diffusion model family that the paper's fine-tuned models build on.","marker":"[31]"},{"why":"Is the pretrained SD-Inpaint model that the paper fine-tunes for object removal and insertion.","marker":"[9]"},{"why":"Serves as a leading insertion baseline and synthetic-data training approach that ORIDa aims to replace.","marker":"[46]"},{"why":"Serves as a zero-shot insertion baseline compared in the user study.","marker":"[3]"},{"why":"Serves as an insertion baseline that trains on synthetic compositing data.","marker":"[38]"},{"why":"Serves as a removal baseline compared on shadow and reflection erasure.","marker":"[40]"},{"why":"Serves as an instruction-based editing baseline for the removal comparison.","marker":"[10]"}],"fun_headline_variants":["Real photos teach AI to insert objects with true shadows","200 objects, 30k shots: new dataset for object compositing","Real-captured pairs unlock object removal and insertion","ORIDa: 30k real images for object placement and removal"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The background-only image is a true counterfactual: between the background capture and the four object-present captures, nothing changes except the object's presence.","fun_headline_variants_meta":{"raw":{"variants":["Real photos teach AI to insert objects with true shadows","200 objects, 30k shots: new dataset for object compositing","Real-captured pairs unlock object removal and insertion","ORIDa: 30k real images for object placement and removal"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000175,"raw_usage":{"total_tokens":1275,"prompt_tokens":927,"completion_tokens":348,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":543,"completion_tokens_details":{"reasoning_tokens":277}},"tokens_in":543,"tokens_out":348,"duration_ms":3900,"temperature":1.0,"reasoning_tokens":277,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T04:57:23.302915+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare each background-only image with the object-present images in regions away from the object: if pixel differences outside the object mask are consistently nonzero and grow with capture time, the counterfactual assumption fails. A direct version of this test is to place a static fiducial marker in each scene and measure whether its appearance drifts across the five captures.","supporting_citations":[{"cited_title":"Microsoft coco: Common objects in context","cited_arxiv_id":null,"evidence_quote":"Provides the real COCO images used by the paper to train object insertion while preserving object identity."},{"cited_title":"High-resolution image synthesis with latent diffusion models","cited_arxiv_id":null,"evidence_quote":"Is the latent diffusion model family that the paper's fine-tuned models build on."},{"cited_title":"Inpainting with diffusers","cited_arxiv_id":null,"evidence_quote":"Is the pretrained SD-Inpaint model that the paper fine-tunes for object removal and insertion."},{"cited_title":"Anydoor: Zero-shot object-level im- age customization","cited_arxiv_id":null,"evidence_quote":"Serves as a zero-shot insertion baseline compared in the user study."},{"cited_title":"Object- stitch: Object compositing with diffusion model","cited_arxiv_id":null,"evidence_quote":"Serves as an insertion baseline that trains on synthetic compositing data."},{"cited_title":"Resolution-robust large mask inpainting with fourier convolutions","cited_arxiv_id":null,"evidence_quote":"Serves as a removal baseline compared on shadow and reflection erasure."}],"review_version":1}