{"id":"a3040ae1-8714-48b6-8a26-e5c3940e1b29","arxiv_id":"2507.07519","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"MUVOD provides multi-view video panoptic masks for 17 dynamic scenes (459 instances, 73 categories) and benchmarks for multi-view video and 3D object segmentation.","lead":"MUVOD is a new multi-view video dataset for 4D object segmentation, with 17 scenes from diverse camera rigs and panoptic masks tracked across time and views. It also defines a multi-view video object segmentation task with an evaluation metric and a 3D segmentation benchmark, including baseline results.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Ground-truth reliability is the load-bearing assumption: GT masks are generated by the same XMem propagation used as the benchmark baseline, no independent mask-quality check is reported, and the abstract's image count conflicts with the stated annotation protocol.","rationale":"The reader's weakest assumption is close to this, but not identical: the reader emphasizes missing quantitative quality checks; the sharper problem is the overlap between the annotation generator and the evaluated baseline. Even an internal visual inspection by the dataset authors may not break the circularity; only independent re-annotation can. The internal count inconsistencies (abstract 7830 vs. the 3-views-per-scene protocol implying 1530, Section I's 21 frames vs. Section III.C's description, and the Table VI vs. Table VIII caption mismatch for the 3D benchmark) make the reported statistics unreliable and strengthen the need for release-level auditing rather than acceptance purely on paper claims. This does not require rejecting the dataset idea; the paper can be accepted conditionally on releasing data, scripts, and a ground-truth validation protocol. I agree partially with the reader: same general concern, but my emphasis is the circularity and the count conflicts rather than only the absence of manual verification.","tokens_in":21443,"tokens_out":9337,"duration_ms":99350,"concrete_test":"Download the released MUVOD masks and count the actual number of annotated RGB images per scene and per view; reconcile the count with the abstract's 7830 and Section I's stated 21 frames. Then select a random stratified sample of roughly 100 non-keyframe frames from non-initial views, have two independent annotators re-segment all labelled objects from scratch, and compute mean IoU between MUVOD GT and each re-annotation. Rerun the Section IV.C baseline against the corrected re-annotated GT. If mean IoU falls materially below 0.90, or if the baseline J&F changes by more than 3 points, or if per-scene rankings change, the propagated-GT concern is confirmed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central contribution of MUVOD is a set of ground-truth masks that are spatiotemporally consistent and dense enough to train or evaluate 4D segmentation. The weakest link is that these masks are not independently verified. In Section III.C, non-initial camera keyframes and all intermediate frames are produced by propagating manually refined cini keyframes with XMem/SAM; the baseline in Section IV.C is likewise \"applying XMem [57] both spatially and temporally.\" Evaluating a propagation model against labels produced by the same propagation model is circular: systematic XMem errors (label switches, boundary drift, identity loss across wide-baseline views, cf. Fig. 7) can be embedded in the \"ground truth,\" inflating the Table III/IV scores and biasing comparisons in the Section V 3D benchmark. No inter-annotator agreement or error-rate measurement is reported. The reported scale is additionally inconsistent: the abstract says 7830 RGB images (30 frames per video), but the protocol annotates three views per scene, giving 17×3×30 = 1530 annotated frames, while Section I says 21 manually refined frames per view. Thus the dataset's central value as accurate ground truth is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents MUVOD, a multi-view video dataset of 17 realistic scenes with panoptic-style instance annotations in space and time, a semi-automatic annotation pipeline (manual keyframes, SAM, XMem spatial/temporal propagation, LightGlue geometric cues), an evaluation metric J&F_N for multi-view video object segmentation, and a baseline that applies XMem across views and time. It also introduces a 3D object segmentation benchmark with 50 target objects from 12 scenes and evaluates ISRF, SA3D, SAGA, and Gaussian Grouping. The central claims are that MUVOD is a large, diverse, richly annotated multi-view video dataset and that the proposed benchmark enables meaningful evaluation of 4D/3D object segmentation.","tokens_in":21674,"tokens_out":10093,"duration_ms":105259,"significance":"The dataset addresses a real gap: prior multi-view video datasets are limited to specific domains such as street scenes, furniture, or human-only scenes. The inclusion of 17 scenes from four sources with 9 to 46 views, 459 instances, and 73 categories, together with a defined task, a clearly stated metric, a baseline, and a 3D benchmark with four state-of-the-art methods, is a potentially useful community resource. The paper provides a dataset URL and the benchmark is a concrete contribution. The paper is not a theoretical derivation; its value depends on whether the released ground-truth masks are accurate and consistent. That dependency is currently unmet, so the significance is conditional on the requested validation and on reconciling the reported dataset scale.","major_comments":[{"comment":"The headline dataset scale is internally inconsistent. The abstract states that MUVOD provides 7830 RGB images (30 frames per video), but Section I says only three camera views per scene are annotated and that each view contains 21 frames with manually refined ground-truth masks, and Section III.C describes sampling keyframes every 10 frames from a single initial camera. Under the stated protocol, 17 scenes times 3 views times 21 frames gives 1071 manually refined frames, and even taking 30 frames per video gives 1530, not 7830. The number 7830 is close to 375 available views times 21 frames, suggesting that all views rather than three are annotated, which conflicts with Section I. Please reconcile the count and specify, for each released mask, whether it is manually refined, spatially propagated, or temporally propagated.","section":"Abstract; Section I; Section III.C"},{"comment":"Ground-truth mask accuracy is not validated, and the baseline evaluation is at risk of self-agreement. In the annotation pipeline, only keyframes of the initial camera cini are manually refined; non-initial camera keyframes and all intermediate frames are produced by propagating masks with XMem and SAM. The baseline method in Section IV.C is exactly XMem applied spatially and temporally. Systematic XMem errors, such as label switches or boundary drift, can therefore be embedded in the 'ground truth' and inflate the scores reported in Tables III and IV; Fig. 7 documents that masks can be applied to nearby objects of the same class, and no evidence is given that the ground-truth column is not affected by the same failure mode. No inter-annotator agreement, per-frame error rate, or independent quality check is reported. Please add a quantitative validation study, for example human correction rates on a random sample of frames or a comparison against independently produced masks, and report how much of the baseline score is self-agreement with the automatically propagated labels.","section":"Section III.C; Section IV.C; Fig. 7"},{"comment":"The 3D object segmentation benchmark's per-condition conclusions rest on very few objects. Table VI shows that occluded, small-scale, and complex-structure categories each contain only 3 objects, and Table VIII reports mean IoU for these conditions without variance or per-object results. With n=3 per condition, the paper's claim of a 'more comprehensive analysis' of method strengths and limitations is stronger than the evidence supports. The final paragraph of Section V.C acknowledges the small subset, but the quantitative tables should also report per-object IoU or confidence intervals, and the conclusions should be correspondingly qualified.","section":"Section V.B; Section V.C; Table VIII"}],"minor_comments":[{"comment":"The first sentence has a subject-verb agreement error: 'The application of methods ... have steadily gained popularity' should be 'has steadily gained popularity'.","section":"Abstract"},{"comment":"The caption contains a typo: 'SPATIO-TEMPORAL COSISTENCE' should read 'CONSISTENCE'.","section":"Table I"},{"comment":"The name of the 3D segmentation method by Ye et al. is rendered inconsistently as 'LeRF-Mask' in Section V.A and 'LERF-Mask' in Section II.B; please use one spelling throughout.","section":"Section II.B; Section V.A"},{"comment":"The axis labels and tick labels in Figure 5 appear as garbled placeholder glyphs in the manuscript, making the figure difficult to read; please regenerate it with standard fonts.","section":"Figure 5"},{"comment":"The benchmark description does not mention whether the evaluation prompts (points or strokes) are provided under the same protocol across the four methods, which matters for fairness; please specify the prompting procedure.","section":"Section V.B"}],"recommendation":"major_revision","confidential_remarks":"The paper's core problem is not the idea but the evidence: the reported scale cannot be derived from the described protocol, and the ground-truth masks are produced by the same propagation model used as the baseline, without independent verification. If the authors can supply a frame-level validation study and correct the scale statement, the dataset could be a solid contribution; if not, the resource's reliability as ground truth remains unresolved. The fit with the journal is appropriate, and I did not find citation or circularity problems of a different kind."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: MUVOD is a genuinely useful resource in an area that lacks one, but the paper as written doesn't establish the reliability of its own ground truth. The abstract's image count is inconsistent with the annotation protocol, and the baseline shares its propagation engine with the label generation. Both need fixing before the dataset numbers can be trusted.\n\nThe new thing here is real: 17 multi-view dynamic scenes, 459 instances across 73 categories, with panoptic masks that are meant to stay consistent across camera views and over time. That's not present in the datasets they compare against. The 3D segmentation benchmark, with 50 objects in four difficulty classes, is a smaller but still useful addition. The annotation pipeline is thoughtful: manual keyframes, SAM, XMem propagation, depth-layered occlusion handling, and 3D geometric cues for cross-view consistency. The evaluation metric is a sensible multi-view extension of J&F, and the baseline (spatial+temporal XMem) is a reasonable first reference.\n\nThe soft spots are real. First, the numbers don't add up. The abstract says 7830 RGB images at 30 frames per video, but the protocol annotates three views per scene and only 21 manually refined frames per view. 17×3×30 is 1530, not 7830. Either the abstract is wrong or the protocol description is missing something. Second, and more important, the ground-truth masks are produced by SAM and XMem propagation, and the baseline is XMem applied both spatially and temporally. Evaluating XMem against labels that XMem itself helped generate tests the model against its own output. Systematic errors—label switches, boundary drift, identity loss across wide-baseline views—will be baked into the 'ground truth' and inflate the Tables III/IV scores. The paper reports no inter-annotator agreement or error-rate measurement on a manually verified subset. Figure 7 actually shows baseline confusion on similar objects; that's fine as a baseline failure, but it also signals where the propagation-based GT could be wrong in the same places.\n\nThese are fixable. Add a validation subset where humans correct the propagated masks and measure the difference; report the corrected IoU; and clarify the dataset statistics. The 3D benchmark is small and excludes fisheye scenes, but that's acceptable given the constraints.\n\nWho is this for? People working on 4D segmentation, dynamic NeRF/3DGS, and video object segmentation. They'll want this dataset even with caveats. It deserves a serious referee; I'd send it out, but expect major revision. The dataset contribution is worth engaging with, and the quality evidence needs to be supplied before the benchmark numbers are used.","headline":"MUVOD fills a real gap in multi-view video segmentation, but the paper must fix inconsistent image counts and show the XMem-based ground truth is not circularly inflating its own baseline.","tokens_in":22188,"tokens_out":3768,"would_cite":false,"duration_ms":37822,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MUVOD is a multi-view video dataset whose panoptic masks keep 459 object instances identifiable across time and across camera views.","keywords":["multi-view video dataset","4D object segmentation","panoptic segmentation","video object segmentation","3D segmentation benchmark","semi-automatic annotation","dynamic scenes"],"falsifier":"Manually re-segment a random sample of frames from multiple annotated views, then compute the mean IoU between those fresh manual masks and the dataset's propagated masks; if that agreement is below roughly 90% IoU, the propagated masks are too unreliable to support the benchmark comparisons.","tokens_in":21279,"feed_emoji":"🎥","tokens_out":14728,"duration_ms":146425,"temperature":0.7,"pith_summary":"MUVOD is built to fill a gap: dynamic real-world scenes with many interacting objects have had no richly annotated multi-view video resource for 4D object segmentation. The paper assembles 17 scenes from synchronized camera rigs with 9 to 46 views each, reporting 7,830 RGB frames and panoptic masks, a format in which every pixel is assigned either a unique instance ID or a background class. The masks are designed to be spatio-temporally coherent, so a given object keeps the same identity across frames of one view and across different views of the same rig. The paper also defines a multi-view video object segmentation task with an averaged J&F metric, reports a baseline on it, and turns a 50-object subset into a 3D segmentation benchmark with four difficulty classes. If the annotations are accurate, this gives the NeRF and 3D Gaussian Splatting communities a training and evaluation ground for dynamic multi-object scenes.","feed_headline":"MUVOD gives 459 objects stable masks across time and views","feed_subtitle":"Panoptic labels follow every object through frames and up to 46 camera views per scene.","key_machinery":"The load-bearing mechanism is a spatio-temporal annotation pipeline that turns sparse manual input into dense ground-truth masks: for the initial camera, keyframes are manually boxed and masks are produced by a promptable segmentation foundation model, then refined; those masks are propagated to neighboring cameras with a long-term video object segmentation model, with view-frustum similarity, triangulated 3D point clouds, and projected-point prompts providing geometric guidance; bidirectional temporal propagation fills the remaining frames. A depth-layer ordering for static and environmental objects lets the pipeline treat occlusion deterministically by overlaying masks from farthest to nearest. On the evaluation side, the paper's metric averages the standard 2D video object segmentation score, the mean of region similarity $J$ and contour accuracy $F$, over the $N$ cameras being used: $J\\&F^N = \\frac{1}{N} \\sum_{i=0}^{N-1} (J\\&F)_{c_i}$.","core_discovery":"The paper's central claim is that MUVOD provides a resource that combines synchronized multi-view video with dense, instance-level, view-consistent annotations in real dynamic scenes. It contains 459 instances across 73 categories, with each object assigned a semantic label, a unique instance ID, and a motion status of dynamic, static, or environmental; dynamic and static objects are treated as things, while environmental objects are treated as stuff. The annotations are produced by a semi-automatic pipeline that manually refines keyframes of one initial camera and propagates masks to other views and other frames, using depth layers to resolve occlusions deterministically. On top of the dataset, the paper defines a multi-view video object segmentation task and a camera-averaged metric, and it reports a baseline reaching 79.4% J&F on three views averaged over all scenes. The paper further presents a 3D segmentation benchmark of 50 objects and evaluates four state-of-the-art methods on it, with the best method averaging 78.8% IoU across the four object difficulty classes.","pith_inferences":["Beyond the paper, a human-agreement study on a held-out set of frames would test the central accuracy assumption, since the pipeline only refines keyframes of one camera.","Beyond the paper, the depth-layer representation suggests a testable variant: estimate depth ordering automatically and measure how much annotation time and mask accuracy change.","Beyond the paper, the sparse-view experiment implies that camera-rig geometry is itself a performance variable, so future benchmarks might report results as a function of view density rather than for a single rig configuration."],"forward_implications":["4D object segmentation methods can be trained and compared on real dynamic scenes with many objects instead of isolated humans or street views.","The camera-averaged J&F metric gives the field a common yardstick for spatio-temporal mask propagation across viewpoints.","The 3D segmentation benchmark exposes where current lifting methods lose accuracy, particularly on occluded, small-scale, and complex-structure objects.","Downstream applications such as 4D scene editing and VR/AR interaction become testable because object identity is already linked across space and time."],"supporting_citations":[{"why":"Supplies the long-term video object segmentation model used for spatial and temporal mask propagation in the annotation pipeline and as the task baseline.","marker":"[57]"},{"why":"Supplies the promptable segmentation model that generates initial object masks from bounding-box prompts and projected 3D point prompts.","marker":"[82]"},{"why":"Supplies the base 124-class taxonomy that MUVOD extends and revises into its 73 annotated categories.","marker":"[81]"},{"why":"SemanticKITTI is one of the existing panoptic multi-view datasets compared with MUVOD, showing a street-only class gap.","marker":"[30]"},{"why":"KITTI-360 is the other street-only panoptic comparison dataset used to motivate broader scene coverage.","marker":"[31]"},{"why":"Gaussian Grouping is the top-performing method on the proposed 3D segmentation benchmark and also serves as a prior multi-object benchmark reference.","marker":"[14]"},{"why":"ISRF is one of the four state-of-the-art methods evaluated on the new 3D benchmark, providing a comparison baseline.","marker":"[15]"},{"why":"SA3D is one of the four evaluated methods and achieves the second-best average on the 3D benchmark.","marker":"[16]"},{"why":"SAGA is one of the four evaluated methods, showing large degradation on the harder object categories.","marker":"[17]"}],"fun_headline_variants":["MUVOD: 459 objects tracked across views and time","New dataset links multi-view video to 4D masks","Benchmark for consistent segmentation in dynamic scenes","MUVOD: view-consistent masks for up to 46 cameras","Track 459 object instances through multi-view video"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The propagated masks are accurate enough to count as ground truth even though only keyframes of one camera are manually refined.","fun_headline_variants_meta":{"raw":{"variants":["MUVOD: 459 objects tracked across views and time","New dataset links multi-view video to 4D masks","Benchmark for consistent segmentation in dynamic scenes","MUVOD: view-consistent masks for up to 46 cameras","Track 459 object instances through multi-view video"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000317,"raw_usage":{"total_tokens":1868,"prompt_tokens":1093,"completion_tokens":775,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":709,"completion_tokens_details":{"reasoning_tokens":695}},"tokens_in":709,"tokens_out":775,"duration_ms":8201,"temperature":1.0,"reasoning_tokens":695,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:38:36.074849+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Manually re-segment a random sample of frames from multiple annotated views, then compute the mean IoU between those fresh manual masks and the dataset's propagated masks; if that agreement is below roughly 90% IoU, the propagated masks are too unreliable to support the benchmark comparisons.","supporting_citations":[{"cited_title":"Xmem: Long-term video object segmentation with an atkinson-shiffrin memory model,","cited_arxiv_id":null,"evidence_quote":"Supplies the long-term video object segmentation model used for spatial and temporal mask propagation in the annotation pipeline and as the task baseline."},{"cited_title":"Large-scale video panoptic segmentation in the wild: A benchmark,","cited_arxiv_id":null,"evidence_quote":"Supplies the base 124-class taxonomy that MUVOD extends and revises into its 73 annotated categories."},{"cited_title":"SemanticKITTI: A Dataset for Semantic Scene Under- standing of LiDAR Sequences,","cited_arxiv_id":null,"evidence_quote":"SemanticKITTI is one of the existing panoptic multi-view datasets compared with MUVOD, showing a street-only class gap."},{"cited_title":"Kitti-360: A novel dataset and bench- marks for urban scene understanding in 2d and 3d,","cited_arxiv_id":null,"evidence_quote":"KITTI-360 is the other street-only panoptic comparison dataset used to motivate broader scene coverage."},{"cited_title":"Interactive segmen- tation of radiance fields,","cited_arxiv_id":null,"evidence_quote":"ISRF is one of the four state-of-the-art methods evaluated on the new 3D benchmark, providing a comparison baseline."},{"cited_title":"Segment anything in 3d with nerfs,","cited_arxiv_id":null,"evidence_quote":"SA3D is one of the four evaluated methods and achieves the second-best average on the 3D benchmark."}],"review_version":1}