{"id":"e0cad39e-5af2-4e56-a373-85eed4f096af","arxiv_id":"1908.10899","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A crowd-sourced dataset of 5,668 phone videos shot from upper-story windows improves activity classification on the VIRAT security-video benchmark by 8.3% when added to its training set.","lead":"The paper introduces a crowd-sourced video dataset called Out the Window, in which workers film themselves acting out everyday outdoor chores from a raised window view. Adding these videos to a standard security-video training set improved activity classification accuracy by 8.3%, suggesting a cheaper way to gather diverse training data for surveillance systems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 8.3% LOSO improvement is not trustworthy because the paper admits early stopping on the evaluation scene; re-run with a nested validation split before relying on the claim.","rationale":"The reader's verdict of CONDITIONAL is appropriate, but the single most load-bearing weakness is not the unvalidated annotation alignment; it is the evaluation protocol's admitted test-set early stopping. Even with perfect OTW annotations, the 8.3% improvement could be an artifact of selecting the best epoch on the evaluation scene. This is directly attested in Section 5 and is therefore a concrete, fixable flaw rather than a speculative mismatch. The annotation-alignment concern is real and would also need checking, but it is secondary because the observed transfer to VIRAT would be hard to explain if OTW labels were badly misaligned. The dataset contribution itself is potentially valuable and the LOSO idea is sound, so the paper should not be rejected; it should be conditional on a properly nested cross-validation and on reporting uncertainty. Hence I keep the verdict unchanged and partially agree with the reader's stated weakest assumption.","tokens_in":11008,"tokens_out":5662,"duration_ms":58095,"concrete_test":"Re-run the LOSO comparison with a nested validation split: in each of the five folds, hold out one of the four training scenes for early stopping and checkpoint selection, then evaluate the selected checkpoint on the left-out scene. Report per-fold accuracy and a bootstrap confidence interval (or paired significance test) on the mean VIRAT+OTW minus VIRAT-only difference. If the 8.3% advantage is not robust across folds or its confidence interval includes zero, the central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 5.3, Figure 8 right) is that adding OTW to VIRAT training yields an 8.3% mean classification accuracy gain under Leave-One-Scene-Out cross-validation. This number rests on an evaluation protocol that the paper itself flags as unfair: Section 5 states the TSN is trained 'terminating when loss on the validation set plateaued,' and in LOSO the 'validation set' is exactly the left-out scene whose accuracy is later reported. Selecting checkpoints on the test scene and then scoring on that same scene is a form of test-set overfitting. It can inflate both compared models, and because VIRAT-only and VIRAT+OTW have different training-set sizes and loss dynamics, the inflation can differ between them, making the 8.3% difference an artifact of checkpoint selection rather than a measure of OTW's contribution. The paper reports no confidence intervals, significance tests, or per-fold breakdowns, so neither the 8.3% nor the post-hoc 12.5% 'challenging activities' figure can be distinguished from noise. A secondary unvalidated premise is OTW/VIRAT label alignment (Section 4 only 'tried to keep these definitions close'), but the evaluation-protocol flaw alone is sufficient to undermine the headline number.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Out the Window (OTW) dataset, a crowdsourced security-video activity dataset containing 5,668 instances of 17 activities drawn from the VIRAT/ActEV taxonomy. Data is collected by instructing Amazon Mechanical Turk workers to lean a phone against an upper-story window and act out scenarios, yielding weakly trimmed clips with diverse scenes, actors, and viewpoints. Annotation combines expert keyframe labeling with Mask R-CNN detection and SORT tracking for interpolation. The authors evaluate a Temporal Segment Network on VIRAT Ground 2.0 under a default train/validation split and under Leave-One-Scene-Out cross-validation, reporting that adding OTW to VIRAT training improves mean classification accuracy by 8.3%, and by 12.5% on the most challenging activities. The dataset is released under CC BY 4.0.","tokens_in":11362,"tokens_out":3498,"duration_ms":35961,"significance":"If the claims hold, the OTW dataset and the ``scenario acting'' collection methodology are a meaningful contribution: they provide a security-specific activity dataset that is substantially larger and more diverse in scenes and actors than VIRAT, and the public release under CC BY 4.0 is valuable to the community. The LOSO evaluation protocol is a strong step toward controlling for scene bias, which is a well-known weakness of small security-video benchmarks. The main results, however, rest on an evaluation protocol the manuscript itself acknowledges as unfair, and the current statistical evidence is insufficient to separate the reported gains from noise or checkpoint-selection artifacts.","major_comments":[{"comment":"The early-stopping rule stated in Section 5 (``terminating when loss on the validation set plateaued'') is applied in LOSO cross-validation where, per Section 5.3, the left-out scene is declared to be the validation set for each fold. This means the same scene is used both to select the checkpoint and to report the final accuracy, which is a form of test-set overfitting. Because the VIRAT-only and VIRAT+OTW training sets differ in size and loss dynamics, the inflation from this selection can differ between the two compared models, so the reported 8.3% improvement may be an artifact of checkpoint selection rather than a measure of OTW's contribution. Please re-run the experiments with a nested validation split (e.g., hold out the test scene, use one of the training scenes for early stopping) and report the results obtained with checkpoints selected without any information from the left-out scene.","section":"Section 5 and Section 5.3"},{"comment":"The central comparison is reported only as a mean accuracy improvement over five LOSO folds, without per-fold numbers, error bars, confidence intervals, or significance tests. With only five folds, the 8.3% mean difference is not distinguishable from random variation. In addition, the 12.5% improvement is computed on a post-hoc subset of activities whose accuracy fell below 40%, selected from the same results that are being summarized, with no correction for multiple comparisons or a pre-specified definition of ``challenging.'' Please report the per-fold accuracy matrix, provide confidence intervals or a paired test over scenes or classes, and define the challenging subset before running the evaluation.","section":"Section 5.3, Figure 8"},{"comment":"The utility claim for OTW presupposes that OTW activity labels are semantically consistent with the VIRAT labels used for evaluation, but the manuscript only states that ``we tried to keep these definitions close'' and provides no inter-annotator agreement, no frame-level comparison of automatically interpolated boxes against human annotations, and no explicit reconciliation of the 17 OTW labels with their VIRAT counterparts. If the OTW labels are systematically noisier or semantically shifted, the reported transfer improvement could reflect label mismatch rather than learned activity structure. Please report annotation-quality statistics (e.g., temporal-boundary agreement, box IoU between interpolation and human re-annotation on a held-out sample) and describe the label alignment procedure in detail.","section":"Section 4"}],"minor_comments":[{"comment":"In the bullet on efficient annotation, ``automataed object detection'' should be ``automated object detection.''","section":"Section 1"},{"comment":"In the description of scenario acting, ``enteirng a car'' should be ``entering a car.''","section":"Section 3.1"},{"comment":"``classiﬁcation ac curacies'' should be ``classiﬁcation accuracies.''","section":"Section 5.3"},{"comment":"The caption states ``8 object label,'' but the table lists nine object types; this should be corrected to ``9 object labels.''","section":"Figure 6 caption"},{"comment":"The right panel of Figure 7 reports the default-split comparison described in Section 5.2, but the caption does not state this explicitly; clarify that Figure 7 uses the train/validation split from Section 5.1, while Figure 8 uses LOSO, to avoid confusion.","section":"Section 5.2 and Figure 7"},{"comment":"References [25] and [26] both describe Temporal Segment Networks; the text in Section 5 cites [26] while the related work section cites [25]. Please cite the ECCV version consistently for the architectural details and the TPAMI version for the journal extension, or otherwise disambiguate the two references in the text.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The dataset and collection methodology are potentially valuable, but the headline empirical claim needs to be re-established with a statistically valid evaluation. The nested-validation issue is the most serious, and the annotation-quality validation is also important for the utility claim. I would encourage the editor to request a revision rather than reject, since the dataset release itself is a contribution and the evaluation flaw is fixable within the manuscript's scope."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The Out the Window dataset is a real contribution, and the paper is worth reading for the collection methodology alone. The scenario-acting protocol—having Turkers film from an upstairs window while performing everyday outdoor activities—is a clever way to generate security-style videos at scale, and the public release under CC BY 4.0 makes it a useful resource for the community. The annotation pipeline with keyframe labeling and tracking is also sensible, and the paper is refreshingly honest about its limitations.\n\nThe headline result is another matter. The 8.3% improvement under Leave-One-Scene-Out is not trustworthy as reported. The paper states that training terminates when loss on the validation set plateaus, and in LOSO the validation set is exactly the held-out scene that is later used for evaluation. That is test-set overfitting, and the paper even acknowledges it is \"a bit unfair.\" Because the two compared models (VIRAT-only and VIRAT+OTW) have different training-set sizes and loss dynamics, the checkpoint selection can inflate one model more than the other, so the 8.3% difference could be an artifact of when training stopped rather than a measurement of OTW's contribution. There are no error bars, significance tests, or per-fold results, and the 12.5% \"challenging activities\" figure is a post-hoc subgroup, so it's hard to distinguish any of these numbers from noise.\n\nThe deeper premise—that OTW annotations are semantically aligned with VIRAT's activity labels—is also unvalidated. Section 4 only says they \"tried to keep these definitions close,\" and there is no inter-annotator agreement or check on the tracking-based interpolation. That matters, because if the labels are systematically different, the apparent gain could reflect label mismatch rather than learned activity structure.\n\nNone of this sinks the dataset. The central contribution is the collection method and the data itself, and that holds up. But the quantitative evidence for its utility needs rework: a nested validation split for checkpoint selection, per-fold accuracy, and ideally a small annotation agreement study. With those, the paper would be much stronger.\n\nThis is a paper I'd send to a serious referee. The dataset deserves a place in the literature, and the evaluation flaws are fixable. I'd want the revision to address the data-snooping issue head-on before accepting the headline claim.","headline":"The OTW dataset is a genuinely useful crowdsourced collection for security video, but the headline 8.3% LOSO improvement is undermined by early stopping on the evaluation scene and missing uncertainty estimates.","tokens_in":11783,"tokens_out":3334,"would_cite":true,"duration_ms":28013,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that adding a crowdsourced dataset of window-view phone videos to standard surveillance training data improves activity classification accuracy by 8.3 percent, and by 12.5 percent on the hardest classes.","keywords":["crowdsourced dataset","activity classification","security video","scenario acting","surveillance video","video annotation","transfer learning","Temporal Segment Network"],"falsifier":"Take a random sample of OTW clips, have independent annotators re-label them using the benchmark's exact activity definitions, and measure frame-level agreement with the released annotations; low agreement would show the improvement may be label noise. A complementary check is to shuffle OTW activity labels while keeping scenes fixed: if the accuracy gain survives shuffling, it comes from scene and viewpoint diversity rather than from activity semantics.","tokens_in":10797,"feed_emoji":"🎥","tokens_out":8611,"duration_ms":84014,"temperature":0.7,"pith_summary":"This paper introduces Out the Window (OTW), a crowdsourced dataset of 5,668 instances of 17 security-video activities, and claims that adding it to standard surveillance training data materially improves activity classification on novel scenes. The collection method asks online workers to lean a phone against an upper-story window and act out everyday scenarios, producing natural videos with far more actor, viewpoint, and scene variety than studio-style security collections. Under leave-one-scene-out evaluation on the VIRAT benchmark, training with OTW plus VIRAT beats training on VIRAT alone by 8.3 percent mean classification accuracy, and by 12.5 percent on the activities whose baseline accuracy is below 40 percent. The reason to care is that security-video models currently generalize poorly because their training data comes from a few scenes and actors; a cheap, scalable source of diverse security-style video would address that bottleneck.","feed_headline":"Crowdsourced window videos lift security-video activity scores by 8.3%","feed_subtitle":"Phone clips filmed out a window add enough scene and actor variety to improve hardest-class accuracy by 12.5%.","key_machinery":"The mechanism that carries the argument is 'scenario acting': each crowd worker receives a broad everyday errand, such as going grocery shopping or going on vacation, and performs it in their own yard or driveway while their phone records from a window. Because workers are never told which discrete activities matter, the videos contain natural, varied instances of the 17 target classes without the stiffness of scripted acting. The annotation pipeline then keeps cost low: analysts mark only the first and last frame of each activity and its participating objects, and a detector plus tracker propagates boxes through the intermediate frames, with the activity bounding box defined as the convex union of object boxes. Evaluation uses a Temporal Segment Network classifier and leave-one-scene-out cross-validation, so the reported gains measure generalization to actors, viewpoints, and scenes absent from training rather than memorization of scene context.","core_discovery":"The paper's central claim is that a crowd-sourced collection protocol can supply the diversity that security-video activity models lack, and that the resulting dataset demonstrably improves a standard classifier on a standard benchmark. On the VIRAT Ground 2.0 benchmark, using the OTW dataset together with VIRAT training data raises mean classification accuracy from the VIRAT-only baseline by 8.3 percent under leave-one-scene-out cross-validation, where each model is tested on a scene it never saw in training. For the most difficult activities, those with baseline accuracies below 40 percent, the average improvement is 12.5 percent. The evaluation is closed-set classification of trimmed clips, so the reported gain is in recognizing an activity once it has been localized, not in detecting it in untrimmed video.","pith_inferences":["If the label-alignment assumption holds, the same scenario-acting recipe could be pointed at other rare or underrepresented security activities, including ones that are almost absent from social video, without studio-scale costs.","The 15 percent gap the paper notes between OTW evaluation and VIRAT leave-one-scene-out evaluation suggests that factors other than viewpoint and actor diversity, such as camera resolution and scene clutter, still separate crowdsourced phone video from true surveillance footage; that gap is a testable target for ablation.","The symmetric activity pairs the paper highlights (opening/closing, entering/exiting, loading/unloading) are hard because they differ mainly in temporal order, so the OTW benefit might be explained partly by restoring temporal-direction cues; flipping frame order in the OTW clips would test that."],"forward_implications":["OTW data plus VIRAT training improves mean accuracy by 8.3 percent over VIRAT-only training under leave-one-scene-out evaluation.","The hardest activities, those scoring below 40 percent on the baseline, improve by 12.5 percent on average.","The collection produced over 200 examples for every VIRAT-relevant class except pulling, riding, and cell-phone activities, cutting into the class imbalance that dominates security-video benchmarks.","Sparse scenes and weakly trimmed clips make tracking-based annotation feasible, keeping human annotation to a few keyframes per activity while still producing spatiotemporal boxes.","Because the gains appear under leave-one-scene-out evaluation, the benefit transfers to scenes the model has not seen, which is the setting that matters for deployed security networks."],"supporting_citations":[{"why":"Supplies the benchmark dataset, the activity definitions, and the scene split used for evaluation.","marker":"[19]"},{"why":"Supplies the classifier architecture and training settings used in every comparison run.","marker":"[26]"},{"why":"Supplies the crowdsourced scripted-home-video collection model that the protocol adapts to outdoor security views.","marker":"[20]"},{"why":"Supplies the object detector used to propagate bounding boxes between keyframes.","marker":"[12]"},{"why":"Supplies the online tracker used to link detected objects across frames in annotation.","marker":"[2]"},{"why":"Supplies the pretrained model that initializes the classifier before fine-tuning.","marker":"[15]"},{"why":"Supplies the annotation tool used to record keyframe boxes.","marker":"[5]"}],"fun_headline_variants":["Window-shot crowdsourced clips boost activity recognition by 8.3%","Phone-from-window videos improve security activity scores 8.3%","Crowdsourced window videos add 8.3% to VIRAT activity accuracy","Window videos from crowds raise hardest activity scores 12.5%","Crowd-shot window clips lift tough VIRAT activities 12.5%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that OTW's crowd annotations are accurate and mean the same thing as the benchmark's activity definitions; if the labels are systematically wrong or systematically different, the 8.3 percent gain could be an artifact of label mismatch rather than learned activity structure.","fun_headline_variants_meta":{"raw":{"variants":["Window-shot crowdsourced clips boost activity recognition by 8.3%","Phone-from-window videos improve security activity scores 8.3%","Crowdsourced window videos add 8.3% to VIRAT activity accuracy","Window videos from crowds raise hardest activity scores 12.5%","Crowd-shot window clips lift tough VIRAT activities 12.5%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000308,"raw_usage":{"total_tokens":1712,"prompt_tokens":848,"completion_tokens":864,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":464,"completion_tokens_details":{"reasoning_tokens":762}},"tokens_in":464,"tokens_out":864,"duration_ms":7839,"temperature":1.0,"reasoning_tokens":762,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T10:30:44.917008+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a random sample of OTW clips, have independent annotators re-label them using the benchmark's exact activity definitions, and measure frame-level agreement with the released annotations; low agreement would show the improvement may be label noise. A complementary check is to shuffle OTW activity labels while keeping scenes fixed: if the accuracy gain survives shuffling, it comes from scene and viewpoint diversity rather than from activity semantics.","supporting_citations":[{"cited_title":"T., M UKHERJEE , S., A GGARWAL , J., L EE, H., D AVIS, L., ET AL","cited_arxiv_id":null,"evidence_quote":"Supplies the benchmark dataset, the activity definitions, and the scene split used for evaluation."},{"cited_title":"Temporal segment net- works: Towards good practices for deep action recogni- tion","cited_arxiv_id":null,"evidence_quote":"Supplies the classifier architecture and training settings used in every comparison run."},{"cited_title":"A., V AROL , G., W ANG , X., F ARHADI , A., L APTEV , I., AND GUPTA, A","cited_arxiv_id":null,"evidence_quote":"Supplies the crowdsourced scripted-home-video collection model that the protocol adapts to outdoor security views."},{"cited_title":"Mask r-cnn","cited_arxiv_id":null,"evidence_quote":"Supplies the object detector used to propagate bounding boxes between keyframes."},{"cited_title":"Simple online and realtime tracking","cited_arxiv_id":null,"evidence_quote":"Supplies the online tracker used to link detected objects across frames in annotation."},{"cited_title":"VGG image annotator (VIA)","cited_arxiv_id":null,"evidence_quote":"Supplies the annotation tool used to record keyframe boxes."}],"review_version":1}