{"id":"5fa2176f-089a-42d8-ab7b-356406feef69","arxiv_id":"2502.06681","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A benchmark dataset with 963,554 bounding boxes of 22 people across seven cameras over seven months shows current Re-ID models drop to 18.81% CMC@1 on long-term indoor re-identification.","lead":"CHIRLA is a new video dataset for long-term person re-identification, recorded over seven months in an indoor office with seven synchronized cameras. It provides about one million labeled bounding boxes for 22 people and benchmark protocols for tracking and re-identification under clothing changes.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Identity-label reliability is unvalidated, and the 'fully annotated' claim conflicts with the paper's own exclusion of unidentifiable people; until an independent label-consistency check is reported, the long-term benchmark numbers are not trustworthy.","rationale":"The reader's weakest assumption is that identity labels are correct across the seven-month capture, and my analysis agrees that this is the most load-bearing condition. The central claim of a usable long-term, video-based indoor Re-ID benchmark depends entirely on whether the same ID across sequences denotes the same person. The paper provides no quantitative label-quality evidence, only a description of the semi-automatic pipeline and manual review. This is not an accusation of fraud; it is an ordinary, testable requirement for any dataset paper. In addition, the manuscript contains an internal inconsistency: it calls the dataset 'fully annotated' while also stating that labels are removed when a person cannot be sufficiently identified. That inconsistency further weakens the central claim and would need to be corrected even if identity labels are otherwise reliable. Because this concern is already reflected in the reader's CONDITIONAL verdict, I do not propose a different verdict; the paper should be accepted only after the authors provide an independent label-consistency check and revise the 'fully annotated' wording.","tokens_in":19692,"tokens_out":4366,"duration_ms":44915,"concrete_test":"Sample approximately 500 tracklets across sequences and cameras, stratified by ID, containing both face and body crops. Have at least two independent annotators, blind to the existing CHIRLA labels, link each tracklet to an identity using face and body appearance; compute Fleiss' kappa and the per-ID label disagreement rate relative to the released annotations. Independently, run a face-recognition model (e.g., ArcFace) on the same face crops to measure cross-sequence face-feature consistency for each released ID. If annotator agreement is below 0.9 or label disagreement exceeds 1-2%, the long-term and multi-camera benchmarks must be re-validated, and the 'fully annotated' claim must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that CHIRLA's identity labels are correct across seven months, ten sequences, and seven cameras. The Procedure section states that IDs were produced by YOLOv8 detections, Deep SORT tracklets, and manual review in a custom GUI, but the paper reports no inter-annotator agreement, no label error rate, and no independent check that the same ID across sequences refers to the same person. Deep SORT is known to make association errors, especially across long gaps, and manual correction of roughly one million bounding boxes is not verifiable from the description alone. If identity labels are inconsistent across months, the reported long-term result (18.81% CMC@1) describes annotation noise rather than appearance change, and the benchmark's central value collapses. Additionally, the manuscript itself says 'individuals are labeled only if they can be sufficiently identified; otherwise, their labels are removed,' which directly contradicts the abstract's 'fully annotated' claim and makes tracking ground truth incomplete, inflating false-negative and false-positive counts in Table 3. Both issues are load-bearing, with identity-label consistency being the more fundamental one.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CHIRLA, a video-based person re-identification dataset recorded over seven months in an indoor office environment using seven synchronized cameras. It contains 10 sequences, 22 individuals, roughly 963,000 annotated bounding boxes, and semi-automatic identity labels produced with YOLOv8 detections, Deep SORT tracklets, and manual GUI review. The authors define separate benchmark protocols for person tracking (brief occlusions and multiple-person occlusions) and for Re-ID (reappearance, long-term, multi-camera, and multi-camera long-term), and they report extensive evaluations with trackers and CNN/Transformer Re-ID models. The main empirical claims are that the dataset is challenging, with the best long-term Re-ID result at 18.81% CMC@1, and that current appearance-based trackers degrade substantially in occlusion-focused scenarios.","tokens_in":19941,"tokens_out":3303,"duration_ms":33708,"significance":"If the identity labels are reliable, CHIRLA fills a real gap: it is a long-term, video-based, indoor Re-ID dataset with temporal coherence, clothing variation, multi-camera views, and roughly 44,000 bounding boxes per identity, which is substantially richer per ID than most existing clothing-change benchmarks. The public data and code, synchronized capture, ethical consent, and the breadth of evaluated models are concrete strengths. The low CMC and MOTA values support the paper's difficulty claim qualitatively. However, the central value of the dataset rests on the correctness of identity labels over seven months, and the manuscript provides no quantitative annotation-quality evidence for that assumption. The benchmark conclusions can therefore only be provisionally accepted until label consistency and annotation completeness are demonstrated.","major_comments":[{"comment":"The load-bearing assumption that the same ID refers to the same person across all seven months, ten sequences, and seven cameras is not validated. The Procedure section states that IDs were produced by YOLOv8 detections, Deep SORT tracklets, and manual review in a custom GUI, but no inter-annotator agreement, label error rate, or independent cross-sequence identity check is reported. Because the long-term Re-ID numbers in Table 4 (e.g., 18.81% CMC@1 for ResNet101) are exactly the quantities most sensitive to cross-month identity confusion, annotation noise could masquerade as appearance change. I ask the authors to add a quantitative validation: for example, re-annotate a random sample of frames by at least two annotators and report per-sequence and per-camera ID agreement, and/or run an independent face or whole-body verification check on cross-sequence ID pairs, with the resulting error rates and any corrected labels documented.","section":"Methods, Procedure"},{"comment":"The manuscript's claim that CHIRLA is 'fully annotated' is contradicted by the sentence in the Procedure section: 'individuals are labeled only if they can be sufficiently identified; otherwise, their labels are removed.' Removing unidentifiable individuals means the ground truth is incomplete, which directly affects the tracking metrics in Table 3: detections of unlabeled people are counted as false positives, and missing ground truth inflates false negatives and distorts MOTA and IDF1. The paper should report how many frames and bounding boxes were removed for this reason, per sequence and camera, quantify the resulting impact on the metrics, and revise the 'fully annotated' claim to a precise statement of annotation coverage.","section":"Methods, Procedure and Table 3"},{"comment":"The introduction lists 'Open-set evaluation with distractors' as a unique feature of CHIRLA, but the experimental section states that 'We restrict the experiments to closed-set evaluation... distractors (unknown IDs) provided by CHIRLA are excluded from the test set. We leave the open-set evaluation for future studies.' As a result, the claimed open-set protocol is not actually evaluated or validated anywhere in the paper. The authors should either provide an open-set evaluation with appropriate threshold-based metrics (e.g., identification rate at a controlled false-alarm rate) or remove the open-set claim from the dataset's advertised contributions.","section":"Technical Validation, Re-ID Benchmark and Table 4"},{"comment":"The benchmark results are averaged over manually selected train/test subsets, but no variance or per-subset statistics are reported. The selection procedure is described as 'prioritizing challenging scenarios,' which introduces a potential selection bias, and with only 7-10 subsets for the Re-ID scenarios, the averaged CMC and mAP values in Table 4 may not be stable enough to support fine-grained comparisons between models. I request per-subset results, standard deviations, or confidence intervals for at least the headline numbers (long-term and multi-camera long-term), so readers can judge whether observed differences are meaningful.","section":"Benchmark Data Preparation and Table 4"}],"minor_comments":[{"comment":"The Subject IDs column includes values 24, 25, and 26 while the text states the dataset has 22 individuals; the authors should clarify the ID numbering scheme and which IDs are unused or reserved.","section":"Data Records, Table 2"},{"comment":"There are several typos and formatting issues: 'CA VIAR' should be 'CAVIAR', 'A VI' should be 'AVI', and 'Video' entries for CHIRLA should be consistent with the '70 videos' count mentioned in the table caption.","section":"Table 1 and Data Records"},{"comment":"Equation (6) contains the typo 'top-krank' and the CMC definition would be clearer if it explicitly stated that the indicator is evaluated over a single ground-truth match per query, as the surrounding text already notes.","section":"Evaluation Metrics"},{"comment":"The sentence 'we present well-defined benchmarks, to rigorously tested our dataset' is grammatically incomplete and should be rewritten.","section":"Technical Validation"},{"comment":"The exclusion of train_0 and test_0 from evaluation is stated only in the main text; this important protocol detail should also appear in the table caption.","section":"Table 4 caption"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a resource paper, so the absence of a derivation is not a concern. The key risk is annotation quality, not circularity: if the identity labels are inconsistent across the seven-month capture, the headline long-term CMC numbers would describe annotation noise. This is fixable within the scope of the paper by adding a label-validation study and by quantifying the removed-label fraction, so I do not recommend rejection. I would also encourage the editor to weigh the open-set claim against the fact that only closed-set experiments are reported, since this affects the dataset's advertised novelty."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth knowing: CHIRLA fills a real slot in the Re-ID benchmark zoo—long-term (seven months), video-based, synchronized multi-camera, indoor, with clothing change and higher-resolution crops than DeepChange. The Table 1 comparison is convincing, and the dataset itself is a serious resource: roughly 964K boxes, 22 identities, public code, and a GUI for labeling. Credit where due: the authors ship the data and the benchmark code, evaluate a broad set of trackers and Re-ID models, and honestly report that the long-term scenario is hard (best CMC@1 18.81%). The low numbers support the claim that the benchmark is challenging. The soft spots are real but fixable. The stress-test note is on target about label reliability: the paper describes YOLOv8 + Deep SORT + manual GUI review but gives no inter-annotator agreement, no label error rate, and no independent check that IDs are consistent across months. Deep SORT does make association errors, and manual correction of a million boxes cannot be verified from the description alone. That said, I would not call the whole thing collapsed—the dataset still has value as a resource, and the manual review process is at least described. The more clear-cut problem is the 'fully annotated' claim in the abstract and contribution list, which directly conflicts with the Procedure statement that individuals are labeled only if they can be sufficiently identified and otherwise removed. That overclaim needs to be corrected. Also, the train/test splits were manually selected, and the paper reports no error bars across subsets, so the measured difficulty is not yet a stable reference. The open-set distractor feature is advertised as a contribution but not evaluated—the paper explicitly leaves it for future work, which is honest but means the claim should be softened. The central combination—long-term, video, synchronized cameras, clothing change—does hold up as new. The paper deserves a serious referee, but the review should require: (1) a quantitative label-consistency check (e.g., a sample re-verified by independent annotators, or at minimum per-ID agreement statistics), (2) a correction of the 'fully annotated' phrasing, (3) error bars or at least subset-level variance for the benchmark numbers, and (4) either an open-set evaluation or a revised contribution list. This is a useful dataset paper, not a definitive benchmark paper yet.","headline":"CHIRLA is a genuinely new long-term video Re-ID benchmark combination, but the label-quality evidence is missing and the fully annotated claim is contradicted by the paper's own procedure.","tokens_in":725,"tokens_out":1716,"would_cite":true,"duration_ms":28892,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces CHIRLA, a seven-month multi-camera indoor video dataset for long-term person re-identification, and reports that the best model tested achieves 18.81% rank-1 accuracy in its long-term scenario.","keywords":["person re-identification","long-term Re-ID","clothing change","video dataset","multi-camera benchmark","indoor surveillance","occlusion","tracking"],"falsifier":"Sample bounding boxes assigned to the same ID from different sequences and have several independent human annotators judge whether they depict the same person; if a meaningful share of matches disagree with the dataset labels, the long-term benchmark scores partly reflect annotation noise rather than appearance change.","tokens_in":19537,"feed_emoji":"🎥","tokens_out":9882,"duration_ms":76350,"temperature":0.7,"pith_summary":"The paper introduces CHIRLA, a dataset for long-term person re-identification recorded over seven months in a four-room indoor office with seven synchronized cameras. It contains 22 people, more than 5.5 hours of video, and about 963,000 annotated bounding boxes, with identity labels kept consistent across cameras and recording sessions and substantial clothing variation between sessions. The authors also define benchmark protocols for two tracking scenarios and four Re-ID scenarios — reappearance, long-term, multi-camera, and multi-camera long-term — and evaluate many CNN and transformer models on them. The best model reaches 18.81% CMC@1 and 23.24% mAP on the long-term scenario, and tracking accuracy drops sharply under occlusion. If the dataset and its labels are sound, CHIRLA provides a benchmark where appearance change, occlusion, and camera transitions can be tested together in continuous video.","feed_headline":"Seven-month indoor person re-ID benchmark tops out at 18.81%","feed_subtitle":"CHIRLA records seven months of indoor video; best models re-identify only about one in five people after long gaps.","key_machinery":"The central object is the CHIRLA dataset itself: seven synchronized 1080×720 cameras recording four connected indoor spaces over seven months, yielding 10 video sequences and 963,554 annotated bounding boxes for 22 identities. Identity labels were produced by a semi-automatic pipeline combining YOLOv8 person detections, Deep SORT tracklets, and manual review in a custom GUI, with IDs kept consistent across all cameras and sequences. The benchmark protocols are the other load-bearing component: tracking tasks built from occlusion intervals of one to five seconds, and Re-ID tasks built from reappearances, long temporal gaps, cross-camera views, and their combination, scored with CMC, mAP, MOTA, and IDF1.","core_discovery":"CHIRLA is claimed to be the first fully annotated, long-term, video-based person Re-ID dataset captured in an indoor environment, combining temporal continuity with deliberate clothing change. It offers more bounding boxes per identity than previous clothing-change datasets, about 44,000 per ID, and higher-resolution person crops than existing long-term benchmarks, while the synchronized seven-camera setup supports cross-camera evaluation. Benchmark experiments show that current Re-ID models, including CNN and transformer backbones, reach only 18.81% rank-1 accuracy in the long-term scenario and 24.31% in the multi-camera long-term scenario, and that occlusion-focused tracking results are far worse than full-test-set results, with best MOTA dropping from 68.84 to 12.42 under brief occlusions. The paper takes these numbers as evidence that long-term indoor Re-ID with clothing changes remains unsolved and that CHIRLA offers a realistic testbed for it.","pith_inferences":["Editorial inference: if the long-term scores are representative, practical indoor identity systems should not rely on clothing or body appearance alone; face, gait, or other stable cues would be needed for months-long tracking.","Editorial inference: the dataset's roughly 44,000 boxes per identity make it possible to test tracklet-level or temporal-aggregation Re-ID models that exploit within-session appearance consistency, an avenue the paper does not explore.","Editorial inference: because the paper reports no inter-annotator agreement or label-error rate, a random-subset re-labeling study would calibrate how much of the benchmark difficulty is real appearance change versus annotation noise."],"forward_implications":["If CHIRLA is a valid long-term benchmark, state-of-the-art Re-ID models are not ready for months-long indoor identity matching: the best long-term CMC@1 is 18.81% and the best mAP 23.24%.","Appearance-based tracking helps overall, but not after occlusions: the best full-test MOTA is 68.84, while the best MOTA on brief-occlusion and multi-person-occlusion subsets falls to 12.42 and 21.83, so maintaining identity through occlusion remains an open problem.","The open-set protocol with unknown distractor IDs provides a concrete way to evaluate decision thresholds for rejecting strangers, something closed-set CMC and mAP cannot measure.","Because CHIRLA provides continuous synchronized video and higher-resolution crops, methods that combine temporal continuity, face appearance, and body appearance can be assessed together, which cropped-image clothing-change datasets do not permit."],"supporting_citations":[{"why":"The main existing long-term clothing-change dataset, used to argue that its cropped-image-only release limits temporal and face-aware evaluation.","marker":"[16]"},{"why":"A multi-day, multi-camera clothing-change dataset that is image-based, providing the comparison point for CHIRLA's continuous video.","marker":"[15]"},{"why":"A two-outfit clothing-change dataset with limited viewpoints, used to position CHIRLA's greater appearance variation.","marker":"[14]"},{"why":"The standard video-based Re-ID benchmark, establishing the temporal-coherence property CHIRLA extends to the long-term setting.","marker":"[11]"},{"why":"Supplies the person detections that seed the semi-automatic annotation pipeline.","marker":"[34]"},{"why":"Provides the initial identity tracklets that were manually reviewed and corrected.","marker":"[35]"},{"why":"Defines MOTA and the CLEAR MOT metrics used to evaluate the tracking benchmark.","marker":"[37]"},{"why":"Defines the CMC rank metric used to evaluate the Re-ID benchmark.","marker":"[39]"}],"fun_headline_variants":["Seven-month indoor benchmark: re-ID models hit only 18.8%","CHIRLA: 7 months, 22 people, 1M boxes, 18.8% top accuracy","Re-ID models fail long-term: 18.8% rank-1 on CHIRLA","Long-term indoor re-ID dataset: best models 18.8%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the identity labels are correct: every bounding box assigned to an ID across the seven-month capture really shows that same person, and no labeling error systematically swaps or confuses identities.","fun_headline_variants_meta":{"raw":{"variants":["Seven-month indoor benchmark: re-ID models hit only 18.8%","CHIRLA: 7 months, 22 people, 1M boxes, 18.8% top accuracy","Re-ID models fail long-term: 18.8% rank-1 on CHIRLA","Long-term indoor re-ID dataset: best models 18.8%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000638,"raw_usage":{"total_tokens":2939,"prompt_tokens":947,"completion_tokens":1992,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":563,"completion_tokens_details":{"reasoning_tokens":1897}},"tokens_in":563,"tokens_out":1992,"duration_ms":15029,"temperature":1.0,"reasoning_tokens":1897,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T14:40:33.436113+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Sample bounding boxes assigned to the same ID from different sequences and have several independent human annotators judge whether they depict the same person; if a meaningful share of matches disagree with the dataset labels, the long-term benchmark scores partly reflect annotation noise rather than appearance change.","supporting_citations":[{"cited_title":"InProceedings of the Asian Conference on Computer Vision (ACCV), https://doi.org/10.1007/978-3-030-69535-4_5 (2020)","cited_arxiv_id":null,"evidence_quote":"A multi-day, multi-camera clothing-change dataset that is image-based, providing the comparison point for CHIRLA's continuous video."},{"cited_title":"& Paulus, D","cited_arxiv_id":null,"evidence_quote":"Provides the initial identity tracklets that were manually reviewed and corrected."},{"cited_title":"& Tao, H","cited_arxiv_id":null,"evidence_quote":"Defines the CMC rank metric used to evaluate the Re-ID benchmark."}],"review_version":1}