{"id":"50144b70-eeba-46c1-8767-9cf38ff75672","arxiv_id":"2505.20540","paper_version":2,"verdict":"REJECT","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"A survey of causal reasoning for video person re-identification that reviews DIR-ReID, identity-shuffle GANs, and causal transformers, but contains unverified performance claims.","lead":"This survey argues that video-based person re-identification should move from correlation-based models to causal reasoning that separates identity cues from confounders such as clothing and background. It reviews methods, datasets, and metrics for causal Re-ID, but the specific performance gains and a shopping mall case study are presented without verifiable sourcing.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's central quantitative evidence is unverifiable: the +11.2%, +15.3%, and +7.8% gains attributed to DIR-ReID, IS-GAN, and UCT do not trace to any source containing those methods or results, so the conclusion's case for causal Re-ID lacks a factual foundation.","rationale":"The reader's verdict correctly identifies the core problem: the survey's central claim rests on specific performance numbers that are not supported by the cited literature. My independent check of the reference list confirms an even more concrete failure than 'missing protocol': reference [18], which is repeatedly used to support IS-GAN and its +15.3% clothing-change gain, is actually the STMN paper, a spatial-temporal memory network with no identity-shuffle GAN and no DeepChange evaluation. Reference [5], cited for DIR-ReID's causal intervention and +11.2% gain, is an image-based domain-generalization paper that does not use the causal-intervention framing the survey attributes to it. Reference [15], used for UCT, is an image-based visible-infrared method, not a video transformer, and the claimed 62.7% Rank-1 is not checked against its reported protocol. The Section 5.2 deployment story compounds this: a detailed 67% to 89% improvement in a European shopping mall is presented with citations to generic retail-tracking and edge-device papers that do not contain that measurement. Since the strongest claim in the conclusion is explicitly quantitative, these unverifiable numbers are not ancillary; they are the entire evidentiary basis for the recommended paradigm shift. The appropriate remedy is to reject the survey in its current form and require the authors to replace each unsupported number with a correctly cited, protocol-backed result or to remove the quantitative claims entirely. This is not a disagreement with the plausible general idea that causal reasoning might improve robustness; it is a statement about the internal traceability of the paper's own evidence. I agree with the reader's weakest-assumption assessment and see no additional load-bearing concern beyond it, though the citation mismatches make the problem more severe than a simple missing protocol.","tokens_in":24483,"tokens_out":2513,"duration_ms":27709,"concrete_test":"For each flagship claim, retrieve the actual cited paper and check three things: (1) whether the paper's method name matches the survey's name (DIR-ReID in arXiv:2103.15890; IS-GAN in reference [18], currently Eom et al. STMN; UCT in reference [15], currently Yuan et al. unbiased feature learning); (2) whether the exact numbers +11.2% cross-domain, +15.3% DeepChange, and +7.8% SYSU-MM01 appear with a described evaluation protocol and dataset split; and (3) whether the evaluation uses video tracklets rather than single images. If any of these checks fails, that claim must be removed or replaced with a correctly cited, protocol-backed number before the survey's central conclusion can stand.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim in the conclusion is that causal Re-ID models 'achieve substantial gains' of +11.2% Rank-1 in cross-domain generalization, +15.3% on clothing-change datasets, and +7.8% in cross-modality tasks. These are the only quantitative supports offered for the survey's paradigm-shift thesis, yet none is traceable. In Section 4.2, DIR-ReID is cited to reference [5], but that paper ('Learning Domain Invariant Representations for Generalizable Person Re-Identification') is an image-based domain-generalization method that does not present itself as a causal video Re-ID model and does not report the claimed +11.2% figure as an intervention result. IS-GAN is cited to reference [18], which is actually Eom et al.'s 'Video-based Person Re-identification with Spatial and Temporal Memory Networks' (STMN), not an 'Identity Shuffle GAN' and not a clothing-change model; the DeepChange +15.3% result therefore has no supporting source. UCT is cited to reference [15], 'Unbiased Feature Learning with Causal Intervention for Visible-Infrared Person Re-identification', which is an image-based cross-modality method, not the 'Unbiased Causal Transformer' described, and the 62.7% / +7.8% SYSU-MM01 numbers are not verified in the text. The shopping-mall case study in Section 5.2 is similarly presented as fact with references that are generic retail/tracking papers and contain no such deployment protocol. Because the survey's thesis is explicitly quantitative and the numbers are load-bearing, the absence of verifiable sources for all three flagship gains is a decisive correctness problem, not a stylistic one.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a survey of causal methods for video-based person re-identification (Re-ID). It proposes a taxonomy of causal approaches (generative disentanglement, domain-invariant modeling, causal transformers), reviews datasets and evaluation metrics, and argues that causal models outperform correlation-based methods, with specific quantitative claims such as +11.2% Rank-1 for DIR-ReID, +15.3% on DeepChange for IS-GAN, and +7.8% for UCT. It also presents a real-world shopping mall deployment case study and concludes by advocating a paradigm shift from correlation to causation. The paper is structured as an expository review rather than a technical contribution.","tokens_in":24785,"tokens_out":5140,"duration_ms":49208,"significance":"If the quantitative claims were substantiated, the survey would provide a useful systematic overview of an emerging subfield and a strong argument for causal methods. The paper does offer a broad taxonomy, a compilation of datasets, and a discussion of causal concepts, which are of some value. However, the load-bearing empirical evidence is unverifiable: the cited references do not contain the claimed methods or results. The central thesis—that causal interventions yield substantial practical gains—is therefore unsupported, making the paper's significance contingent on unverifiable assertions.","major_comments":[{"comment":"The central quantitative claims of the survey are not traceable to the cited literature. Section 4.2 attributes to DIR-ReID [5] a Rank-1 of 75.2% on Market-1501→DukeMTMC-ReID and a +11.2% improvement over non-causal baselines, but reference [5] (Zhang et al., \"Learning Domain Invariant Representations for Generalizable Person Re-Identification\") is an image-based domain-generalization method that does not present causal interventions and does not report these figures. Section 4.2 attributes to IS-GAN [18] a +15.3% Rank-1 improvement on DeepChange, but reference [18] is Eom et al.\"s \"Video-based Person Re-identification with Spatial and Temporal Memory Networks\" (STMN), not an \"Identity Shuffle GAN\" and not a clothing-change model. Section 3.1 additionally claims a +15.7% improvement under occlusion for IS-GAN, again citing [18]. Section 4.2 attributes to UCT [15] a 62.7% Rank-1 on SYSU-MM01 and a +7.8% improvement, but reference [15] (Yuan et al.) is an image-based visible-infrared method named \"Unbiased Feature Learning with Causal Intervention for Visible-Infrared Person Re-identification\", not the \"Unbiased Causal Transformer\" described in the text, and the numbers are not reported there. Because the conclusion (Section 9) rests its paradigm-shift argument on these specific gains, the failure of traceability undermines the paper's central claim.","section":"4.2 (also 3.1, 5.1, 9)"},{"comment":"Section 5.2 describes a \"large European shopping mall deployment\" in which replacing a correlation-based system with a causal DIR-ReID model improved cross-camera re-identification accuracy from 67% to 89% and jacket-removal cases from 51% to 83%. The text cites references [5,86,91,92,93,43,48,94], but none of these is a case study of this mall deployment: [5] is the image-based DIR-ReID paper, [86] is a general retail open-world Re-ID paper, [91] is a community-college surveillance case study, [92] is about multi-resolution Re-ID, [93] is about edge computing, and [43] is about causal intervention for clothes-changing Re-ID without this deployment. No measurement protocol, dataset split, or system description is provided. This unsourced empirical narrative is presented as fact and is load-bearing for the survey's practicality claims.","section":"5.2"},{"comment":"The survey's scope is internally inconsistent: it uses the term \"causal video-based person Re-ID\" but applies it to methods that are not video-based. DIR-ReID operates on single images, and UCT is an image-based visible-infrared method. Table 5, the summary of recent video-based Re-ID methods, contains no causal video method other than STMN (which is not causal), while Table 4 lists DIR-ReID, IS-GAN, DCR-ReID, and UCT as causal methods despite not satisfying the video-based definition. The absence of explicit inclusion criteria for \"causal video-based Re-ID\" makes the taxonomy ambiguous and undermines the survey's central organizational claim.","section":"4 (overall taxonomy; Table 5)"},{"comment":"Section 7 states as facts that compression and hardware optimizations introduce accuracy trade-offs of 5-15%, demographic error rate disparities reach 23%, privacy-preserving methods drop accuracy by 10-15%, and real-world deployments suffer 30-40% accuracy drops, with none of these figures cited. Section 8 similarly presents unsourced projections (e.g., 60% computation reduction, 20x throughput, 70-80% labelled-data reduction, 8-12% out-of-domain gains, 15-20% multimodal reliability improvements). In a survey, these quantitative statements require references; without them they appear invented and compound the traceability problem already present in Sections 4 and 5.","section":"7 (also 8)"}],"minor_comments":[{"comment":"The header \"Journal Not Specified\" and the line \"Submitted toJournal Not Specified for possible open access publication\" indicate that the manuscript has not been processed by an actual journal; the authors should provide the publication venue or remove the placeholder.","section":"Header / metadata"},{"comment":"There are inconsistencies in capitalization and spelling, such as \"video-based person Re-ID\" versus \"video-based person Re-ID\" and \"Labratory\" in the affiliation; a careful proofreading pass is needed.","section":"2.2 and throughout"},{"comment":"The Figure 4 caption states that the violin plot shows 32% versus 8% \"not as experimental values\", but the surrounding text cites performance improvements (e.g., +11.2%) without clarifying which numbers are illustrative and which are empirical; the distinction should be made explicit.","section":"3.1, Figure 4 caption"},{"comment":"Reference [24] (Wang et al., \"Causal disentanglement for semantics-aware intent learning in recommendation\") is about recommender systems, not person Re-ID; citing it in the context of causal disentanglement for Re-ID is inappropriate and weakens the survey's credibility.","section":"References"},{"comment":"The formal definition of an SCM as a tuple G = (V, E) is incomplete; a structural causal model normally includes exogenous variables, structural equations, and a distribution over exogenous noise, which are not mentioned.","section":"3.2, SCM definition"}],"recommendation":"reject","confidential_remarks":"The manuscript's reference list does not support the central numerical claims; the mismatches are extensive and affect the paper's core argument. Even a major revision would need to replace the unsupported numbers with verified results or restructure the survey to avoid claiming quantitative evidence. Given the survey's explicitly quantitative thesis, a reject is warranted. The authors may benefit from consulting the actual papers they cite and ensuring that each attributed result appears in the cited source."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a survey with a plausible central idea — that causal reasoning could address the brittleness of video-based person Re-ID — but the quantitative evidence it leans on does not hold up. The three flagship numbers (+11.2% cross-domain, +15.3% clothing-change, +7.8% cross-modality) and the Section 5.2 shopping mall case study are either mis-attributed or unsourced. That matters because the conclusion uses those numbers as the core case for a \"paradigm shift.\"\n\nWhat the paper does well: it assembles a broad, current bibliography on causal methods in Re-ID, organizes the space into generative disentanglement, domain-invariant modeling, and causal transformers, and includes a useful dataset table. The worked example in Section 3.4 (the red-jacket/blue-jacket illustration) is genuinely clear and would help newcomers. The discussion of fairness, privacy, and interpretability is reasonable, if shallow.\n\nThe soft spots are not minor. The claims about DIR-ReID, IS-GAN, and UCT don't trace to the cited sources: reference [5] is an image-based domain generalization paper that doesn't report the +11.2% figure; reference [18] is STMN, a memory network, not an \"Identity Shuffle GAN\"; and reference [15] is an image-based cross-modality method, not the \"Unbiased Causal Transformer.\" The 67% to 89% shopping mall deployment in Section 5.2 is presented as fact with references to generic retail-tracking papers that contain no such protocol. The \"causal-specific robustness measures\" promised in the abstract are not new definitions; they're existing metrics from other papers. In other words, the survey's empirical foundation is not merely weak — it is unverifiable.\n\nI'd still send this to peer review, because the topic is real and the flaws are fixable in principle. But the authors need to verify or retract every quantitative claim, replace the bogus citations, and either source the case study or drop it. If that can't be done, the correct outcome is rejection.\n\nWho gets value from this? Readers looking for a conceptual map of causal ideas in Re-ID, and perhaps a cautionary example of how citation errors can undermine a survey. I wouldn't cite it in my own work until the numbers are fixed.\n\nRecommendation: serious referee, with the expectation of major revision or rejection.","headline":"A useful survey of causal ideas for video Re-ID that is undermined by untraceable performance numbers and an unsourced case study.","tokens_in":25368,"tokens_out":3627,"would_cite":false,"duration_ms":34474,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This survey argues that video-based person re-identification fails in the wild because it learns correlations, and that causal modeling—treating identity as a cause of appearance and intervening on confounders—can replace those fragile…","keywords":["video-based person re-identification","causal inference","structural causal models","counterfactual reasoning","domain generalization","clothing-change robustness","disentangled representation learning","surveillance applications"],"falsifier":"Re-run the cited methods under their original protocols: train DIR-ReID on Market-1501 and evaluate on DukeMTMC-ReID to check whether the cross-domain Rank-1 gain over the non-causal baseline is about 11.2%; train the identity-shuffling GAN on DeepChange to check whether the clothing-change Rank-1 gain is about 15.3%; and run UCT on SYSU-MM01 to check for about 7.8% over DIR-ReID and IS-GAN. Also inspect reference [86] and the IS-GAN references to see whether the shopping mall deployment and the identity-shuffling results actually appear there; if the numbers do not reproduce or the sources do not contain them, the survey's central quantitative claim is unsupported.","tokens_in":24252,"feed_emoji":"🎥","tokens_out":9390,"duration_ms":93305,"temperature":0.7,"pith_summary":"Video-based person re-identification models that look strong on benchmarks repeatedly fail in real deployments because they learn correlations—clothing, background, lighting—that break under new cameras, viewpoints, and outfits. This survey argues that the fix is causal: explicitly treat identity as a generative cause of appearance, model confounders with structural causal models, and train with interventions and counterfactuals that force identity predictions to stay stable when non-identity attributes change. If the reported numbers are right, the payoff is large: roughly 11.2% higher Rank-1 accuracy in cross-domain transfer, 15.3% on clothing-change datasets, and 7.8% in visible-to-infrared matching, where correlation-based systems degrade sharply. The survey organizes these efforts into a taxonomy of generative disentanglement, domain-invariant modeling, and causal transformers, reviews metrics and datasets, and concludes that the field should shift from correlation-based to causal learning. Why it matters: surveillance, retail analytics, and forensics depend on matching people across cameras, and current systems fragment identities exactly when conditions vary.","feed_headline":"Causal reasoning lifts person re-ID in the wild, survey finds","feed_subtitle":"Intervening on clothing, lighting, and background adds 11–15% Rank-1 accuracy where correlation-based models collapse.","key_machinery":"The central machinery is the Structural Causal Model (SCM), a directed graph whose nodes are identity, appearance, and confounders such as clothing, background, camera, and occlusion, and whose edges encode the generative statement that identity causes appearance. The load-bearing operation is the intervention, typically written $P(\\mathrm{ID} \\mid do(\\mathrm{Clothing}=c)) = \\sum_z P(\\mathrm{ID} \\mid \\mathrm{Clothing}=c, Z=z)P(Z=z)$, which removes the backdoor paths that let clothing or background masquerade as identity. Training then uses counterfactual generation $X' = f(I, D')$ (same identity, altered domain) plus a consistency loss $\\mathcal{L}_{\\text{causal}} = d(f_{\\mathrm{ID}}(A), f_{\\mathrm{ID}}(A'))$, and adversarial identity shuffling in IS-GAN, to enforce invariance. DIR-ReID and UCT apply the same machinery to domain features and visible-infrared modality shifts.","core_discovery":"The paper's central claim is that the brittleness of video Re-ID is structural: models trained to minimize $P(Y \\mid X)$ on curated tracklets will always exploit spurious cues, so benchmark success does not transfer to the wild. Causal Re-ID addresses this by building a structural causal model in which Identity causes Appearance and confounders such as clothing, background, camera, and occlusion also act on appearance, then applying interventions to block those confounder paths. The paper surveys three families of implementation: DIR-ReID's domain-feature intervention for cross-domain generalization, IS-GAN's identity-shuffling generative disentanglement for appearance change, and the UCT causal transformer for cross-modality matching. It reports concrete gains for these models and concludes that real-world robustness, fairness, interpretability, and privacy all improve when identity is learned as a cause rather than a correlation.","pith_inferences":["A reader wanting to build on the survey will need to locate the original protocols for the headline numbers: the paper does not give the split or measurement setup behind the 11.2%, 15.3%, and 7.8% gains, and the reference cited for IS-GAN is the STMN paper, not the identity-shuffling GAN.","If the causal framing is doing real work, then a model that literally swaps clothing or background during training should recover most of the reported gains even without a formal SCM; if it does, the causal vocabulary may be a useful scaffold for a data-augmentation effect rather than a separate mechanism.","A natural next experiment the paper does not run is to compare counterfactual positive pairs against standard augmentations in a self-supervised pretraining loop on DeepChange or a similar clothing-change benchmark, and measure whether the intervention-style pairs give the out-of-domain gains the survey anticipates.","The retail deployment claim is checkable: if a public or independently licensed retail dataset reproduces the 67%-to-89% jump when switching from a correlation baseline to a body-shape-and-gait causal model, that would turn the survey's strongest anecdote into a transferable result."],"forward_implications":["If the causal claim is right, a Re-ID model trained with clothing and background interventions should keep its identity embedding stable when a person changes outfits, and Rank-1 accuracy on clothing-change benchmarks such as DeepChange should rise by roughly 15%.","Cross-domain and cross-modality evaluations (e.g., Market-1501 to DukeMTMC, visible to infrared on SYSU-MM01) should show gains near the reported 11.2% and 7.8% over correlation-based baselines.","Evaluation practice should widen beyond CMC and mAP to include counterfactual consistency, causal saliency ranking, and intervention-based score shift, because those metrics directly test whether a model is using identity causes rather than shortcuts.","Deployments that replace correlation-based trackers with causal disentanglement models should see fewer identity switches across camera transitions, as in the reported retail deployment where cross-camera accuracy rose from 67% to 89%.","A shift to causal Re-ID would also change the field's stated goals: fairness by intervening on protected attributes, privacy by learning minimal identity representations, and interpretability through counterfactual explanations become part of the standard design."],"supporting_citations":[{"why":"Supplies DIR-ReID, the SCM-based domain-invariant method whose reported 11.2% cross-domain Rank-1 gain anchors the survey's main quantitative claim.","marker":"[5]"},{"why":"Supplies the structural causal model and do-calculus framework used to define interventions and backdoor adjustment.","marker":"[11]"},{"why":"Supplies UCT, the causal transformer whose reported 7.8% cross-modality Rank-1 gain supports the survey's cross-modal claims.","marker":"[15]"},{"why":"Cited in the survey for IS-GAN's identity-shuffling disentanglement and the reported 15.3% clothing-change gain; the listed article is the STMN memory-network paper.","marker":"[18]"},{"why":"Provides the DeepChange clothing-change benchmark on which the 15.3% improvement is reported.","marker":"[42]"},{"why":"Supports the claim that clothing acts as a confounder and that causal cloth-debiasing improves cloth-changing person Re-ID.","marker":"[7]"},{"why":"Cited for the European shopping mall deployment and the 67% to 89% cross-camera accuracy figures.","marker":"[86]"}],"fun_headline_variants":["Survey: Causal models beat correlations for video person re-ID","Causal re-ID: Blocking confounders boosts real-world accuracy","Why person re-ID fails in the wild: Causal survey offers fix","Video re-ID needs causal logic, not surface cues, says survey","Causal reasoning key to person re-ID beyond benchmarks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's quantitative case rests on the accuracy and traceability of specific gains attributed to DIR-ReID, IS-GAN, and UCT (11.2%, 15.3%, and 7.8% Rank-1) and on the reality of the Section 5.2 European shopping mall deployment with its 67% to 89% accuracy figures; if those numbers are wrong or cannot be traced to reproducible experiments, the central argument loses most of its force.","fun_headline_variants_meta":{"raw":{"variants":["Survey: Causal models beat correlations for video person re-ID","Causal re-ID: Blocking confounders boosts real-world accuracy","Why person re-ID fails in the wild: Causal survey offers fix","Video re-ID needs causal logic, not surface cues, says survey","Causal reasoning key to person re-ID beyond benchmarks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00017,"raw_usage":{"total_tokens":1256,"prompt_tokens":924,"completion_tokens":332,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":243}},"tokens_in":540,"tokens_out":332,"duration_ms":3725,"temperature":1.0,"reasoning_tokens":243,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T13:51:52.951064+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the cited methods under their original protocols: train DIR-ReID on Market-1501 and evaluate on DukeMTMC-ReID to check whether the cross-domain Rank-1 gain over the non-causal baseline is about 11.2%; train the identity-shuffling GAN on DeepChange to check whether the clothing-change Rank-1 gain is about 15.3%; and run UCT on SYSU-MM01 to check for about 7.8% over DIR-ReID and IS-GAN. Also inspect reference [86] and the IS-GAN references to see whether the shopping mall deployment and the identity-shuffling results actually appear there; if the numbers do not reproduce or the sources do not contain them, the survey's central quantitative claim is unsupported.","supporting_citations":[{"cited_title":"Learning Domain Invariant Representations for Generalizable Person Re-Identification","cited_arxiv_id":"2103.15890","evidence_quote":"Supplies DIR-ReID, the SCM-based domain-invariant method whose reported 11.2% cross-domain Rank-1 gain anchors the survey's main quantitative claim."},{"cited_title":"Unbiased Feature Learning with Causal Intervention for Visible- Infrared Person Re-identification","cited_arxiv_id":null,"evidence_quote":"Supplies UCT, the causal transformer whose reported 7.8% cross-modality Rank-1 gain supports the survey's cross-modal claims."},{"cited_title":"Video-based Person Re-identification with Spatial and Temporal Memory Networks","cited_arxiv_id":null,"evidence_quote":"Cited in the survey for IS-GAN's identity-shuffling disentanglement and the reported 15.3% clothing-change gain; the listed article is the STMN memory-network paper."},{"cited_title":"DeepChange: A Large Long-Term Person Re-Identification Benchmark with Clothes Change","cited_arxiv_id":"2105.14685","evidence_quote":"Provides the DeepChange clothing-change benchmark on which the 15.3% improvement is reported."},{"cited_title":"Person detection and re-identification in open-world settings of retail stores and public spaces","cited_arxiv_id":"2505.00772","evidence_quote":"Cited for the European shopping mall deployment and the 67% to 89% cross-camera accuracy figures."}],"review_version":1}