{"id":"38701be0-d366-4470-911c-c33fab72cce9","arxiv_id":"2607.14821","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":9,"one_line_summary":"A unified taxonomy and survey of single- to multi-modal person ReID, plus a Transformer-based VI-ReID baseline that is solid but not state-of-the-art.","lead":"This preprint surveys person re-identification across modalities—from RGB to infrared, text, sketch, and even radio signals—and proposes a Transformer baseline for visible-infrared matching. It offers a single map of a fragmented field and a starting point for building cross-modal identity systems.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Baseline 'superior overall performance' contradicted by methods in the paper's own Tables 3 and 4; survey coverage claim itself is not undermined.","rationale":"The reader's weakest assumption correctly identifies the baseline evidence as the load-bearing weak point. My reading of Tables 3, 4, and 12 strengthens that concern: the claim of 'superior overall performance on both datasets' is not merely under-supported by missing error bars and code; it is contradicted by the paper's own comparison tables. DEN alone outperforms the proposed supervised method on every reported metric, and SDCL/MCL outperform it on every reported SYSU-MM01 metric in the unsupervised setting. This is a concrete, internally checkable inconsistency rather than a matter of taste or external consensus. I therefore agree with the reader that the conditional verdict is appropriate. I do not see a comparable problem with the survey's central coverage claim: the taxonomy and organization across VI, TI, sketch, NLOS, tri-spectral, and multi-modal settings appear to deliver the claimed breadth, and independent support comes from the paper's extensive dataset/method tables and references. The baseline issue is significant enough to require revision of the experimental claims, but it does not by itself overturn the survey contribution. Hence the reader's CONDITIONAL verdict should remain unchanged.","tokens_in":44672,"tokens_out":6109,"duration_ms":50463,"concrete_test":"Rebuild Table 12 by including every supervised and unsupervised method whose full metrics appear in Tables 3 and 4 (at minimum DEN, PartMix, PMCM, MAUM, SDCL, and MCL), rerun the proposed method under the identical protocol with at least three random seeds, and report mean ± std. If any included method exceeds the proposed method on either dataset's Rank-1 or mAP, the 'superior overall performance' claim should be revised or removed. Additionally, perform the cross-dataset check recommended in Section 2.1.1 — e.g., train on SYSU-MM01 and test on RegDB/LLCM — to determine whether the baseline generalizes rather than overfitting benchmark-specific biases.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The survey's coverage contribution is plausible and independently useful, but the paper's second stated contribution — the Transformer VI-ReID baseline — makes a claim that is internally inconsistent with the paper's own tables. Section 5 states that the proposed method 'demonstrates superior overall performance on both the RegDB and SYSU-MM01 datasets' under supervised and unsupervised settings. Table 12, however, omits methods from Tables 3 and 4 that outperform it. For supervised VI-ReID, DEN [54] in Table 3 achieves SYSU-MM01 All R-1/mAP of 76.36/71.38 and Indoor 83.56/84.65, and RegDB V->I 95.34/90.21 and I->V 94.98/90.24; the proposed method obtains 69.93/68.91, 76.07/81.50, 93.48/88.72, and 92.61/87.72 respectively — lower than DEN on all six metrics. PartMix [55] and PMCM [36] also beat or match it on SYSU-MM01. For unsupervised VI-ReID, SDCL [99] and MCL [122] in Table 4 beat the proposed method on all SYSU-MM01 metrics, e.g., SDCL All R-1 64.49 vs. 60.87 and Indoor R-1 71.37 vs. 66.13. Thus 'superior overall' is unsupported by the paper's own evidence; at best the baseline is competitive on RegDB and below several published methods on SYSU. The absence of code, error bars, and cross-dataset evaluation compounds this, especially since Section 2.1.1 itself recommends cross-dataset and camera-disjoint testing over within-dataset Rank-1/mAP alone. This concern does not invalidate the survey's coverage or taxonomy, but it does undermine the baseline claim as stated.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper surveys person re-identification from single-modal to cross-modal and multi-modal settings. It organizes VI-ReID, TI-ReID, Sketch-ReID, NLOS-ReID, tri-spectral ReID, and multi-modal ReID under a taxonomy based on retrieval protocol and learning objective, and reviews datasets and representative methods. As a second contribution, it proposes a Transformer-based VI-ReID baseline with supervised and unsupervised variants, evaluated on SYSU-MM01 and RegDB. The paper claims to be the first survey covering this combination of scenarios.","tokens_in":45184,"tokens_out":8185,"duration_ms":57965,"significance":"If the survey's coverage claim holds, the paper provides a useful structured reference: it brings together six task families that are usually treated separately, includes extensive method and dataset tables, and is candid about benchmark-specific biases and the need for stronger evaluation protocols. The proposed taxonomy is a reasonable organizing principle. The baseline contribution, however, is not established as claimed: the experimental evidence in the paper's own Tables 3 and 4 contradicts the 'superior overall performance' statement, and no code or error bars are provided. The survey content remains defensible, but the experimental claim requires substantial revision.","major_comments":[{"comment":"The claim in §5 that the proposed baseline 'demonstrates superior overall performance on both the RegDB and SYSU-MM01 datasets' under supervised and unsupervised settings is contradicted by the paper's own tables. In the supervised setting, DEN [54] (Table 3) outperforms the baseline on every reported metric: SYSU-MM01 All 76.36/71.38 vs 69.93/68.91, Indoor 83.56/84.65 vs 76.07/81.50, RegDB V→I 95.34/90.21 vs 93.48/88.72, and I→V 94.98/90.24 vs 92.61/87.72. PartMix [55] also beats the baseline on SYSU-MM01 All and Indoor. In the unsupervised setting, SDCL [99] (Table 4) exceeds the baseline on SYSU-MM01 All (64.49/63.24 vs 60.87/59.56) and Indoor (71.37/76.90 vs 66.13/73.21). None of these methods appears in Table 12. The comparison is therefore selective, and the 'superior overall' statement should be replaced with a qualified claim or the table should include the full set of methods fr","section":"§5 / Table 12 vs Tables 3 and 4"},{"comment":"The experimental support for the baseline is limited to single-run within-dataset Rank-1 and mAP on SYSU-MM01 and RegDB. No cross-dataset evaluation, camera-disjoint testing, or robustness to degraded modalities is reported, despite §2.1.1 explicitly recommending that 'future studies should report cross-dataset evaluation, camera- or environment-disjoint testing' rather than relying solely on within-dataset Rank-1 and mAP. The absence of error bars or multiple-seed results further weakens the 'superior' and 'significantly surpass' wording. These limitations are structurally separate from the survey's coverage contribution and should be fixed by strengthening the evaluation or by re-scoping the claim.","section":"§5 / §2.1.1 Evaluation protocol"},{"comment":"The ablations in Table 13 report single-run numbers without variance or statistical tests. Since the reported gaps among some variants are small (e.g., supervised SYSU R-1 69.93 vs 67.21), it is not possible to determine whether the differences are meaningful. In addition, the ablation does not include a standard Transformer ReID baseline (e.g., DC-Former [238], which the design explicitly draws on), so the incremental contribution of the proposed modules is not isolated from the gains of the backbone. Please provide repeated runs, variance, or a stronger baseline comparison.","section":"§5.2 Ablation study"}],"minor_comments":[{"comment":"The paragraph 'Identity-aware Foundation Models' appears twice verbatim in the Future Research section; remove the duplicate.","section":"§6"},{"comment":"The Wi-PER81 dataset is introduced twice through the same reference [214]; the later 'More recently, Cascio et al. [214] further advanced' sentence should be merged with the earlier description.","section":"§2.4"},{"comment":"For PD [160], the RSTPReid mAP is listed as '????'; please supply the value or mark it as not reported consistently with other entries.","section":"Table 6"},{"comment":"The header 'Identitiesr' contains a typo; should be 'Identities'. Also the table caption says 'Low-light' but the column is not defined in the main text.","section":"Table 2"},{"comment":"The text refers to 'Fig. 1, D' when discussing noisy correspondence in TI-ReID; the intended figure appears to be Fig. 5. Please correct the cross-reference.","section":"§2.2"},{"comment":"The NLOS ReID section would benefit from a performance summary table analogous to Tables 3, 4, 6, 8, 9, and 11; currently the relative strengths of ReID3D, mmWave, and RF methods are described only qualitatively.","section":"§2.4"}],"recommendation":"major_revision","confidential_remarks":"The survey portion is likely to be of interest to the journal's readership. My main concern is that the baseline claim is presented more strongly than the evidence supports, and the selective Table 12 may create an impression of cherry-picking even if unintended. I would encourage the editor to request a full comparison with the methods listed in Tables 3 and 4, or a clearly scoped claim, and ideally code release for reproducibility."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on 2607.14821. The survey part is genuinely useful. As far as I can tell, it is the first to put VI, TI, sketch, and NLOS ReID in one frame, and the organizing principle—retrieval protocol plus learning objective rather than just modality count—is sensible. The NLOS coverage is a nice service to the community, and the tri-spectral and multi-modal fusion discussion is more than a placeholder. If someone is new to cross-modal ReID, this paper will save them a lot of literature hunting. I would use it as a reference for the landscape.\n\nThe baseline, though, is where the paper trips. The text says the proposed Transformer \"demonstrates superior overall performance\" on SYSU-MM01 and RegDB. That is contradicted by the paper's own Tables 3 and 4. On SYSU-MM01 supervised, DEN gets 76.36/71.38 All R-1/mAP and 83.56/84.65 Indoor; the proposed baseline gets 69.93/68.91 and 76.07/81.50. PartMix and PMCM also beat it. On unsupervised, SDCL and MCL beat it on SYSU-MM01 (e.g., SDCL All R-1 64.49 vs. 60.87). Table 12 simply omits those methods. So the \"superior\" claim is not just unsubstantiated—it is contradicted by the paper's own numbers. At best the baseline is competitive on RegDB and below state of the art on SYSU. And for a paper that rightly tells the field to move toward cross-dataset and camera-disjoint evaluation, offering only single-run within-dataset results, in a table that omits the stronger comparators, is an inconsistency. No code, no error bars.\n\nAlso some editorial issues: the future directions section repeats itself verbatim, and there are typos like \"ldentification\" in the title. Minor, but worth cleaning.\n\nI would not call this fatal for the survey. The taxonomy and coverage are the contribution, and they hold up. The baseline can be reframed as a modest reference model—simple and reproducible in principle—without claiming superiority. The authors should fix the claim, add the missing comparisons, and ideally release code and report variance.\n\nVerdict: send it to peer review as a survey, but expect the experimental section to be revised. I'd be fine citing the survey.","headline":"Useful survey of cross-modal ReID that overstates its own baseline; the survey half deserves a serious referee, the baseline claim does not.","tokens_in":45634,"tokens_out":2412,"would_cite":true,"duration_ms":20412,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that person re-identification has become a single spectrum of cross-modal matching tasks, from visible–infrared to text, sketch, and beyond-visibility sensing, and that a simple dual-path Transformer can serve as a competi","keywords":["person re-identification","cross-modal ReID","visible-infrared ReID","text-image ReID","sketch ReID","non-line-of-sight ReID","modality-invariant representation learning","Transformer baseline"],"falsifier":"Run the proposed baseline, with the paper's own settings, on the low-light LLCM benchmark without adaptation after training on RegDB or SYSU-MM01; if its Rank-1 and mAP fall to the level of a standard CNN baseline under the same protocol, the claim that simple Transformer-based modality confusion yields robust modality-invariant features is not supported.","tokens_in":44595,"feed_emoji":"👤","tokens_out":8778,"duration_ms":72718,"temperature":0.7,"pith_summary":"The paper tries to organize the many branches of person re-identification—matching the same person across non-overlapping cameras—into one map. It argues that the field has moved from single-modal RGB matching to cross-modal settings (visible–infrared, text–image, sketch–photo, and signal-based non-line-of-sight cases), and that newer tri-spectral and multi-modal systems belong on the same conceptual spectrum. Its organizing principle is to classify tasks by retrieval protocol and learning objective rather than by the number of modalities involved. The paper also proposes a Transformer-based visible–infrared baseline and claims it outperforms existing methods on the RegDB and SYSU-MM01 benchmarks in both supervised and unsupervised settings. If the map is right, techniques for reducing modality gaps should transfer across all these subfields instead of being developed in isolation.","feed_headline":"One taxonomy now spans every cross-modal ReID flavor","feed_subtitle":"Visible, textual, sketched, and beyond-sight matching share one framework, plus a new Transformer baseline.","key_machinery":"The survey's organizing device is a two-axis taxonomy: each person re-identification task is classified by its retrieval protocol (which modality queries which) and by its primary learning objective (heterogeneous alignment versus spectral-aware representation versus multi-source fusion). This places VI, text-image, sketch, NLOS, tri-spectral, and multi-modal work under one framework. The experimental machinery is a dual-path Transformer baseline for VI-ReID: separate patch embeddings for visible and infrared images, a shared Transformer encoder whose class token aggregates identity information, feature-level modality confusion to erase modality-specific cues, a shared memory bank for cluste","core_discovery":"The paper's central claim is that existing person re-identification research can be described for the first time by a single taxonomy: cross-modal tasks (VI-ReID, TI-ReID, Sketch-ReID, NLOS-ReID) share the problem of heterogeneous modality alignment; tri-spectral ReID centers on spectral-aware representation; and multi-modal ReID centers on multi-source fusion. The authors further claim that a simple Transformer-based framework—separate patch embeddings for visible and infrared inputs, a shared Transformer encoder with a class token as an identity aggregator, feature-level modality confusion, a shared memory bank, and cluster-contrastive learning—captures these principles and achieves strong","pith_inferences":["Editorial inference: if the taxonomy is adopted, benchmark design could move toward a single multi-modal gallery (RGB, IR, text, sketch, and NLOS-style queries) so that alignment techniques are compared under one protocol instead of separate per-modality-pair datasets.","Editorial inference: a testable extension of the baseline is to train on RegDB and evaluate on the low-light LLCM set without adaptation; if the modality-confusion design is genuinely modality-invariant, the drop should be small, and if not, the paper's own recommended cross-dataset protocol would expose it.","Editorial inference: the paper's future direction on causal representation learning implies that current disentanglement methods, which separate factors without causal structure, may not transfer to unseen sensors; a concrete test would swap sensor type at test time and measure the drop."],"forward_implications":["If the taxonomy is right, a method's worth in cross-modal ReID should be judged by how well it solves heterogeneous alignment rather than by which pair of modalities it handles, encouraging technique transfer across VI, text-image, sketch, and NLOS settings.","The proposed baseline demonstrates that a Transformer with a shared encoder, class token, and feature-level modality confusion can outperform established CNN-based VI-ReID methods on RegDB and SYSU-MM01 in both supervised and unsupervised settings.","IR-guided RGB pseudo-label refinement improves unsupervised performance, supporting the principle that the more stable modality can be used to supervise pseudo-label generation for the noisier one.","On the five-modality ORBench-style protocol, adding infrared and color-pencil queries to a text query yields large mAP gains, while adding sketch to an already rich combination gives only marginal gains, so sensor selection should account for diminishing returns.","The paper's own recommendation that future work report cross-dataset, camera-disjoint, and missing-modality evaluation implies that current within-dataset Rank-1/mAP numbers should be read cautiously."],"fun_headline_variants":["One taxonomy now spans all cross-modal ReID","Single taxonomy ties together every ReID modality","From visible to NLOS: one ReID framework","Cross-modal ReID finally gets a unified taxonomy","Blur boundaries: one survey for all ReID modes"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise for the experimental contribution is that single-run, within-dataset Rank-1 and mAP on RegDB and SYSU-MM01 are sufficient evidence that the proposed Transformer baseline is superior—a premise the paper itself questions in Section 2.1.1, where it calls for cross-dataset and camera-disjoint evaluation rather than reliance on within-dataset numbers alone.","fun_headline_variants_meta":{"raw":{"variants":["One taxonomy now spans all cross-modal ReID","Single taxonomy ties together every ReID modality","From visible to NLOS: one ReID framework","Cross-modal ReID finally gets a unified taxonomy","Blur boundaries: one survey for all ReID modes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000772,"raw_usage":{"total_tokens":3235,"prompt_tokens":705,"completion_tokens":2530,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":449,"completion_tokens_details":{"reasoning_tokens":2457}},"tokens_in":449,"tokens_out":2530,"duration_ms":16832,"temperature":1.0,"reasoning_tokens":2457,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T00:56:39.834465+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the proposed baseline, with the paper's own settings, on the low-light LLCM benchmark without adaptation after training on RegDB or SYSU-MM01; if its Rank-1 and mAP fall to the level of a standard CNN baseline under the same protocol, the claim that simple Transformer-based modality confusion yields robust modality-invariant features is not supported.","supporting_citations":[],"review_version":1}