{"id":"ebc7e3ad-b628-4ef0-95db-8795fa70c51a","arxiv_id":"2412.18874","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"IUST_PersonReId is a new benchmark for person re-identification in modest attire where state-of-the-art models show substantially lower accuracy.","lead":"A new person re-identification (ReID) dataset captures 1,847 identities in Iranian and Iraqi settings, emphasizing modest attire such as hijabs. Standard ReID models show large accuracy drops on this dataset compared to Western and East Asian benchmarks, indicating a cultural gap in current ReID systems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's causal claim that modest attire drives the ReID performance drop is not supported by controlled evidence; IUST_PersonReId differs from Market1501/MSMT17 in cameras, occlusion, and collection setting, and within the dataset attire is never measured separately from visibility.","rationale":"The reader's verdict is conditional, and my stress-test supports that conditionality. The central scientific contribution is the dataset and the documented performance gap; both are credible and useful. The dataset is substantial, the annotation pipeline is described in detail (CVAT, multi-tracker pre-annotation, CVLab-ReID-Tool), and the face-blurring experiment is a practical check. However, the title's causal emphasis—modest attire and cultural clothing as the cause of the drop—is not pinned down by any experiment. Tables 5 and 6 show that large domain gaps exist even among conventional benchmarks, and the within-dataset gender analysis in Tables 7 and 8 is confounded by an uneven visibility distribution. I do not read this as internal inconsistency or as an implausible hypothesis; it is simply an unsupported causal attribution. A within-dataset attire label and visibility-matched evaluation would settle it. Minor issues (Table 1 says 19 cameras while Table 2 says 17; no error bars; data URL validity) are secondary. No verdict change is needed.","tokens_in":13330,"tokens_out":4321,"duration_ms":41072,"concrete_test":"Label each test-set image (or at least every query) with an attire category (hijab/modest versus non-modest, using the public face-blurred release) and a visibility level from Section 4.4.3. Restrict evaluation to a matched subset where gender, number of cameras per identity, and visibility proportions are equal across attire categories, then recompute Rank-1 and mAP for SOLIDER and CLIP-ReID. If the attire-specific gap disappears or shrinks to noise, the modest-attire attribution in Section 4.2 is unsupported; if a large gap remains within visibility-matched strata, the claim is substantially strengthened. As a secondary check, re-run the same fine-tuning protocol on a random subset of IUST_PersonReId matched to Market1501's training-set size to rule out dataset-size effects.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing assumption is in Section 4.2: 'This decline demonstrates the challenges introduced by the new domain, where individuals wear modest attire characteristic of Iranian culture.' The evidence is Table 3's large mAP drop, but IUST_PersonReId differs from Market1501 and MSMT17 in several uncontrolled dimensions: 17/19 cameras versus 6/15; 1,847 identities and 117,455 images versus 1,501/32,668 and 4,101/126,441; locations (mosque, Arbaeen procession, markets, campus) with different occlusion and viewpoint distributions; and a different collection/annotation pipeline (multi-tracker pre-annotation, BRISQUE quality selection, 250 ms temporal sampling). Cross-dataset Tables 5 and 6 show large drops even between existing benchmarks—for SOLIDER, Market-to-MSMT is 16.95 mAP and Market-to-Duke is 52.90—so an additional drop on IUST is not by itself evidence that modest attire is the cause. The within-dataset gender evidence is also confounded: Section 4.4.3 and Table 8 show female queries are mostly occluded (74 female versus 17 male) while male queries are mostly clear (401 male versus 83 female), so the female performance gap in Table 7 could reflect visibility rather than hijab use. No image-level label of 'modest attire' appears in the released dataset, so the proposed mechanism is never directly tested. The claim is plausible but empirically underdetermined; if it fails, the paper's title-level conclusion about cultural attire is weakened, although the dataset itself remains valuable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces IUST_PersonReId, a person re-identification dataset with 1,847 identities and 117,455 images collected from Iranian and Iraqi settings, with emphasis on modest attire such as hijabs. The authors fine-tune SOLIDER and CLIP-ReID on the dataset, report substantial mAP drops compared with Market1501 and MSMT17, and interpret these drops as evidence that modest attire and cultural clothing create distinct ReID challenges. The paper also presents sequence-based majority voting, cross-dataset transfer experiments, gender-based analyses, visibility analyses, and face-blurring/fusion experiments, and it releases the dataset and an annotation tool.","tokens_in":13637,"tokens_out":3315,"duration_ms":33539,"significance":"If the central attribution is substantiated, the dataset would fill a real gap: existing ReID benchmarks are dominated by Western and East Asian clothing, and a publicly available, culturally specific dataset with trained baselines would be a useful resource for fairness and domain-generalization research. The paper's strengths include the public release of the dataset and annotation tool, careful multi-stage annotation with quality filtering, and the observation that sequence-based voting improves performance on this data. However, the paper's title-level claim that modest attire drives the performance drop is currently supported only by uncontrolled cross-dataset comparisons, so the significance of the scientific conclusion is conditional on additional controlled evidence.","major_comments":[{"comment":"The train/test split is described as a temporal division by video duration, not an identity-disjoint split. In standard ReID evaluation, identities in the training set must be disjoint from those in the gallery and query sets. If the first 75% of video footage and the remaining 25% can contain the same tracked person at different times, the evaluation protocol is not identity-disjoint and the reported numbers are not comparable with standard benchmarks. Please clarify whether any identity appears in both subsets, and if so, re-split by identity.","section":"Section 4.1"},{"comment":"The claim that the performance drop 'demonstrates the challenges introduced by the new domain, where individuals wear modest attire characteristic of Iranian culture' is not supported by the presented evidence. IUST_PersonReId differs from Market1501 and MSMT17 in camera count (17 vs. 6 and 15), identity count, image count, collection locations (mosque, Arbaeen procession, market, campus), camera types, and annotation pipeline. Tables 5 and 6 themselves show large cross-dataset drops among existing benchmarks (e.g., SOLIDER Market-to-MSMT is 16.95 mAP), so an additional drop on IUST_PersonReId does not isolate attire as the cause. A controlled comparison is needed, such as matching visibility/occlusion distributions, evaluating on a non-modest-attire subset from the same cameras, or using attire labels as a covariate.","section":"Section 4.2, Tables 3, 5, 6"},{"comment":"The gender performance analysis is confounded with visibility. Table 8 shows that 74 of 91 occluded queries are female, whereas 401 of 484 clear queries are male. The lower female performance in Table 7 could therefore be explained by occlusion rather than by hijab or modest attire per se. Please report performance stratified by both gender and visibility category, and avoid attributing the aggregate female deficit to cultural attire without controlling for the visibility distribution.","section":"Section 4.4.3, Table 8"},{"comment":"The released dataset does not contain an image-level label for 'modest attire' or 'hijab,' so the proposed mechanism is never directly tested. The paper infers modest attire from the collection context and from gender, but this is not a measurement. To support the central claim, the authors should either release attire annotations or design an experiment in which attire is explicitly varied while other factors are held fixed.","section":"Section 3.5 and Section 4.4.3"}],"minor_comments":[{"comment":"The dataset name is written inconsistently as 'IUST PersonReId', 'IUSTPersonReId', and 'IUST PersonReId' in the abstract, Table 1, and the main text; please standardize the spelling.","section":"Throughout"},{"comment":"'CV AT' should be 'CVAT', and 'builtin trackers' should be 'built-in trackers'.","section":"Section 3.3"},{"comment":"The y-axis label 'Overal' contains a typo; it should be 'Overall'.","section":"Figure 5"},{"comment":"The face-recognition experiment uses ArcFace but does not specify which pre-trained weights or training data were used for the face model; please provide this information for reproducibility.","section":"Section 4.4.4"},{"comment":"The parenthetical drop values for the diagonal elements are not drops, and the notation '(↑ 0.01)' is confusing; please use a clearer convention or remove the parenthetical values from diagonal entries.","section":"Table 5"}],"recommendation":"major_revision","confidential_remarks":"The paper's core dataset contribution is likely useful and the measurements appear internally consistent, but the central scientific claim about cultural attire is under-supported by the current experimental design. The issues are fixable: the authors can re-run or clarify the identity split, add visibility-controlled and/or attire-labeled analyses, and soften the causal language. I do not see a load-bearing error that forces rejection, but the revision must address the confounds rather than only adding caveats."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing to know: IUST_PersonReId is a serious dataset contribution, and the paper's central causal claim about modest attire is not backed by controlled evidence. Those two statements can both be true, and a good referee will keep them apart.\n\nWhat's actually new: this is the first ReID dataset aimed at Islamic modest clothing, collected from five real-world settings in Iran and Iraq (campus, market, mosque, Arbaeen procession). 1,847 identities, 117,455 images, 19 cameras. The collection pipeline looks careful: multi-tracker pre-annotation, BRISQUE quality selection, 250ms temporal sampling, a released annotation tool, and face blurring for privacy. The paper also reports sequence-based re-ID, which improves mAP substantially, and cross-dataset transfer tables that include existing benchmarks, which is honest in that it lets the reader see how large domain gaps already are between Market, MSMT, and Duke.\n\nThe soft spot is exactly where the stress-test puts it: Section 4.2 states the performance drop 'demonstrates the challenges introduced by the new domain, where individuals wear modest attire,' but IUST differs from Market1501/MSMT17 in camera count, identity count, occlusion distribution, and collection setting. The cross-dataset tables show that even Market-to-MSMT transfer costs 16.95 mAP for SOLIDER, so an additional drop on IUST cannot be attributed to clothing without a controlled comparison. The gender analysis is confounded too: Table 8 shows female queries are mostly occluded (74 vs 17 male) and male queries mostly clear (401 vs 83 female), so the female gap in Table 7 could be visibility, not hijab. And there is no per-image label of 'modest attire,' so the mechanism is never directly tested.\n\nOther fixable issues: no error bars, the dataset URL in the paper contains a space and does not resolve, and the fine-tuning protocol is under-specified. These matter for a benchmark paper.\n\nThe dataset itself remains valuable even if the modest-attire hypothesis is weakened. The right fix is not rejection; it is a revision that either tempers the causal language to 'domain shift' or adds a controlled experiment that isolates attire from other confounds.\n\nWho should read it: anyone working on ReID domain generalization or fairness evaluation. Send it to review; the resource deserves referee time, and the reviewers can push the authors to separate what they built from what they claim about why it is hard.","headline":"The dataset is a real and useful resource, but the modest-attire explanation for the performance drop is not controlled for; review should focus on separating the resource from the causal claim.","tokens_in":14183,"tokens_out":2748,"would_cite":true,"duration_ms":24770,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces IUST_PersonReId, a benchmark of 1,847 identities from Iranian and Iraqi scenes, and claims that current state-of-the-art re-identification models lose tens of mAP points on modest attire compared with standard…","keywords":["person re-identification","dataset benchmark","modest attire","cultural clothing bias","domain gap","occlusion","gender bias","surveillance"],"falsifier":"Train the same models on the same identities photographed in both modest and non-modest clothing under identical camera, lighting, and occlusion conditions; if the mAP gap between clothing conditions is near zero, the benchmark-level drop is not attributable to attire.","tokens_in":13150,"feed_emoji":"🧕","tokens_out":5797,"duration_ms":52936,"temperature":0.7,"pith_summary":"This paper argues that person re-identification models fail to transfer to societies where modest dress, especially the hijab, is common, and that this failure is a distinct dataset-level problem the field has not addressed. To make the case, it builds IUST_PersonReId, a 1,847-identity benchmark from Iranian and Iraqi settings including a campus, a fruit shop, a hypermarket, a mosque, and the Arbaeen street procession, then tests two state-of-the-art models on it. Both drop sharply: SOLIDER's mAP falls from 93.10 on Market1501 to 42.35, and CLIP-ReID's from 89.67 to 51.60; relative to MSMT17 the drops are 23.01 and 21.74 points. Sequence-based majority voting over multiple frames partially recovers the loss, and gender and visibility ablations show that female identities, more often occluded by head coverings, are the hardest to re-identify. If the paper is right, cultural clothing variation is a measurable and under-represented axis of re-identification generalization, and this dataset provides the resource to train and test against it.","feed_headline":"State-of-the-art person ReID drops on modest-attire data","feed_subtitle":"New benchmark from Iran and Iraq shows why current ReID models fail on hijab-wearing pedestrians, with mAP losses up to 51 points.","key_machinery":"The load-bearing object is the dataset itself: IUST_PersonReId, with 1,847 identities, 117,455 annotated bounding boxes, roughly 19 cameras, and five locations in Iran and Iraq, sampled at 250-millisecond intervals and filtered by the BRISQUE quality measure, then split 75/25 by video duration. Around the dataset, the argument runs through three instruments: single-image fine-tuning of SOLIDER (a semantic self-supervised human representation model) and CLIP-ReID (a vision-language re-identification method), sequence-based re-identification with majority voting over multiple frames of the same identity, and ablations on cross-dataset transfer, gender balance, keypoint-based visibility, and face blurring or fusion. The dataset carries the claim, while the ablations are meant to show that the performance gap is tied to occlusion, female clothing, and limited facial cues rather than only to generic dataset difficulty.","core_discovery":"On its own terms, the paper's central discovery is that person re-identification models trained on existing benchmarks do not generalize to modest Islamic clothing: SOLIDER's mAP is 42.35 on IUST_PersonReId versus 93.10 on Market1501 and 65.36 on MSMT17, and CLIP-ReID's mAP is 51.60 versus 89.67 and 73.32. The paper reads these gaps as evidence that hijab-based modest attire creates occlusions and removes distinctive cues such as hair, body contours, and full face exposure, which current models implicitly rely on. It further finds that using multiple frames of the same person with majority voting substantially raises mAP (SOLIDER from 42.35 to 69.6, CLIP-ReID from 51.6 to 71.0), that female re-identification is harder even after balancing gender proportions, and that blurring faces barely hurts performance while fusing face-recognition distances with body re-identification does not help. The intended upshot is that cultural clothing variation is a genuine and under-served re-identification domain, not a minor annotation artifact.","pith_inferences":["The same cross-dataset protocol could isolate the attire effect by synthetically re-dressing Western-benchmark pedestrians in modest clothing and checking whether mAP drops by comparable amounts; the paper does not run that controlled experiment.","Because the dataset's temporal split keeps training and testing videos non-overlapping in time, it offers a natural testbed for domain adaptation and continual learning under distribution shift, which the paper only partially exploits.","The face-blur and face-fusion outcomes suggest that privacy-aware training objectives could be built directly on this dataset, but the paper does not propose such an objective; that is a direct next step."],"forward_implications":["Re-identification systems deployed in modest-attire regions should be trained or adapted on culturally matched data; cross-dataset testing shows transfer mAP below 14 percent for both models on IUST_PersonReId.","Video and sequence-based re-identification is a practical mitigation: majority voting over tracklets lifts SOLIDER mAP from 42.35 to 69.6 and CLIP-ReID from 51.6 to 71.0, so deployment should use multiple frames rather than single detections.","Benchmarking protocols should report gender- and visibility-disaggregated metrics, because overall mAP hides the sharp female and occluded-subgroup drops that the ablations expose.","Future re-identification models and domain-adaptation methods can use IUST_PersonReId as an explicit cultural-domain target, since current models leave tens of mAP points unclaimed there.","The face-blurring and face-fusion results imply that body-based cues are the reliable signal in this domain, so privacy-preserving re-identification does not have to sacrifice accuracy to avoid using facial identity."],"supporting_citations":[{"why":"Market-1501 is the primary Western/East Asian benchmark against which the paper measures SOLIDER's 50.75-point mAP drop.","marker":"[5]"},{"why":"MSMT17 is the second benchmark against which the paper measures 23.01 and 21.74-point mAP drops for SOLIDER and CLIP-ReID.","marker":"[6]"},{"why":"SOLIDER is one of the two state-of-the-art models whose performance on IUST_PersonReId provides the main evidence of the domain gap.","marker":"[31]"},{"why":"CLIP-ReID is the other evaluated model, and its drop and sequence-based recovery are central to the paper's claims.","marker":"[32]"},{"why":"LUPerson supplies the pre-training initialization that the paper uses when fine-tuning SOLIDER on the new dataset.","marker":"[7]"},{"why":"The CLIP model provides the pre-trained vision-language features that CLIP-ReID builds on in the evaluations.","marker":"[33]"},{"why":"AlphaPose keypoint extraction is used to categorize queries into clear, partial, and occluded visibility, supporting the occlusion and hijab-related analysis.","marker":"[35]"}],"fun_headline_variants":["ReID models tumble 50 points on modest-attire benchmark","Hijab-wearing pedestrians break state-of-the-art ReID","New Iranian benchmark shows cultural blind spot in person re-ID","Modest clothing causes major mAP drops in person re-identification","Temporal cues recover ReID accuracy lost on modest attire"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument assumes the large mAP drop is caused by modest attire, but IUST_PersonReId also differs from Market1501 and MSMT17 in camera count, scene types, occlusion distribution, and identity count, and no experiment controls for those variables.","fun_headline_variants_meta":{"raw":{"variants":["ReID models tumble 50 points on modest-attire benchmark","Hijab-wearing pedestrians break state-of-the-art ReID","New Iranian benchmark shows cultural blind spot in person re-ID","Modest clothing causes major mAP drops in person re-identification","Temporal cues recover ReID accuracy lost on modest attire"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000204,"raw_usage":{"total_tokens":1459,"prompt_tokens":1081,"completion_tokens":378,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":697,"completion_tokens_details":{"reasoning_tokens":291}},"tokens_in":697,"tokens_out":378,"duration_ms":3838,"temperature":1.0,"reasoning_tokens":291,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:21:28.471904+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the same models on the same identities photographed in both modest and non-modest clothing under identical camera, lighting, and occlusion conditions; if the mAP gap between clothing conditions is near zero, the benchmark-level drop is not attributable to attire.","supporting_citations":[{"cited_title":"Zheng, L","cited_arxiv_id":null,"evidence_quote":"Market-1501 is the primary Western/East Asian benchmark against which the paper measures SOLIDER's 50.75-point mAP drop."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"MSMT17 is the second benchmark against which the paper measures 23.01 and 21.74-point mAP drops for SOLIDER and CLIP-ReID."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"SOLIDER is one of the two state-of-the-art models whose performance on IUST_PersonReId provides the main evidence of the domain gap."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"CLIP-ReID is the other evaluated model, and its drop and sequence-based recovery are central to the paper's claims."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"LUPerson supplies the pre-training initialization that the paper uses when fine-tuning SOLIDER on the new dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"AlphaPose keypoint extraction is used to categorize queries into clear, partial, and occluded visibility, supporting the occlusion and hijab-related analysis."}],"review_version":1}