{"id":"5c3a94d2-e212-42a7-b56c-31a256ff3b7b","arxiv_id":"2411.17305","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"iCarB is a new public multimodal biometric dataset collection for face, fingerprint, and voice recognition captured from 200 consenting drivers inside a car.","lead":"Researchers introduce iCarB, a set of three biometric datasets (face, fingerprint, and voice) collected from 200 volunteers inside a car, with detailed protocols and metadata. These datasets aim to advance in-vehicle driver recognition and multimodal biometric research by providing a large public in-car collection.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The three iCarB datasets intentionally lack cross-modal subject correspondences, so genuine same-person face+fingerprint+voice driver recognition cannot be evaluated; only synthetic pseudo-identities are possible, weakening the central multimodal utility claim.","rationale":"The reader's verdict correctly identifies missing baseline recognition experiments and reported data-quality issues as risks to the utility claim. I agree with that concern. However, the more fundamental and less easily fixed issue is the deliberate absence of cross-modal identity correspondence. The paper frames the contribution as a multimodal benchmark for driver recognition, yet the released data prevent any genuine same-person multimodal evaluation. This is not a speculative external critique; Section I states it directly and Section III.D shows that all provided protocols are single-modality. The pseudo-identity workaround is offered as a way to train fusion algorithms, but it does not establish that the datasets can validate actual driver recognition, because the paired modalities come from different physical subjects. Adding baseline unimodal recognition results would improve the paper but would not resolve this gap. The verdict should remain CONDITIONAL, with the condition being that the authors either provide a principled cross-modal identity linkage (with privacy safeguards) or explicitly reposition the contribution as three independent unimodal in-vehicle datasets plus a synthetic pseudo-identity fusion benchmark, rather than a true multimodal driver-recognition resource.","tokens_in":8594,"tokens_out":4463,"duration_ms":47200,"concrete_test":"Download the released iCarB metadata and protocol files and check whether any subject identifier or metadata field can link a physical person across iCarB-Face, iCarB-Fingerprint, and iCarB-Voice. Then attempt to construct a same-person multimodal verification set by forming enrollment/probe pairs that use face, fingerprint, and voice data from the same subject. If no cross-modal identity correspondence exists or can be recovered from the provided metadata, the central multimodal driver-recognition claim is unsupported and the paper must be re-scoped to three independent unimodal datasets with pseudo-identity fusion as a separate synthetic benchmark.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that iCarB provides the largest multimodal in-vehicle biometric dataset for advancing driver recognition. For that claim to hold, a researcher must be able to evaluate biometric matching across the three modalities for the same person. Section I explicitly states that the identities of data subjects do not correspond across iCarB-Face, iCarB-Fingerprint, and iCarB-Voice (e.g., ID 1 is a different person in each dataset). No cross-dataset identity mapping is provided. Consequently, no released protocol can enroll a face, fingerprint, and voice sample from the same individual and probe them as one driver identity. The paper's proposed remedy is to \"create multimodal pseudo-identities\" by pairing unrelated subjects across modalities. That procedure does not test real multimodal driver recognition; it tests fusion on artificially linked biometric samples from different people. The paper does not state or validate the assumption that score-level fusion behavior on such pseudo-identities matches genuine same-person multimodal data. This is the most load-bearing weakness because it affects the central multimodal contribution independently of the data-quality concerns in Section IV: even perfectly clean per-modality data would not support the claimed multimodal driver-recognition benchmark without a true cross-modal identity link. Section III.D further confirms that all provided protocols are per-modality and single-sensor, with no cross-modal protocol included.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces three datasets (iCarB-Face, iCarB-Fingerprint, iCarB-Voice) of biometric data collected inside a car from 200 consenting volunteers per modality, with detailed capture protocols, metadata, and predefined evaluation protocols. The authors claim these are the largest and most diverse publicly available in-vehicle biometric datasets, with iCarB-Fingerprint being the first public in-vehicle fingerprint dataset. They also propose using the three datasets together via 'multimodal pseudo-identities' to train and test fusion algorithms. The paper provides file counts (3,600 face videos, 9,521 fingerprint images, 36,000 voice samples), internal consistency of protocols, and documented limitations.","tokens_in":8847,"tokens_out":5319,"duration_ms":50187,"significance":"If the datasets are usable for recognition, they would be a valuable resource for the biometrics community: they are ethically sourced (explicit consent, IRB approval), gender-balanced, span the Fitzpatrick scale, include multiple controlled and uncontrolled conditions, and are accompanied by protocols and metadata for bias and robustness studies. The explicit documentation of data quality limitations is commendable. However, the absence of any recognition baselines and the lack of cross-modal identity correspondence substantially weaken the central claims of benchmarking utility and multimodal driver recognition.","major_comments":[{"comment":"The identities of data subjects do not correspond across the three datasets, so the iCarB datasets do not support evaluation of genuine same-person multimodal (face+fingerprint+voice) recognition. The proposed 'multimodal pseudo-identities' pair unrelated subjects and are not validated as a substitute for real multimodal fusion; the paper gives no evidence that fusion behavior on pseudo-identities matches that on authentic multimodal data. This is a load-bearing weakness because the title and abstract frame the datasets as supporting driver recognition across three modalities. The authors should either release a subset with true cross-modal correspondences or substantially revise the multimodal utility claim.","section":"Section I, Multimodality bullet; Section III.D"},{"comment":"No recognition experiments are reported for any of the three modalities. The paper asserts that the datasets can be used to evaluate and benchmark face, fingerprint, and voice recognition systems, but does not demonstrate that the data support such evaluation. Section IV documents overexposed and underexposed face videos, missing fingerprint images for older/dry-skin subjects, and wind-distorted voice recordings; without baseline results, a reader cannot judge whether standard matchers achieve usable accuracy on these data. At least one baseline experiment per modality (e.g., a face verification network, a minutiae-based fingerprint matcher, an i-vector/x-vector speaker verification system) is needed to substantiate the datasets' benchmarking value.","section":"Sections II and IV"},{"comment":"The claim that iCarB constitute 'the largest and most diverse publicly available in-vehicle biometric datasets' is not substantiated. The comparison in Section I is limited to two datasets (VFPAD and 3DMAD) and only considers subject counts; no comprehensive survey of in-vehicle biometric datasets is provided, and 'diversity' is asserted from demographic metadata without quantitative measures. The claim should be supported by a systematic comparison with prior work, or softened to 'among the largest' with a precise statement of what was compared.","section":"Abstract and Section I"}],"minor_comments":[{"comment":"The word 'renumerated' should be 'remunerated'.","section":"Section II.B"},{"comment":"The total of 180 voice samples per subject is initially confusing because the protocol lists 90 condition-sentence combinations; the text should explicitly state that each sentence was captured simultaneously by two microphones, yielding two files per sentence.","section":"Section II.B (Voice)"},{"comment":"It is unclear whether the '200 volunteers' are the same individuals across the three datasets or three distinct cohorts. Section II says '200 volunteers for each piece of biometric data,' which suggests different subject pools; this should be clarified, as it affects the interpretation of the demographic statistics in Section III.","section":"Section II.B and Section III"},{"comment":"The phrase 'no intra-sensor variations were considered' is ambiguous; it likely means that each protocol uses data from a single sensor and does not mix sessions, but this should be restated more clearly.","section":"Section III.D"},{"comment":"For the microphones, the sampling rate and bit depth are not reported; these are essential for reproducible speaker recognition experiments.","section":"Table III"},{"comment":"The statement that the near-infrared camera has 'two illumination circuits' is not explained; specify whether they are used for different purposes (e.g., still frames vs. video) or are redundant.","section":"Section II.A"},{"comment":"The abbreviation 'VFPAD' is used without expansion; define it at first use (e.g., 'Vehicle Face Presentation Attack Detection').","section":"Section I"}],"recommendation":"major_revision","confidential_remarks":"The paper reports a dataset collection effort with careful protocols and a willingness to document limitations, which is a strength. The two principal concerns are (a) the absence of any recognition baselines, which makes the claimed benchmarking utility unverified, and (b) the cross-modal identity mismatch, which undercuts the 'multimodal' aspect of the contribution. The former is addressable by adding experiments; the latter may require either releasing a small same-person subset or reframing the contribution as three co-registered unimodal datasets rather than a multimodal driver-recognition benchmark. I would also advise the editor to request a more systematic comparison supporting the 'largest and most diverse' claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the iCarB paper. The datasets are real, documented, and likely useful as per-modality in-car biometric benchmarks. The paper's multimodal claim, however, is materially weaker than the abstract suggests, because subject identities are deliberately not linked across the three datasets. That is the one thing to know before citing or planning experiments.\n\nActual strengths: iCarB-Fingerprint is, as far as I can tell, the first public in-vehicle fingerprint dataset. The acquisition is consent-based under Swiss data protection rules, with 200 subjects, a 50/50 gender split, Fitzpatrick skin types I-VI, and ages from 18 to 60+. File counts check out (3,600 face videos, 9,521 fingerprint images, 36,000 voice samples), filenames are parseable, and each dataset ships pre-defined protocols plus metadata. The limitations section is candid about overexposed and underexposed face captures, missing fingerprints for older or dry-skin subjects, and wind-distorted voice samples. That honesty is a real plus.\n\nSoft spots, in order. First and most importantly, the multimodal promise. Section I states plainly that ID 1 in iCarB-Face is not the same person as ID 1 in iCarB-Fingerprint or iCarB-Voice, and no cross-modal identity map is released. So no protocol can enroll face plus fingerprint plus voice for a single driver and probe that identity. The proposed workaround, multimodal pseudo-identities, pairs unrelated people across modalities. That can test score-level fusion mechanics, but it cannot validate genuine same-person multimodal recognition, and the paper never tests the assumption that fusion behavior on pseudo-identities matches real multimodal data. This weakens the central contribution more than the data-quality issues do; even perfectly clean per-modality data would not fix it. Second, there are no recognition baselines. A dataset paper can survive that if the protocols are solid and the data files are inspectable, but here it means no evidence about how usable the noisy samples actually are. Third, \"largest and most diverse\" is asserted rather than demonstrated with a systematic comparison; it may well be true, but the evidence is thin.\n\nVerdict: the per-modality datasets deserve serious engagement, and a good referee should push for baseline experiments and an honest reframing of the multimodal scope. For readers working on driver recognition or demographic bias, the per-modality resources are still valuable.","headline":"Useful per-modality in-car biometric datasets, but the multimodal driver-recognition story is materially weakened because identities are deliberately unlinked across modalities.","tokens_in":9361,"tokens_out":2650,"would_cite":true,"duration_ms":43749,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The iCarB project presents three new biometric datasets—face, fingerprint, and voice—collected inside a car from 200 consenting volunteers, claiming they are the largest and most diverse publicly available in-vehicle biometric datasets…","keywords":["in-vehicle biometrics","driver recognition","face recognition","fingerprint recognition","voice recognition","multimodal biometrics","biometric datasets","demographic bias"],"falsifier":"Compute verification accuracy on iCarB-Face's indoor, no-accessory, frontal protocol with any standard face recognition system; if the equal error rate is not clearly better than chance, the dataset would fail to support the claimed benchmarking utility.","tokens_in":109,"feed_emoji":"🚗","tokens_out":5567,"duration_ms":106565,"temperature":0.7,"pith_summary":"This paper presents three biometric datasets collected inside a car from 200 consenting volunteers: face videos, fingerprint images, and voice samples. The authors claim these are the largest and most diverse publicly available in-vehicle biometric datasets, and that the fingerprint set is the first public in-car fingerprint dataset. The data was gathered in indoor and outdoor sessions with deliberately introduced variations such as masks, sunglasses, hats, dry or moist fingers, and background noise, to imitate real driver-recognition conditions. If these claims hold, the datasets give the research community a consent-based, multimodal benchmark for evaluating and comparing recognition systems and for studying demographic and environmental bias.","feed_headline":"Largest in-car biometric dataset covers face, fingerprint, voice","feed_subtitle":"Two hundred consenting volunteers, three modalities, built-in noise for driver-recognition benchmarking.","key_machinery":"The load-bearing mechanism is the data-collection protocol, which imposes controlled variations and records metadata per capture. Each subject is assigned an anonymous ID, and every file's name encodes the session, sensor, and variation, such as accessory and action for face, scanner and finger condition for fingerprint, and microphone, window state, and noise type for voice. The protocol also defines structured evaluation protocols in the form of enrolling and probing CSV files, so that recognition systems can be tested globally or per-variation. This design is what makes the three datasets usable as a benchmark and is the basis for the claimed breadth and diversity.","core_discovery":"The central discovery is the datasets themselves: iCarB-Face, iCarB-Fingerprint, and iCarB-Voice, collected from 200 volunteers (100 male, 100 female) seated in the driver's seat, with skin tones spanning the full Fitzpatrick scale and ages from 18 to 60+. iCarB-Face contains 3,600 near-infrared face videos covering indoor and outdoor lighting, with and without masks, hats, and sunglasses, and with different head movements. iCarB-Fingerprint contains 9,521 fingerprint images from two scanners (one thermal, one optical) under normal, dry, moist, hot, and cold finger conditions. iCarB-Voice contains 36,000 voice samples recorded by two microphones under noiseless, traffic, music, discussion, and impulsive-noise conditions, with windows open and closed. The paper argues that because all three modalities share the same automotive environment and come with evaluation protocols and demographic and environmental metadata, the datasets enable unimodal and multimodal benchmarking, presentation-attack development, and bias analysis that existing single-modality in-vehicle datasets cannot.","pith_inferences":["If the datasets live up to their quality claims, they could serve as a common testbed for comparing unimodal versus multimodal driver recognition, which the field currently lacks.","Because the three datasets are not identity-linked across modalities, researchers cannot directly study cross-modal matching; the paper's implicit assumption that pseudo-identities suffice for fusion evaluation is a limitation worth testing.","A natural next step is to run baseline recognition experiments, which the paper does not provide, to establish the achievable performance on each protocol; without these, the claimed benchmarking value remains unquantified.","The deliberate noise conditions, such as dry fingers and windy voice captures, could be used to stress-test the robustness of recognition systems in ways that standard datasets do not."],"forward_implications":["Researchers can benchmark face, fingerprint, and voice recognition systems on data from the same automotive environment, using the provided per-variation protocols.","The datasets support construction of multimodal pseudo-identities for training and testing fusion algorithms, since IDs are not linked across modalities.","The included metadata enables demographic and environmental bias evaluations, for instance by skin type, age, gender, weather, or fingerprint condition.","The data can be used to generate presentation attacks for evaluating presentation-attack-detection algorithms on face, fingerprint, and voice modalities.","The fingerprint dataset, if confirmed as the first public in-vehicle one, fills a gap for automotive biometric research."],"supporting_citations":[{"why":"Describes VICAR, an audio-visual speech corpus in a car environment, cited as a prior multimodal in-vehicle dataset whose link is broken, serving as a comparison point for the claimed first multimodal in-car dataset.","marker":"[1]"},{"why":"Presents VFPAD, a face presentation attack detection dataset in NIR, cited as an existing in-vehicle biometric dataset with fewer subjects (40) and only face modality.","marker":"[2]"},{"why":"Introduces 3MDAD, a multimodal driver distraction dataset, cited as another in-vehicle dataset with only face data and smaller subject counts (50 daytime, 19 nighttime).","marker":"[3]"}],"fun_headline_variants":["iCarB: in-car biometric datasets for face, fingerprint, voice","Largest in-vehicle biometric datasets now include fingerprints","First in-car fingerprint dataset joins face and voice biometrics","200 diverse volunteers, three biometrics, one car environment"],"cache_read_input_tokens":11520,"weakest_assumption_plain":"The claim that these datasets will be valuable for biometrics research assumes the collected data is actually usable for recognition benchmarking, yet the paper itself reports overexposed and underexposed face videos, missing fingerprint images for older and dry-skin subjects, and wind-distorted voice recordings, and it provides no baseline recognition experiments.","fun_headline_variants_meta":{"raw":{"variants":["iCarB: in-car biometric datasets for face, fingerprint, voice","Largest in-vehicle biometric datasets now include fingerprints","First in-car fingerprint dataset joins face and voice biometrics","200 diverse volunteers, three biometrics, one car environment"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00041,"raw_usage":{"total_tokens":2211,"prompt_tokens":1118,"completion_tokens":1093,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":734,"completion_tokens_details":{"reasoning_tokens":1024}},"tokens_in":734,"tokens_out":1093,"duration_ms":10048,"temperature":1.0,"reasoning_tokens":1024,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:14:41.914753+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compute verification accuracy on iCarB-Face's indoor, no-accessory, frontal protocol with any standard face recognition system; if the equal error rate is not clearly better than chance, the dataset would fail to support the claimed benchmarking utility.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes VICAR, an audio-visual speech corpus in a car environment, cited as a prior multimodal in-vehicle dataset whose link is broken, serving as a comparison point for the claimed first multimodal in-car dataset."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Presents VFPAD, a face presentation attack detection dataset in NIR, cited as an existing in-vehicle biometric dataset with fewer subjects (40) and only face modality."},{"cited_title":"Kotwal, S","cited_arxiv_id":null,"evidence_quote":"Introduces 3MDAD, a multimodal driver distraction dataset, cited as another in-vehicle dataset with only face data and smaller subject counts (50 daytime, 19 nighttime)."}],"review_version":1}