{"id":"d6d340ed-2b23-4f35-a415-baaf9b5e8db7","arxiv_id":"2607.22302","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A 62,856-sample fMRI dataset of full-HD digital-human face videos plus a geometry-guided video-diffusion decoder that reconstructs facial identity and motion from brain signals.","lead":"Researchers scanned three people watching 2,174 photorealistic digital-human face videos and built a system that reconstructs short face videos from the recorded fMRI signals. The work introduces a 62,856-sample benchmark and reports large gains over prior brain-decoding baselines.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Test-set independence is not established: with 31 shared motion configurations and a small shared face-shape/base-mesh library, held-out videos may reuse training identities/motions, so the reported decoding gains could reflect stimulus-manifold interpolation rather than fMRI decoding of unseen face","rationale":"The paper introduces a genuinely useful controlled dataset and a well-structured decoding framework, and the ablations are internally consistent. However, the strongest empirical claim—reconstruction of unseen identities and dynamics from fMRI—rests on the assumption that the 180-video test split is independent of the training split at the level of the components that define the stimuli. The supplementary material shows those components are few: 31 motion configurations, 136 artistic face shapes reused across surface-attribute variations, and a parametric base-mesh interpolation set. A video-level split is therefore not shown to be an identity-level or motion-level split, so the reported gains could stem from retrieval or interpolation within a low-dimensional stimulus manifold rather than from neural decoding of novel faces. This is the same load-bearing weakness the reader identified, and the proposed split-level re-evaluation would settle whether the concern actually lands. Because the reader's conditional verdict already encodes this uncertainty, I do not move the verdict; I would strengthen the conditions to require the leak-controlled split or explicit overlap statistics before accepting the generalization claim.","tokens_in":27784,"tokens_out":8738,"duration_ms":82413,"concrete_test":"Using the Supp. 1.1 face-shape and motion annotations, construct a split in which no face-shape ID (and ideally no motion configuration) appears in both train and test sets; for example, hold out a subset of the 136 artistic shapes and a subset of the 31 motion types. Retrain/evaluate fMRI2Face and the Tab. 2 baselines on this leak-controlled split. If the ID-CSIM, LMD, and FVD advantages shrink or disappear, the original numbers are inflated by the shared stimulus manifold; if the advantages persist, the decoding result generalizes. At minimum, report the overlap counts between train and test for face-shape IDs and motion IDs.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that fMRI2Face decodes identity and dynamics from fMRI—depends on the 180-video held-out set (Sec. 3.2.4) measuring responses to unseen faces/motions. The paper never states that test identities or motion configurations are disjoint from training, and the supplementary details indicate they are not. Supp. 1.1 reports only 31 motion configurations (16 female, 15 male) and only 136 face shapes for the 1,694 artistic digital humans, with each shape fixed and reused across skin-texture/hairstyle/clothing variations; the 480 parametric videos are generated by interpolating a shared base-mesh library with four fixed hairstyles and two skin textures. A random video-level holdout of 180 among 2,174 therefore almost certainly includes test videos whose face geometry and motion type already appear in training. Under these conditions the appearance-context tokens and 3D DECA control can retrieve or interpolate familiar shape/motion combinations; the reported PSNR/ID-CSIM/FVD gains over baselines need not come from true neural decoding of novel identity/dynamics. The failure cases in Fig. 12 (red hair reconstructed as blonde) are consistent with the model defaulting to common training statistics. This is a missing-support problem: the benchmark's central generalization claim is not backed by a leak-controlled split.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces fMRI-Face, a large fMRI dataset of 62,856 paired samples obtained while three participants watched 2,174 controlled 8-second full-HD digital-human facial videos, and fMRI2Face, a decoding framework that predicts appearance-context tokens and DECA-based 3D facial parameters from fMRI and uses them to condition a pretrained video diffusion transformer. Quantitative results on a held-out set of 180 videos show consistent improvements over MindVideo, NeuroPictor, and MindEye2 across reconstruction, identity, geometry, and motion metrics. The paper also presents ablations, ROI analyses, failure cases, and supplementary details on the stimulus-generation pipeline.","tokens_in":28183,"tokens_out":8134,"duration_ms":69064,"significance":"If the reported results hold, fMRI-Face would be a large, controlled benchmark for dynamic face decoding, and fMRI2Face would demonstrate a useful integration of 3D morphable face control with video diffusion for fMRI-conditioned reconstruction. The paper's strengths are the scale and controllability of the stimulus set, the explicit two-stream design (appearance context plus 3D control), the subject-wise results, and the careful loss ablations for the 3D stream. However, the central generalization claim is not yet adequately supported because the test split may not be independent of the training manifold, and the comparisons lack statistical uncertainty.","major_comments":[{"comment":"The held-out set does not establish independence of test identities/motions. Supp. 1.1 reports only 31 motion configurations (16 female, 15 male) for all 2,174 videos, 136 fixed face shapes for the 1,694 artistic digital humans, and 480 parametric videos generated by interpolating a shared base-mesh library with fixed hairstyles and skin textures. A random holdout of 180 videos therefore almost certainly includes test videos whose face shape and motion type already appear in training. Under these conditions, the appearance-context tokens and the DECA 3D control can retrieve or interpolate familiar shape/motion combinations; the gains in Tab. 2 (e.g., PSNR 18.32 vs 15.71, FVD 82.7 vs 361.3) need not reflect fMRI decoding of unseen identities or dynamics. The authors should specify whether test videos are disjoint in shape and motion by construction, or re-split/re-evaluate with disjoint s","section":"Sec. 3.2.4 / Supp. 1.1"},{"comment":"The headline comparisons are reported as single means without error bars or significance tests. Given only 180 test videos and three subjects, metrics such as FVD and ID-CSIM have substantial sampling noise; Tab. 6 shows subject-level variation but no variance within subject. To support \"consistently improves,\" the authors should report confidence intervals (e.g., bootstrap over test clips) and/or paired significance tests against each baseline, and provide subject-wise baseline numbers rather than only our method.","section":"Sec. 5.3.1 / Tabs. 2 and 6"},{"comment":"The ablations are run only on Subject 1, while the main quantitative claims are means over three subjects (Tab. 6). The contribution of each component (context-token count, 3D control, auxiliary latent, mask) may be subject-dependent. Please provide per-subject ablations, or at least justify why Subject 1 is representative; otherwise the design conclusions are not supported on the full dataset.","section":"Sec. 5.4"},{"comment":"The paper's first contribution is the fMRI-Face dataset, but no data availability statement, URL, or code release is included. Without a concrete release plan, the benchmark cannot be used by the community and the central numbers cannot be independently verified. A dataset paper should state the intended availability and license.","section":"Abstract / Sec. 3"}],"minor_comments":[{"comment":"The paper describes a full-HD dataset but reconstructs at 480×832. Clarify whether ground-truth videos are downsampled before metric computation and how landmark metrics are scaled across resolutions.","section":"Sec. 5.1"},{"comment":"The relationship between 2,174 unique videos, 154 repeated clips, and the reported 2,012 training / 316 test pairs is ambiguous. Specify how repeats were allocated between train and test.","section":"Sec. 3.2.4"},{"comment":"The qualitative figures would benefit from showing the rendered 3D control stream for the same examples, and from including baseline failure cases for comparison.","section":"Figs. 5 and 12"},{"comment":"Parameter-level supervision uses DECA parameters extracted from the GT videos; since DECA is trained on real faces, the paper should report DECA accuracy on digital-human stimuli or otherwise justify the reliability of this supervision.","section":"Sec. 4.4.3"}],"recommendation":"major_revision","confidential_remarks":"The test-split concern is the gating issue: if the authors can demonstrate that the held-out set is disjoint in face shape and motion configuration, or re-run with a disjoint split, the contribution would be much stronger. A retrieval-based upper bound (e.g., nearest-neighbor in training by shape/motion) would also help interpret the reported gains. The paper currently reads more like a strong system description than a verified benchmark, so the revision should focus on split construction and statistical rigor."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a large, carefully built fMRI-face-video dataset and a reasonable decoding pipeline, but the paper never establishes that the 180 held-out videos test truly unseen identities or motions. Given the supplementary details—31 motion configurations total, 136 face shapes reused across videos—the reported gains could come from interpolating a small stimulus manifold rather than decoding novel faces.\n\nWhat is genuinely new: fMRI-Face is the first full-HD (1920×1080) fMRI dataset paired with controllable digital-human face videos, with 62,856 paired samples across three subjects. That is a real step up from VanRullen 2019, HYPER, and NFED. The fMRI2Face framework—fMRI-conditioned appearance tokens plus DECA-rendered 3D control injected into a video diffusion transformer—is a sensible combination, and the ablations show each component helps, at least on subject 1. The ROI attribution and reliability analyses are a nice extra, not fluff.\n\nSoft spots, in order of importance. First is the split. Sec. 3.2.4 says only \"we use 1,994 videos for training and holdout 180 videos for testing.\" The supplementary (1.1) reveals the entire dataset is built from 31 motion configurations and 136 artistic face shapes, with parametric faces interpolating a shared base mesh. A random video-level holdout will therefore place multiple motion/shape combinations from training into the test set. The paper never checks whether test identities/motions are disjoint, and it does not report performance on novel factors. The failure case in Fig. 12 (red hair reconstructed as blonde) is consistent with the model defaulting to training statistics. This is a missing-support problem for the central claim, not a minor nuisance. Second, there are no error bars, significance tests, or subject-wise breakdowns for the baselines; the ablations run on subject 1 only. Third, no data or code is released, so the numbers cannot be checked. The architecture itself is not circular—the control stream is supervised with DECA parameters from targets, but the renderings at test time come from fMRI predictions, so the metric computation is not illicit. The paper's own Limitations section is honest about micro-motions and digital-human realism, but it does not mention the factor-overlap issue.\n\nBottom line: if you work on fMRI decoding or face reconstruction, this is worth reading as a dataset proposal and a proof-of-concept decoder. It does not yet demonstrate decoding of unseen identity/dynamics, and the authors should be pushed to release data and rerun with disjoint factor splits. Yes, send it out; a good referee can force the split and statistics fixes.","headline":"Valuable dataset and a plausible decoder, but the held-out set reuses the same face shapes and motions as training, so the headline decoding claim is not yet supported.","tokens_in":28656,"tokens_out":3643,"would_cite":false,"duration_ms":32547,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces fMRI-Face, the first fMRI dataset paired with controllable full-HD (1920×1080) digital human facial videos, and fMRI2Face, a geometry-guided framework that reconstructs dynamic facial videos from brain activity while p","keywords":["neural decoding","fMRI dataset","face reconstruction","digital human","video generation","3D morphable face model","video diffusion","face perception"],"falsifier":"Compare fMRI2Face's reconstructions against a nearest-neighbor baseline that retrieves the most similar training-video latent for each test fMRI window. If the retrieval baseline matches or beats the reported PSNR, ID-CSIM, and FVD on test clips whose identity or motion type also appears in training, the decoding claim collapses; retraining with identities and motions held out by construction would settle it.","tokens_in":27734,"feed_emoji":"🧠","tokens_out":8325,"duration_ms":65693,"temperature":0.7,"pith_summary":"The paper aims to show that dynamic human faces can be reconstructed from fMRI signals at full-HD resolution when the decoder is given explicit 3D face geometry as a guide. To make this possible it introduces fMRI-Face, a dataset of 62,856 paired fMRI-video samples in which three participants viewed background-free, full-HD digital human faces with controlled identity, expression, and head pose. Its fMRI2Face framework splits decoding into two neural controls: appearance context tokens carrying global identity-related attributes, and morphable 3D facial parameters rendered into geometry-consistent motion guidance, both injected into a pretrained video diffusion model. On every reported metric the framework outperforms three representative decoding baselines, sharply improving identity similarity and temporal coherence. If the result holds, the field gains a controlled benchmark for dynamic face perception and a usable path from brain signals to controllable digital-human video.","feed_headline":"fMRI2Face decodes 1080p face videos from brain scans","feed_subtitle":"A geometry-guided decoder beats prior fMRI-to-face baselines on every reported metric, preserving identity and motion.","key_machinery":"Two complementary fMRI-derived controls carry the argument. Brain-derived Appearance Context is a set of 64 learnable query tokens that a transformer decoder extracts from an fMRI window and feeds into the diffusion transformer's context-conditioning interface, replacing text prompts with brain-derived appearance cues. Morphable 3D Facial Control predicts DECA morphable-face parameters—a standard parametric face model—from the same fMRI window, renders them into geometry-consistent facial control frames, and injects them through a parallel control branch with zero-initialized residual bridges. An auxiliary latent predictor completes visual context outside the rendered face region. The diffus","core_discovery":"The central claim is that fMRI responses from visual cortex carry enough information to reconstruct a perceived digital human's identity and motion, provided the decoder separates appearance from geometry. fMRI2Face predicts fMRI-conditioned context tokens for global appearance and, in parallel, a sequence of DECA morphable-face parameters (shape, expression, pose, albedo, detail, illumination, image transform) that are rendered into a 2D facial guidance video. A frozen video-diffusion transformer is conditioned on both streams through zero-initialized residual bridges and an auxiliary latent completion, producing 480×832 video clips at 15 fps. The paper reports consistent gains over baselin","pith_inferences":["The paper does not state that the 180 held-out videos use identities and motion types disjoint from training; a test split that excludes them by construction would establish whether the decoder generalizes to unseen faces or interpolates within the 31 shared motion configurations.","Because stimuli are background-free digital humans under fixed lighting, the dataset measures decoding inside a synthetic stimulus universe; transferring the same architecture to natural face videos or photographs would test its generality.","The residual design that anchors identity-related parameters to the first frame suggests a clean cross-subject test: ablating the appearance stream should still leave shape and albedo-derived identity cues in the geometry stream, revealing which stream actually carries identity.","The authors list rapid micro-motions as a limitation; combining fMRI with a faster neural recording modality could feed the geometry-control stream with temporal priors and likely recover blinks and subtle expression changes."],"forward_implications":["fMRI-Face provides a controlled benchmark that lets future work isolate how identity, expression, and pose are encoded in visual cortex.","If the reported gains hold, explicit parametric 3D face modeling should become a standard component of fMRI-to-face decoders.","The framework shows that a pretrained video diffusion model can be steered entirely by brain-derived conditions, extending fMRI-conditioned video generation beyond faces.","The ROI analyses suggest early visual areas and motion-selective areas differentially drive appearance versus geometry, giving testable predictions for neuroscience.","The pipeline points toward controllable digital-human animation from brain activity, with potential as a communication channel for people who cannot speak or move."],"fun_headline_variants":["Brain scans become 1080p face videos","fMRI2Face: full-HD faces from brain activity","Geometry-guided AI rebuilds face motion from fMRI","New dataset and decoder for dynamic face reconstruction","Decoding identity and expression from fMRI signals"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The held-out test videos are independent of the training videos, so reported scores measure decoding of unseen faces rather than retrieval from a small stimulus manifold built from 31 shared motion types and a shared base-mesh library.","fun_headline_variants_meta":{"raw":{"variants":["Brain scans become 1080p face videos","fMRI2Face: full-HD faces from brain activity","Geometry-guided AI rebuilds face motion from fMRI","New dataset and decoder for dynamic face reconstruction","Decoding identity and expression from fMRI signals"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0002,"raw_usage":{"total_tokens":1251,"prompt_tokens":820,"completion_tokens":431,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":564,"completion_tokens_details":{"reasoning_tokens":359}},"tokens_in":564,"tokens_out":431,"duration_ms":3989,"temperature":1.0,"reasoning_tokens":359,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T05:10:24.459632+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Compare fMRI2Face's reconstructions against a nearest-neighbor baseline that retrieves the most similar training-video latent for each test fMRI window. If the retrieval baseline matches or beats the reported PSNR, ID-CSIM, and FVD on test clips whose identity or motion type also appears in training, the decoding claim collapses; retraining with identities and motions held out by construction would settle it.","supporting_citations":[],"review_version":1}