{"id":"496a1b7e-94cb-4c66-b1cf-f1148a7238eb","arxiv_id":"2508.10786","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"A face-liveness method is claimed, but the manuscript body is an unrelated math paper, leaving the claim unsupported and unverifiable.","lead":"This submission's abstract describes a cooperative face-liveness detection method using optical flow. The included full text is an unrelated mathematics paper on Ricci shrinkers, so the described computer-vision work is absent from the manuscript.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Full text is a Ricci-shrinker mathematics paper, not the described optical-flow liveness method; no architecture, dataset, or evaluation appears anywhere, so the abstract's claim is unsupported.","rationale":"The reader's verdict is UNVERDICTED because the submitted full text is a different mathematical paper with no CV methods or results. This stress-test agrees with that verdict, but identifies the most load-bearing concern as the complete absence of any described system or evidence for the abstract's claim, rather than the specific optical-flow/replay-discrimination assumption highlighted in the reader's weakest_assumption. The absence of the method is more fundamental: even if the replay-flow assumption had been addressed, there is no architecture, training protocol, dataset, or experiment to evaluate. The concrete test is a straightforward text search to confirm the manuscript body does not contain the claimed CV content, and a check for possible updated versions. Since the verdict is already UNVERDICTED and this concern does not suggest a different disposition, the recommended verdict remains UNCHANGED.","tokens_in":2436,"tokens_out":2326,"duration_ms":24732,"concrete_test":"Retrieve the published arXiv:2508.10786 PDF and perform a full-text search for the strings 'liveness', 'optical flow', 'presentation attack', 'classifier', and 'dataset'. If none appear in the main text, the document does not contain the experimental method claimed by the abstract. As a secondary check, inspect the arXiv abs page for a replacement/updated version; if a corrected paper exists, evaluate whether its experiments actually compare cooperative approaching-face flow against replay attacks.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The manuscript's abstract promises a cooperative, video-based face liveness detection system: users move their face toward the camera, optical flow is estimated by a neural network, and a classifier separates live faces from printed photos, screens, masks, and video replays. The full text, however, is an unrelated mathematics article by different authors (Bertellotti and Buzano, arXiv:2508.10790v1, 'Geometric Structure of Ends of Ricci Shrinkers'). None of the components necessary for the CV claim are present: no optical-flow estimator, no classifier architecture, no interaction protocol, no dataset, no training details, no evaluation metrics. The only equations (e.g., Ric_g + ∇²f = ½g in §1.1) concern Riemannian geometry. Because the described method and its empirical support are entirely absent, the central claim cannot be checked. The load-bearing condition for the abstract's statement—existence of a system and evidence that it discriminates attack types—is not met by this document.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The abstract announces a cooperative video-based face liveness detection method built around a 'controlled approaching face protocol': the user is asked to move their frontal-oriented face toward the camera, neural optical flow is estimated, and a classifier separates live faces from printed photos, screen displays, masks, and video replays. The substantive text supplied under 'Full Text', however, is a mathematics paper on the geometric structure of ends of Ricci shrinkers by Bertellotti and Buzano (arXiv:2508.10790), with no mention of face liveness, optical flow, presentation attacks, datasets, or experiments. The central claim—that the proposed method significantly improves discrimination between genuine and attacked faces—is therefore completely unsupported by the manuscript as submitted.","tokens_in":2717,"tokens_out":3251,"duration_ms":37261,"significance":"If substantiated, the proposed cooperative protocol could be a genuinely useful contribution to face anti-spoofing: exploiting the expansion pattern of a live face approaching the camera is a plausible source of 3D/volume cues and could complement passive methods. However, the submitted manuscript provides no substantiation whatsoever. There is no model architecture, no training procedure, no protocol specification, no dataset, no baseline comparison, and no evaluation metric. The only rigorous content is an unrelated differential-geometry article. The manuscript therefore has no current scientific value for the computer-vision audience and cannot be assessed for correctness, novelty, or reproducibility.","major_comments":[{"comment":"The body of the manuscript is not the described liveness-detection paper. The abstract promises a neural optical-flow estimator and a liveness classifier; the full text is 'Geometric Structure of Ends of Ricci Shrinkers' and its principal equation is Ric_g + Hess f = 1/2 g. None of the claimed components—optical flow, classifier, attack types, interaction protocol—appear anywhere in the full text. This is a load-bearing failure: the central claim cannot be checked because the method itself is absent.","section":"Full text, §1 (Eq. 1.1)"},{"comment":"The abstract asserts 'significantly improving discrimination between genuine faces and various presentation attacks' and states the method is more reliable than passive methods. The manuscript contains no dataset, no experimental setup, no baselines, and no results (no APCER, BPCER, EER, ROC, or any other metric). For a claim of empirical superiority, evaluation is mandatory. As submitted, the assertion is entirely unsupported.","section":"Abstract (claim of improved discrimination)"},{"comment":"The core innovation is described only in the abstract as a two-sentence interaction scenario. The full text contains no definition of the protocol: no instructions for the user motion, no camera configuration, no video duration, no frame-rate or temporal-window choices, and no description of how flow fields and RGB frames are combined. The method is therefore neither reproducible nor testable.","section":"Abstract ('controlled approaching face protocol')"},{"comment":"No neural network architecture, loss function, training procedure, or inference details are given for either the optical flow estimator or the final classifier. Without these, it is impossible to verify whether the flow features provide independent information or simply memorize training-set idiosyncrasies. The absence of any equations or pseudocode describing the liveness pipeline makes the manuscript impossible to evaluate even as a theoretical proposal.","section":"Full text, throughout (unspecified models)"}],"minor_comments":[{"comment":"The reference list consists entirely of mathematics papers on Ricci flow and Ricci shrinkers. There is no citation of prior work on face anti-spoofing, optical-flow-based motion analysis, or presentation-attack detection, which would be expected even in a short conference-style submission.","section":"Full text, references"},{"comment":"The displayed equation in §1 is garbled in the provided text (e.g., 'Ricg +∇2f = 1 2 g'), and several labels contain placeholder characters. This is secondary to the substantive mismatch, but it further obstructs reading.","section":"Full text, formatting"}],"recommendation":"reject","confidential_remarks":"As submitted, the manuscript is not a paper on face liveness detection; the supplied full text is a different mathematics article by different authors. I am not inferring intent, but the editor should verify the uploaded file and the arXiv identifier. The CV claim has no supporting content, so within the scope of this submission there is no fix short of replacing the entire manuscript."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here is my read. The abstract of 2508.10786 promises a cooperative video-based face liveness detection system built on a slow approach-to-camera protocol, neural optical flow, and a classifier separating live faces from prints, screens, masks, and replays. The full text, however, is a mathematics paper by different authors about Ricci shrinkers. So the document as submitted does not contain the method it advertises. I cannot referee the CV claim because there is no architecture, no training detail, no dataset, no evaluation, and no analysis tying the abstract to the body. This is a submission-level failure, not a matter of interpretation.\n\nWhat is actually new, and worth saying, is the abstract's idea. The cooperative protocol—ask the user to move their face toward the camera, then use optical flow to infer facial volume—is not a standard liveness cue in the visibility of the literature to me, and combining it with a learned neural classifier to cover print, screen, mask, and replay attacks is a plausible design. If a real paper were written around that idea, it would be worth a serious referee.\n\nBut on this submission, the soft spots are structural. There is nothing to check. The abstract's claim of 'significantly improving discrimination' is completely unsupported. The only specific concern I would flag for the eventual real paper is the replay attack: a video of a face moving toward the camera produces nearly the same optical expansion as a live face. The abstract does not acknowledge this, let alone provide an argument or experiment showing the learned classifier can separate them. That would be the load-bearing test.\n\nThe Ricci shrinker paper itself may be perfectly good mathematics, but it is not this abstract's paper, and the authors are not even the same. My recommendation for the editor: desk-reject or return to authors to resubmit with the correct manuscript. The CV idea has some promise, but this version is not reviewable. I would not cite this document for anything, and I would not bring it to our reading group as-is.","headline":"The abstract describes a face-liveness method; the full text is an unrelated math paper, so the submission cannot be evaluated for the claimed CV work.","tokens_in":3098,"tokens_out":1485,"would_cite":false,"duration_ms":18707,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The authors claim that a cooperative protocol—having a user slowly move their face toward the camera—lets a neural optical-flow and RGB classifier read 3D facial volume cues, improving liveness discrimination against printed photos, screen","keywords":["face liveness detection","presentation attack","optical flow","cooperative protocol","facial volume","video replay attack","neural classifier","spatio-temporal features"],"falsifier":"Record a high-frame-rate video of a real face slowly approaching a camera, display that video on a phone or tablet, and move the display toward a second camera at the same physical speed; run the trained classifier. If the replay clips are accepted at a rate close to genuine clips, the central discrimination claim fails.","tokens_in":2398,"feed_emoji":"🛡️","tokens_out":5209,"duration_ms":59615,"temperature":0.7,"pith_summary":"This paper is trying to establish that face liveness can be checked by asking the user to move their face slowly toward the camera and analyzing the resulting optical flow. The claim is that this movement makes the 3D shape of a real face visible in the motion pattern, while flat photos, screens, rigid masks, and video replays produce flow fields that a neural classifier can separate from genuine faces. If true, this cooperative protocol would give video-based liveness systems a simple, hardware-free way to reduce presentation-attack success compared with passive methods. The central assumption is that a replay of an approaching face does not fool the classifier.","feed_headline":"Slow face approach defeats photos, screens, masks, and replays","feed_subtitle":"Cooperative optical-flow method extracts 3D facial volume cues that passive liveness checks miss.","key_machinery":"The controlled approaching-face protocol is the load-bearing interaction: it converts liveness detection from observing whatever motion happens to watching a specific, repeatable motion whose optical flow should encode 3D facial volume. The neural optical-flow estimator extracts dense frame-to-frame motion, and the resulting flow maps, concatenated with RGB frames, are fed to a neural classifier that learns spatial-temporal patterns separating live faces from presentation attacks. The protocol is what makes 'volume from motion' available as a discriminative signal.","core_discovery":"The paper's central claim is that depth, or facial volume, can be extracted from motion alone under a controlled interaction: participants are instructed to keep their face frontal and move it steadily toward the camera. In that setting, the optical flow field is not just an overall expansion; because nearby parts of a 3D surface move differently from distant parts, the flow carries parallax information tied to facial geometry. The authors state that neural optical-flow estimation plus a classifier operating on both flow and RGB frames significantly improves discrimination between genuine faces and printed-photo, screen, mask, and video-replay attacks, and that this beats passive methods tha","pith_inferences":["A natural extension is to treat the protocol as a challenge-response: the system could randomize the required speed or direction of approach, forcing an attacker to replay a matching motion; whether such randomization actually increases security is not tested in the paper.","Because the discriminator relies on learned optical flow, its behavior on very close-range faces, motion blur, or atypical skin and lighting is likely to depend on the flow estimator's training distribution; a hand-computed divergence or parallax statistic could offer an interpretable diagnostic or a cheaper baseline.","The same 'volume from controlled motion' cue might transfer to other verification tasks, such as object or document liveness, where a user moves an item toward the camera and depth is inferred from the resulting flow."],"forward_implications":["A liveness check could be administered as a short, natural gesture—'slowly bring your face closer'—requiring no special sensor beyond an RGB camera.","If the flow-based volume cue is as discriminative as claimed, printed photos and rigid masks should be rejected at lower false-acceptance rates than in passive systems, because their motion fields lack the parallax of a genuine 3D face.","The hardest attack class becomes a video replay of someone moving toward the camera; the paper's claim that the classifier still separates such replays is the decisive point to test.","Combining flow and RGB features lets the model exploit both motion geometry and appearance, which should generalize better than appearance-only detectors to unseen attack types."],"supporting_citations":[],"fun_headline_variants":["Optical flow reads face depth from a slow approach","Moving closer reveals real vs fake faces","Face liveness via motion-based 3D volume","Cooperative motion defeats spoof attacks","Slow approach lets optical flow see face depth"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that the optical flow generated by a live 3D face moving toward the camera is different enough from the flow of a flat photo, a screen, a rigid mask, and especially a video replay of an approaching face that a learned classifier can tell them apart; the paper supplies no argument or experiment isolating the replay case.","fun_headline_variants_meta":{"raw":{"variants":["Optical flow reads face depth from a slow approach","Moving closer reveals real vs fake faces","Face liveness via motion-based 3D volume","Cooperative motion defeats spoof attacks","Slow approach lets optical flow see face depth"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000591,"raw_usage":{"total_tokens":2546,"prompt_tokens":621,"completion_tokens":1925,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":365,"completion_tokens_details":{"reasoning_tokens":1857}},"tokens_in":365,"tokens_out":1925,"duration_ms":14437,"temperature":1.0,"reasoning_tokens":1857,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T20:12:56.606819+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record a high-frame-rate video of a real face slowly approaching a camera, display that video on a phone or tablet, and move the display toward a second camera at the same physical speed; run the trained classifier. If the replay clips are accepted at a rate close to genuine clips, the central discrimination claim fails.","supporting_citations":[],"review_version":1}