{"id":"51fde5ea-57c2-411b-8239-431f3fc75463","arxiv_id":"1908.06607","paper_version":3,"verdict":"REJECT","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A GAN pipeline transfers a source person's upper-body motion and facial expressions to a target person using body keypoints and facial action units as intermediate representations.","lead":"This paper proposes a three-stage GAN pipeline that copies a source person's upper-body movements and facial expressions onto a target person in video. It uses body keypoints and facial action units as intermediate representations; the experimental evidence is limited to two qualitative examples without metrics.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim hinges on unverified cross-identity FAUP-to-landmark transfer; the paper's two qualitative examples do not establish it.","rationale":"The reader's weakest_assumption identifies the same load-bearing condition: the FAUP-to-landmark generator must generalize from target-only training to source FAUP inputs. My stress-test confirms that this is the most critical technical assumption, because every downstream stage (synthesized landmarks, concatenated UBKP-FL input, final video generation) inherits any failure in that mapping. The paper explicitly says source FAUP are used directly, with no adaptation or calibration, and offers no quantitative evidence that the source distribution is compatible. The experimental section is also too thin: one target individual, two driving videos, qualitative Figure 4 only. These are not alternative criticisms so much as two facets of the same problem: the claimed effectiveness is neither theoretically established nor empirically demonstrated. I therefore agree with the reader's REJECT verdict and recommend no change to it.","tokens_in":5130,"tokens_out":3589,"duration_ms":39472,"concrete_test":"Collect several target-subject videos with aligned FAUP and ground-truth facial landmarks; train the Section 4 generator on a restricted subset of target FAUP values (for example, neutral or low-expression clips), then drive it with source FAUP values outside that range, and compare synthesized landmarks to the target's ground-truth landmarks for the same expressions. If per-landmark normalized mean error increases substantially or a pretrained face/identity metric degrades for out-of-range source FAUP, the OOD assumption fails; if errors remain small across mismatched source-target pairs, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 4 trains the FAUP-to-landmark generator G only on the target person's own (FAUP, facial-landmark) pairs, and then states: 'During the transferring phase, we directly use source's FAUP to synthesize target's facial landmark.' The central claim, that the synthesized video matches the source in body motion, expression, and pose, is load-bearing on the assumption that OpenFace FAUP values from the source lie inside the target's FAUP training distribution and that G maps them to landmarks preserving the source expression. FAUP estimates are not identity-invariant in practice; they shift with face shape, pose, illumination, and OpenFace calibration. If source FAUP falls outside the target's training range, G's landmark output will be off-manifold or wrong-expression, and stage 3's G2, trained on target UBKP-FL/video pairs, will receive an out-of-distribution input. The paper provides no domain adaptation, augmentation, OOD analysis, or failure cases. Section 6 reports 'two videos of one individual' with only qualitative Fig. 4; no metrics, baselines, or code are provided. Thus the central claim rests on an unvalidated distribution-shift assumption plus anecdotal evidence.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a three-stage GAN-based pipeline for transferring upper-body motion, pose, and facial expression from a source video to a target person. Stage 1 extracts upper body keypoints (UBKP) and facial action units plus pose (FAUP) from the source video. Stage 2 normalizes the UBKP between source and target, and uses a pix2pixHD-style generator to map source FAUP into target facial landmarks. Stage 3 concatenates the normalized UBKP and generated landmarks into an image and uses a second pix2pixHD-style generator, augmented with a temporal smoothing loss, to produce the target video. The authors report experiments on two videos of one target individual, driven by two different source individuals, with only qualitative results shown in Fig. 4. The central claim is that the synthesized target sequences are photorealistic and consistent with the source in body motion, expression, and pose.","tokens_in":5321,"tokens_out":2410,"duration_ms":26433,"significance":"If the method worked as claimed, it would provide a practical way to combine body motion transfer with face reenactment in a single pipeline, using FAUP as an identity-independent intermediate representation to avoid explicit facial landmark normalization. The high-level idea is clearly stated and the pipeline is easy to follow, and the authors are explicit about their use of existing building blocks such as pix2pixHD and OpenFace. However, the significance cannot currently be assessed because the paper provides no quantitative metrics, no comparisons with prior work, no ablations, no failure analysis, no dataset details, and no code. The only evidence of effectiveness is two unquantified example videos. The contribution is therefore unverified as presented.","major_comments":[{"comment":"The paper's central claim of effectiveness rests entirely on two qualitative examples. Section 6 reports only that two videos of one individual were synthesized and that the results 'are photorealistic and consistent with the source sequence,' with Fig. 4 as the sole evidence. There are no quantitative metrics (e.g., FID, LPIPS, user study, keypoint accuracy, landmark distance), no baselines, no ablations, and no failure cases. Because the entire contribution depends on the demonstrated quality of the synthesized videos, this lack of evaluation is load-bearing and prevents verification of the central claim.","section":"Section 6"},{"comment":"The FAUP-to-landmark generator is trained only on target-person (FAUP, facial landmark) pairs, yet at test time it is fed source-person FAUP values ('During the transferring phase, we directly use source's FAUP to synthesize target's facial landmark'). The paper provides no evidence that OpenFace FAUP values are sufficiently identity-invariant or that source FAUP values lie within the target's training distribution. If source FAUP falls outside that distribution, the generated landmarks may not preserve the source expression, and stage 3's generator, trained on target UBKP-FL/video pairs, will receive out-of-distribution input. No domain adaptation, augmentation, or out-of-distribution analysis is provided, so the claimed expression consistency rests on an unvalidated distribution-shift assumption.","section":"Section 4"},{"comment":"The loss formulations contain notation inconsistencies that make the exact training objective ambiguous. Eq. (1) defines the per-scale adversarial loss LGAN(G,D_k), but Eq. (2) writes LGAN(G,D) without specifying the scale or how the three discriminators are combined. Eq. (3) contains a typo 'xx−L' in the first expectation and defines LtS(G,D) with a single discriminator D, while Eq. (4) sums LtS over k=1,2,3, presumably over D_k, but this is never made explicit. Since these losses define the training procedure, this ambiguity affects reproducibility.","section":"Equations (1)-(4)"}],"minor_comments":[{"comment":"The abstract states 'Experimental results demonstrated the effectiveness of our method' before any metrics are presented; consider tempering this claim or making the evaluation explicit.","section":"Abstract"},{"comment":"The description of the temporal setup would benefit from a precise definition of the notation: the text says the first output G(x_{t-L}) is conditioned on a zero image z, but the equations do not include z or explain how the zero image is used.","section":"Section 5"},{"comment":"The captions use inconsistent notation with the main text: Fig. 2 refers to x_t and y_t without defining the subscript, and Fig. 3 says 'The discriminator D attempts...' while Eq. (4) uses D_1, D_2, D_3.","section":"Fig. 2 and Fig. 3 captions"},{"comment":"There is a capitalization error in 'Different from above works, We propose...' where 'We' should be lowercase.","section":"Section 2"}],"recommendation":"reject","confidential_remarks":"The paper is well below the experimental standard expected for a serious journal: the central claim is unsupported by quantitative evaluation, no comparison or ablation is provided, and the cross-identity FAUP transfer assumption is unvalidated. The high-level idea may have merit, and a resubmission with thorough experiments, baselines, and analysis of out-of-distribution FAUP inputs would be a different and stronger paper."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You asked for a read on arXiv:1908.06607. Here's my take.\n\nThe paper is a short, clearly written proposal for upper-body video reenactment. The new bit is using facial action units plus pose (FAUP) as an identity-invariant intermediate: train a GAN to map a target person's own FAUP to their facial landmarks, then at test time feed the source's FAUP into that same generator. That's a sensible way around the identity-dependence of landmarks, and combining it with normalized body keypoints in a temporal GAN is a reasonable integration of existing components. I don't think the novelty is high, but it's a legitimate application-level idea.\n\nWhat it is not, is a demonstrated result. The experimental section is five sentences. Two videos of one target person, driven by two sources, shown as still frames. No quantitative metrics, no baselines, no ablations, no error analysis, no code or data. The paper's own claim of \"photorealistic and consistent\" is doing all the work, and it's unsupported. That alone is enough to reject in any venue that expects evidence.\n\nThe softer issues are the minor notation errors in Eqs. 2 and 3, which are annoying but not damaging. The more substantive flaw is the one the stress-test note hits: the FAUP-to-landmark generator is trained only on the target's own FAUP/landmark pairs, then used on source FAUP values with no domain adaptation or even a discussion of distribution shift. OpenFace AUs are not identity-free in practice; pose, face shape, and lighting all shift them. If the source's FAUP fall outside the target's training distribution, the generated landmarks won't carry the source's expression, and the whole pipeline fails quietly. The authors don't address this at all.\n\nIs the central idea circular or incoherent? No. The training procedure is supervised and the architecture is standard. The writing is readable and the pipeline is easy to follow. So this is not a crackpot paper; it's an under-built one.\n\nWho gets value from it? Somebody working on video reenactment might skim it for the FAUP trick and find it worth a shot. But as a published result, it should not stand on what's here.\n\nMy recommendation: desk-reject this version and suggest the authors resubmit with a real evaluation — multiple target subjects, standard metrics (FID, pose/expression consistency, user study), and at least one baseline comparison. If they do that, the paper could become worth reviewing. As-is, it is not.\n\nBottom line: plausible idea, unsupported claim. Not worth referee time in its current state.","headline":"A plausible three-stage pipeline for upper-body video transfer, but the evidence is a single paragraph with no metrics, baselines, or ablations; the central claim is not established.","tokens_in":5899,"tokens_out":1492,"would_cite":false,"duration_ms":17956,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims a three-stage GAN pipeline that transfers a source video's upper-body motion, facial expression, and pose onto a target person, using facial action units and poses as identity-free intermediate representations.","keywords":["video synthesis","generative adversarial network","upper body motion transfer","facial action units","facial landmark generation","temporal coherence","pose transfer","face reenactment"],"falsifier":"Run the full pipeline on a source video whose facial action unit values lie far outside the target person's training distribution, then measure whether the synthesized target facial landmarks reproduce the source expression (for example, by comparing landmark displacement under the same action units); if the landmarks wash out or drift, the claimed expression consistency fails.","tokens_in":4893,"feed_emoji":"🎬","tokens_out":6805,"duration_ms":65717,"temperature":0.7,"pith_summary":"This paper is trying to establish that a single pipeline can transfer an entire upper-body performance—body motion, pose, and facial expression—from one person in a video to another person, with the target person's own appearance and a face that stays realistic. The proposed route avoids directly translating source pixels into target pixels: it first reduces the source video to upper-body keypoints plus facial action units and head pose, converts those action units into the target person's facial landmarks, then synthesizes the target video from the keypoints and landmarks. If the claim holds, realistic reenactment no longer needs 3D motion capture or a separate face-transfer network stitched onto a body-transfer network, and identity-independent facial action units become a workable bridge between two people's faces.","feed_headline":"GAN pipeline transfers a source person's motion and face to a target","feed_subtitle":"Facial action units and head pose carry expression across people; temporal smoothing keeps frames coherent.","key_machinery":"The central machinery is the FAUP representation—facial action units plus head pose—used as an identity-free intermediate between source and target. A first pix2pixHD-based generator maps a source FAUP vector, placed in the center of an otherwise empty image, to the target person's facial landmarks; these are concatenated with the source's upper-body keypoints, linearly normalized to the target's body shape, and fed to a second pix2pixHD generator. Temporal coherence is obtained by conditioning each generated frame on the previous generated frames and optimizing a temporal smoothing adversarial loss together with feature-matching and VGG perceptual losses.","core_discovery":"The central claim is that the three-stage pipeline—source keypoint and action-unit extraction, target keypoint normalization with action-unit-to-landmark synthesis, and temporally smoothed image generation—produces a target-person video whose body motion, facial expression, and pose match the source sequence while the face remains photorealistic. The key move is replacing facial landmarks as the transfer medium with facial action units and pose (FAUP), which carry expression information without identity-specific spatial coordinates. This lets the same source FAUP values be mapped once into the target person's landmark space and then rendered by a GAN, rather than trying to normalize the source person's landmarks into the target's geometry directly.","pith_inferences":["An inference the paper does not draw: the FAUP-to-landmark mapping is only as reliable as the target person's training coverage, so source expressions outside that range may break the consistency claim; augmenting the target data with varied FAUP values would be the natural test.","Extending the same logic to full-body transfer is straightforward because the identity-free intermediate is not specific to upper-body keypoints; only the keypoint normalization would change.","A quantitative check the paper leaves for later: measuring landmark alignment or perceptual similarity between synthesized frames and target ground truth under matched expressions would place a bound on how far the photorealistic claim extends beyond the two example sequences."],"forward_implications":["Body motion and facial expression are transferred in the same pipeline, so there is no seam where separately generated face and body images meet.","Because FAUP values carry no spatial coordinates and are not tied to identity, the source's expression can be used directly without facial landmark normalization.","Sequential frame conditioning with the temporal smoothing loss is expected to reduce flicker compared with per-frame image translation.","The method needs only trained 2D pose and action-unit extractors plus paired target data, not 3D motion capture or a multi-person face model."],"supporting_citations":[{"why":"Supplies the conditional GAN backbone and the adversarial, feature-matching, and VGG losses used by both generation stages.","marker":"[17]"},{"why":"Supplies global pose normalization for upper-body keypoints and the sequential-frame temporal smoothing training setup.","marker":"[6]"},{"why":"Produces the facial action unit and pose values that serve as the identity-free expression representation.","marker":"[3]"},{"why":"Estimates the source upper-body keypoints that drive the body motion transfer.","marker":"[5]"}],"fun_headline_variants":["Facial action units carry expression across people in video synthesis","Source to target: GAN uses action units, not landmarks, for face transfer","Upper body video synthesis: action units map expressions across identities","Motion and expression transfer via action units, not direct landmarks","FAUP-based GAN synthesizes target person video with source's expressions"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The facial-action-unit-to-landmark generator is trained only on the target person's own facial action unit and landmark pairs, so it must successfully transfer to the source person's action unit values even when those values are outside the target's training range.","fun_headline_variants_meta":{"raw":{"variants":["Facial action units carry expression across people in video synthesis","Source to target: GAN uses action units, not landmarks, for face transfer","Upper body video synthesis: action units map expressions across identities","Motion and expression transfer via action units, not direct landmarks","FAUP-based GAN synthesizes target person video with source's expressions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000388,"raw_usage":{"total_tokens":1960,"prompt_tokens":774,"completion_tokens":1186,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":390,"completion_tokens_details":{"reasoning_tokens":1097}},"tokens_in":390,"tokens_out":1186,"duration_ms":8962,"temperature":1.0,"reasoning_tokens":1097,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:39:17.983696+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the full pipeline on a source video whose facial action unit values lie far outside the target person's training distribution, then measure whether the synthesized target facial landmarks reproduce the source expression (for example, by comparing landmark displacement under the same action units); if the landmarks wash out or drift, the claimed expression consistency fails.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Estimates the source upper-body keypoints that drive the body motion transfer."}],"review_version":1}