{"id":"282477a2-525f-4039-ae4e-bfb59fc93992","arxiv_id":"2507.14549","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Faces generated on ANN decision boundaries raise inter-individual variability in emotion labeling, and fine-tuning on those labels improves both group-level and individual-level prediction.","lead":"Faces synthesized on an artificial neural network's decision boundary, where the network cannot choose between two emotions, also split human viewers into different emotion judgments. The authors built a 1,678-image dataset from 22,450 online judgments and used it to fine-tune emotion-recognition models toward individual-level human perception.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No unaddressed artifact confound for the core hypothesis: human-ANN shared boundary may be circular because human responses were used to filter both stimuli and fine-tuned models.","rationale":"The reader identified the most consequential issue: the absence of a naturalness check and the lack of a baseline comparison means the measured human variability could be an artifact of generated-image quality. My independent concern sharpens this: the generation pipeline's two-stage diffusion and the 75th-percentile filtering (Eq. 3) both use the same ANN that defines the 'boundary', and then the fine-tuning uses human responses on the same images. This creates a self-referential loop where the classifier that was used to generate and filter images is later fine-tuned on human labels of those images, making the human-ANN alignment result partly tautological. The strongest independent check would be a neutral-guidance control: generate images with the same diffusion and filtering infrastructure but without boundary uncertainty guidance, then compare human disagreement. If disagreement is equal, the core claim that ANN decision boundaries specifically predict human perceptual variability fails. If disagreement is lower, the claim is supported. This is not a fatal flaw; the paper is a reasonable conditional contribution, but the load-bearing assumption of shared emotional ambiguity versus generic generated-face ambiguity is untested.","tokens_in":9456,"tokens_out":1583,"duration_ms":18726,"concrete_test":"Run a control experiment that generates a matched set of images with identical diffusion and filtering steps but with the uncertainty loss replaced by a neutral guidance (e.g., guidance toward a single emotion or toward the RAF-DB mean embedding, matched for the same emotion pairs). Present these control images to the same or an equivalent human participant pool under identical conditions. If the control images show the same magnitude of inter-observer disagreement and the same model-human entropy correlation as the boundary images, the central claim that ANN decision boundaries specifically track human perceptual variability would be weakened. Additionally, compute the human-model entropy correlation both on the filtered images and on natural RAF-DB images; if the correlation is comparable on natural images, the alignment result is not specific to boundary sampling.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that ANN-confusing facial expressions also elicit human perceptual variability, and that fine-tuning on behavioral data aligns ANNs with human perceptual patterns. The cleanest load-bearing assumption is that the varEmotion human responses are a gold standard measured on fixed stimuli, and that the ANN/human correlation reported is not an artifact of selection or training. However, the pipeline selects generated images by conditioning the diffusion on target emotion pairs and then filters them via a 75th-percentile ANN activation criterion (Eq. 3). Human annotation entropy is then measured on exactly those images and used (a) to report 'heightened uncertainty' relative to an implicit baseline, and (b) to fine-tune the same classifiers. If the generated images are unnatural or idiosyncratic in a way that causes all models and humans to be uncertain, the shared-boundary conclusion would be overstated. The paper lacks a comparison against boundary images sampled without uncertainty guidance or against matched natural faces with known ambiguity (e.g., morphed natural expressions). It also lacks a check that the human-model entropy correlation is not driven by trivial factors such as image brightness, low-level artifacts, or atypical face morphology. Importantly, the 75th-percentile filter uses the same ANN whose boundaries are being probed, so the generation/filtering loop could select images that are ambiguous for that ANN for reasons unrelated to shared human emotion boundaries. The strongest concern is that the reported human-ANN alignment and 'heightened uncertainty' result has no baseline or control stimulus set, so it cannot distinguish shared emotional ambiguity from a generic artifact-of-generation effect.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a pipeline for synthesizing facial expression images on the decision boundaries of ANN classifiers, using a two-stage diffusion process with an 'uncertainty guidance' loss and a 75th-percentile activation filter. The resulting 1,678 images were shown to 66 human participants in 22,450 trials, forming the varEmotion dataset. The authors report that these images elicit high perceptual variability in humans (entropy distributions and an approximately 80% success-plus-bias rate), and that fine-tuning CLIP, DAN, and ResEmoNet on group-level and individual-level behavioral data improves their accuracy on varEmotion and varEmotion-i and increases the Spearman correlation between model and human entropy. The paper concludes that ANN decision boundaries are a systematic source of ambiguous stimuli for humans and that behavioral fine-tuning can align models with individual-level perceptual patterns.","tokens_in":9759,"tokens_out":4779,"duration_ms":55954,"significance":"If the central claim were fully supported, the paper would make a useful contribution: it would connect ANN decision boundaries to human perceptual variability, provide a new behavioral dataset (varEmotion), and show that human choice data can personalize emotion classifiers. The dataset collection effort is substantial, and the cross-architecture comparison (CLIP, DAN, ResEmoNet) gives the fine-tuning results some breadth. However, the current evidence is suggestive rather than conclusive: the 'heightened uncertainty' claim lacks a baseline, the fine-tuning evaluations appear to be on training data, and the generation/filtering loop uses the same ANN whose shared boundary with humans is being asserted. These issues are fixable within the manuscript's scope, but they are load-bearing for the abstract and conclusion.","major_comments":[{"comment":"The claim that ANN-boundary images provoke 'heightened' perceptual uncertainty is not supported without a baseline: the entropy distribution is reported only for generated boundary images, and the roughly 80% success-plus-bias rate in Fig. 4(b) counts the 'bias' outcome (all subjects choosing one target) as a positive result even though that outcome reflects agreement rather than variability. I recommend comparing the entropy of these images with natural faces, off-boundary generated faces, or morph continua, and reporting success and bias rates separately.","section":"Sec. IV-A, Fig. 7(a)"},{"comment":"The fine-tuning evaluation does not separate training from test images: GroupNet and IndivNet are fine-tuned on varEmotion and varEmotion-i and then evaluated on those same datasets, so the accuracy improvements are expected from fitting the training labels. Please report performance on held-out images using image-level or subject-level cross-validation, and include BaseNet on the same held-out splits.","section":"Sec. V-A, Fig. 5(a)"},{"comment":"The entropy correlation improvement from ρ = 0.26 to ρ = 0.85 is computed on varEmotion images that were used for group fine-tuning; a model trained to reproduce human choices on those images would be expected to have correlated entropy. Please compute the Spearman correlation on held-out images and report confidence intervals or p-values.","section":"Sec. V-B, Fig. 5(c)"},{"comment":"The filtering criterion uses the same ANN whose decision boundaries are being probed, so the selected images are ambiguous for that specific classifier by construction; this does not by itself establish that the ambiguity is shared with humans. A control with a different classifier architecture, or a comparison of filtered versus unfiltered generated images, is needed to rule out selection artifacts.","section":"Sec. III-C, Eq. (3)"},{"comment":"The manuscript asserts that the generated images retain 'photorealistic authenticity' but provides no human naturalness rating or artifact check. If the boundary images are uncanny or low-quality, the elevated human disagreement could reflect image unnaturalness rather than shared emotional ambiguity. A brief naturalness rating experiment or a comparison with real or morphed faces would resolve this.","section":"Sec. III-B and Introduction"},{"comment":"The uncertainty loss is underspecified: q(y) is never defined, and the notation switches between p(y|x) and p(y). If q(y) is a distribution over the two target emotions, the product form -p(y|x) q(y) is not a standard objective and its optimization behavior is unclear; please define the exact target distribution (e.g., q = 0.5 for each target emotion) and report sensitivity to the guidance strength γ.","section":"Sec. III-B, Eq. (1)"}],"minor_comments":[{"comment":"The caption refers to a 'Digit recognition task,' but the experiment is a facial expression recognition task; please correct this.","section":"Fig. 7 caption"},{"comment":"Entropy is estimated from roughly 13 judgments per image on average across six categories, which can produce biased entropy estimates; please report bias-corrected entropy or bootstrap confidence intervals.","section":"Sec. IV"},{"comment":"The section headings appear mismatched: Sec. III-A is titled 'Generating Images on ANN perceptual boundary' but contains no method details, while Sec. III-B is titled 'Facial expression recognition experiment' but describes the generation procedure; please reorganize the subsection structure.","section":"Sec. III-A and III-B"},{"comment":"Reference [8] is cited as evidence of 'remarkable accuracy' in facial expression recognition, but the listed work is about family interaction and appears unrelated; please verify and replace this citation.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's central claims are undermined by missing baseline comparisons and by evaluating fine-tuned models on the same data used for training, both of which are fixable with additional experiments or re-analysis. The citation error in [8] should also be corrected. I would be supportive after revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I read it. The one thing to know: the generation pipeline is the real contribution. Sampling CLIP embeddings along an ANN decision boundary via diffusion guidance, then filtering by 75th-percentile activation, is a tidy way to get images that the classifier can't pin down. The varEmotion dataset—1,678 images, 22,450 trials across 66 participants—is real work and potentially a useful resource for the affective-computing crowd, provided they release it.\n\nThe soft spot is the central claim. 'Heightened perceptual uncertainty' is asserted without any baseline. There is no comparison to natural faces, off-boundary generated faces, or known ambiguous morphs. Entropy > 0 just says the images aren't unanimous; it doesn't say ANN boundaries generate ambiguity better than a random face would. The 80% success+bias rate is about hitting the guidance targets, not about boundary-vs-off-boundary.\n\nI also worry about the alignment numbers. GroupNet and IndivNet are fine-tuned on varEmotion and then evaluated on varEmotion and varEmotion-i. The entropy correlation jump from 0.26 to 0.85 is on the training data. That's fitting, not prediction. They need a held-out set of stimuli or participants to make the alignment claim stick.\n\nOn circularity: I don't think the human responses are circular—they're measured on fixed images. But the filter uses the same ANN whose boundary is being probed, so the selection may favor images that are ambiguous for that particular model for reasons that have nothing to do with shared human boundaries. A cross-family ANN check or a natural-ambiguous-face control would help.\n\nMinor: Figure 7's caption says 'Digit recognition task.' Looks like a copy-paste remnant. And no code or data are released, which is a practical obstacle to verifying image quality.\n\nBottom line: it's a solid method plus a dataset with an overstated interpretation. It deserves referee time, but the referee should push for baseline comparisons and out-of-sample evaluation. If they add those, the shared-boundary story could be convincing. As it stands, treat the 'heightened uncertainty' claim as not yet demonstrated.","headline":"A genuinely useful generation pipeline and a substantial behavioral dataset, but the central claim of heightened human uncertainty is not yet demonstrated without baseline comparisons and out-of-sample evaluation.","tokens_in":10266,"tokens_out":2716,"would_cite":false,"duration_ms":33344,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Images that an ANN finds ambiguous between two expressions are the faces on which humans disagree, and fine-tuning on human choices makes the network predict individual perception.","keywords":["facial expression recognition","perceptual variability","ANN decision boundary","diffusion model","human-AI alignment","emotion perception","varEmotion dataset"],"falsifier":"Ask a separate group of raters to score naturalness of boundary-generated and non-boundary generated faces; if boundary images are rated less natural and naturalness predicts disagreement entropy, the central claim is an artifact. Alternatively, match boundary and non-boundary faces for naturalness and compare human disagreement entropy; no difference would falsify the claim that ANN boundaries map to human perceptual variability.","tokens_in":9291,"feed_emoji":"🎭","tokens_out":9493,"duration_ms":101408,"temperature":0.7,"pith_summary":"Facial expressions that sit on the decision boundary of an ANN classifier, where the network cannot commit to one emotion, are also the expressions on which human observers disagree most. To test this, the paper builds a perceptual boundary sampling method that generates faces via a two-stage diffusion process guided by an uncertainty loss, then filters the outputs through classifier activation thresholds. The resulting varEmotion dataset, with 1,678 images and 22,450 judgments from 66 participants, shows that these ANN-confusing faces provoke high entropy in human emotion choices, with the guiding emotion pair attracting roughly 80% of images into success or bias categories. Fine-tuning three classifiers on human behavioral data improves both group-level and individual-level prediction of human choices, and raises the correlation between model and human judgment entropy substantially. If correct, ANN decision boundaries are a practical source of emotionally ambiguous stimuli, and behavioral fine-tuning is a route to personalized emotion recognition.","feed_headline":"Faces that stump AI also split human emotion judgments","feed_subtitle":"ANN-ambiguous faces form the varEmotion dataset; fine-tuning on human choices aligns models with individual perception.","key_machinery":"The load-bearing mechanism is uncertainty guidance during diffusion sampling. In a first stage, a diffusion model denoises image embeddings while being steered by the loss $loss(x,y) = -p(y|x)q(y)$, which raises the classifier's probability for two target emotions and lowers it for the rest; the sampling step is $x_{t-1} = \\mathrm{DDPM}^{-}(x_t) - \\gamma\\nabla_{x_t} loss(x_t,y)$ with $\\gamma = 0.5$. The resulting embeddings are rendered into images by a second-stage text-to-image model, and candidates are kept only when both target emotions' activations exceed their 75th percentiles on the RAF-DB dataset. Human choices on the surviving images form the varEmotion dataset. Then an MLP head on each of three network architectures is fine-tuned on mixed group and individual behavioral data, which is the step that transfers human variability back into the model.","core_discovery":"The central discovery is an empirical correspondence: stimuli deliberately placed on ANN classification boundaries transfer their ambiguity to human observers. The paper claims that images whose ANN activations are uncertain between two emotions, such as anger versus fear, yield human choice distributions with high entropy; the guiding emotion pair still dominates the choices for roughly 80% of images, so the disagreement is structured rather than random. It further claims that fine-tuning a network on human trial data moves its predictions toward both the group distribution and each individual's own patterns, with individual-level fine-tuning adding about 1 to 3.5% accuracy over group-level fine-tuning and raising the Spearman correlation between model and human judgment entropy from 0.26 to 0.85 for one architecture. On the paper's terms, ANN decision boundaries and human perceptual boundaries are aligned closely enough that one can be used to find the other.","pith_inferences":["Going beyond the paper: if the boundary-to-variability link is causal rather than correlational, the same two-stage diffusion pipeline should generate ambiguous stimuli in any domain where a classifier's softmax boundary can be defined, not just facial expressions.","Going beyond the paper: because the filtering step uses only ANN activations, the decisive control is to compare boundary-generated faces against non-boundary generated faces matched for human-rated naturalness; if disagreement entropy no longer differs, the shared-boundary claim would reduce to an image-quality effect.","Going beyond the paper: since individual fine-tuning worked with only a few hundred trials per participant, the protocol could become a lightweight calibration step for personalized affective computing rather than requiring large per-person datasets."],"forward_implications":["Boundary sampling can produce on-demand stimuli for any emotion pair that confuse both the ANN and human observers, with nearly 80% of generated images falling into the paper's success or bias categories.","Fine-tuning on human behavioral data does not degrade standard benchmark accuracy on RAF-DB while improving prediction on the ambiguous varEmotion images.","Individual-level fine-tuning with a relatively small number of trials per person outperforms group-level fine-tuning on predicting that person's choices.","Architecture matters: the network with the largest group-level gain improves by 35% on varEmotion, while the smallest gain is 5%, showing that some models can fit human perceptual boundaries better than others.","Model uncertainty becomes human-like after fine-tuning: for one architecture the Spearman correlation between model and human choice entropy rises from 0.26 to 0.85, meaning the model reproduces how unsure people are, not just what they choose."],"supporting_citations":[{"why":"supplies the two-stage generation design (embeddings first, then images) that the paper adapts.","marker":"[9]"},{"why":"provides the first-stage embedding diffusion prior and controllable generation framework used as the base of the sampling pipeline.","marker":"[1]"},{"why":"shows adversarial image manipulations influence both human and machine perception, motivating the shared-boundary hypothesis.","marker":"[10]"},{"why":"demonstrates that robustified ANN perturbations can modulate human percepts, evidence for shared sensitivity between ANNs and humans.","marker":"[11]"},{"why":"introduces model metamers, the conceptual frame for stimuli that are equivalent for ANNs but not for humans.","marker":"[12]"},{"why":"introduces controversial stimuli that elicit divergent judgments across models, the basis for pitting networks against humans.","marker":"[13]"},{"why":"provides the vision-transformer diffusion backbone used in the first stage of generation.","marker":"[38]"},{"why":"motivates the use of diffusion models as natural-image regularizers to avoid unnatural synthetic stimuli.","marker":"[21]"},{"why":"supports the use of diffusion-based natural adversarial examples to keep generated images realistic.","marker":"[22]"}],"fun_headline_variants":["AI's ambiguous faces mirror human emotion variability","Boundary faces: AI uncertainty mirrors human judgment splits","AI's unclear faces predict human emotion disagreements","When AI hesitates on faces, humans split too","AI boundary faces reveal human perceptual variability"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the generated boundary images look like real, natural faces, so the human disagreement they trigger comes from emotional ambiguity rather than from the images being artificial or uncanny, but the pipeline's only filter is an ANN activation threshold with no human naturalness check.","fun_headline_variants_meta":{"raw":{"variants":["AI's ambiguous faces mirror human emotion variability","Boundary faces: AI uncertainty mirrors human judgment splits","AI's unclear faces predict human emotion disagreements","When AI hesitates on faces, humans split too","AI boundary faces reveal human perceptual variability"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000834,"raw_usage":{"total_tokens":3635,"prompt_tokens":933,"completion_tokens":2702,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":549,"completion_tokens_details":{"reasoning_tokens":2633}},"tokens_in":549,"tokens_out":2702,"duration_ms":19879,"temperature":1.0,"reasoning_tokens":2633,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T15:53:07.875758+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Ask a separate group of raters to score naturalness of boundary-generated and non-boundary generated faces; if boundary images are rated less natural and naturalness predicts disagreement entropy, the central claim is an artifact. Alternatively, match boundary and non-boundary faces for naturalness and compare human disagreement entropy; no difference would falsify the claim that ANN boundaries map to human perceptual variability.","supporting_citations":[{"cited_title":"CoCoG: Controllable Visual Stimuli Generation based on Human Concept Representations","cited_arxiv_id":"2404.16482","evidence_quote":"provides the first-stage embedding diffusion prior and controllable generation framework used as the base of the sampling pipeline."},{"cited_title":"Subtle adversarial image manipulations influence both human and machine perception,","cited_arxiv_id":null,"evidence_quote":"shows adversarial image manipulations influence both human and machine perception, motivating the shared-boundary hypothesis."},{"cited_title":"Strong and precise modulation of human percepts via robustified anns,","cited_arxiv_id":null,"evidence_quote":"demonstrates that robustified ANN perturbations can modulate human percepts, evidence for shared sensitivity between ANNs and humans."},{"cited_title":"Model metamers reveal divergent invariances between biological and artificial neural networks,","cited_arxiv_id":null,"evidence_quote":"introduces model metamers, the conceptual frame for stimuli that are equivalent for ANNs but not for humans."},{"cited_title":"Controversial stimuli: Pitting neural networks against each other as models of human cognition,","cited_arxiv_id":null,"evidence_quote":"introduces controversial stimuli that elicit divergent judgments across models, the basis for pitting networks against humans."},{"cited_title":"Adversarial counterfactual visual explanations,","cited_arxiv_id":null,"evidence_quote":"motivates the use of diffusion models as natural-image regularizers to avoid unnatural synthetic stimuli."},{"cited_title":"Advdiffuser: Natural adversarial example synthesis with diffusion models,","cited_arxiv_id":null,"evidence_quote":"supports the use of diffusion-based natural adversarial examples to keep generated images realistic."}],"review_version":1}