{"id":"509438bb-b0fe-443c-815a-34b3031db5f0","arxiv_id":"2502.03420","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Text-to-image portrait generators deviate from the requested age by about 3 to 5 years on average, with larger errors and demographic biases at older ages, so synthetic faces are not yet reliable for precise age tasks.","lead":"This paper tests whether three text-to-image models can generate portraits that look like the age specified in the prompt, using two automated age-estimation systems as judges. The result is a caution: synthetic faces often miss the requested age by several years, especially at older ages, so they should not be trusted for high-stakes age checks without careful filtering.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline MAE values may reflect age-estimator domain shift on synthetic faces rather than generator failure; no human-rater validation or real-data baseline is provided.","rationale":"The reader's weakest assumption correctly identifies the load-bearing point: the evaluation treats MiVOLO-D1 and JAM as valid measures of apparent age on synthetic faces. My independent read reaches the same conclusion, and the paper itself supplies supporting evidence that this is a real risk. In Section III.B the authors write that 'since the images in our study are synthetic and do not correspond to real individuals, we treat the prompted age as the ground truth,' but this only shifts the burden to the estimators. If either estimator is biased on synthetic faces, every MAE, RMSE, and correlation in Table IV mixes generator error with evaluator error. The paper also claims that synthetic performance 'lags behind real-world image data' and that biases are 'more pronounced in synthetic data,' yet no quantitative real-data baseline is reported. The proprietary, self-cited JAM model makes independent replication harder, but that is a reproducibility issue secondary to the validity issue. I would not change the reader's CONDITIONAL verdict: the paper is a useful benchmark, and its cautionary conclusion may well be true, but the central quantitative claim needs a human-rater calibration study and a real-data comparison before it can be accepted as measured. The proposed test would settle the matter directly.","tokens_in":7143,"tokens_out":3007,"duration_ms":30626,"concrete_test":"Select a stratified sample of ~300 synthetic images across the three generators, the 30 prompt ages, and genders. Have 5 independent human raters estimate each face's age from the image alone; average raters per image. Compute human MAE vs prompt age and compare with MiVOLO-D1 and JAM MAEs on the same images. If human MAE is significantly lower than estimator MAE (e.g., by more than 1 year), the paper's reported errors are partly evaluator artifacts and the central claim needs re-benchmarking with a calibrated evaluator; if human MAE is comparable to or higher than estimator MAE, the concern is refuted and the conclusion stands.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central inference, that T2I generators do not reliably depict age, is made by comparing prompt age against outputs of MiVOLO-D1 and JAM (Section III.B and IV.A). This treats the estimators as ground-truth readers of apparent age. But both estimators were trained predominantly on real photographs, and no evidence is given that their errors are unbiased on synthetic faces. Age estimators are known to regress toward the training-set mean on out-of-distribution inputs; if that occurs here, old prompts would be systematically underestimated and young prompts overestimated, inflating MAE and lowering correlation exactly where the paper reports failures. The problem is compounded by one estimator, JAM, being proprietary and self-cited, and by the authors' own statement that 'age estimation accuracy may decrease in older populations' (Ethical Impact C); this prior knowledge of estimator bias is not controlled for. Without calibrating estimators on synthetic images (e.g., with human perceptual ratings) or reporting how the same estimators score real portraits with known ages, the headline 2.99–5.26 year MAE range cannot be attributed to the generators rather than the evaluators.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates whether three text-to-image models (FLUX.1-dev, Stable Diffusion 3.5 Large, and SDXL Epic Realism) can generate portraits matching a specified age. The authors construct 12,960 prompts spanning 30 ages from 10 to 78, 212 nationalities, and both genders; generate one image per prompt; and score the resulting faces with two age estimators, MiVOLO-D1 and JAM. They report MAE, RMSE, and Pearson correlation for each estimator-generator pair, finding the best alignment for Epic Realism (MAE 2.99 with MiVOLO-D1, 3.52 with JAM) and the worst for Flux (MAE 5.26 and 4.68), with larger errors and low correlations at age extremes, demographic biases, and a small fraction of decade-level outliers. They conclude that current synthetic portraits are too unreliable for high-stakes age-related tasks without significant filtering and curation.","tokens_in":7231,"tokens_out":6133,"duration_ms":59260,"significance":"The study addresses a timely and practical question, and its breadth is a genuine strength: 30 ages, 212 nationalities, three generators, two estimators, and demographic subgroup analyses. The authors are appropriately cautious in their conclusion and provide useful diagnostics beyond global metrics, including age-bucket breakdowns, outlier counts, and regression slopes. However, the central claim depends on treating age-estimator outputs on synthetic faces as ground-truth apparent age; no human-rater validation, no calibration of the estimators on synthetic images, and no quantitative real-image baseline are provided. Because both estimators are trained predominantly on real photographs and one is proprietary and co-authored by the same team, the reported MAE range is best interpreted as a joint generator-estimator alignment score. If the estimator domain-shift concern is addressed with perceptual ratings or a real-image control, the paper would be a sound and useful cautionary benchmark for practitioners.","major_comments":[{"comment":"The paper's load-bearing assumption is stated in Section III.B: 'we treat the prompted age as the ground truth,' and then all MAE/RMSE/correlation values in Table IV are computed against the outputs of MiVOLO-D1 and JAM. This assumes both estimators are unbiased readers of apparent age on synthetic faces, but no evidence for that assumption is provided. Age estimators trained on real photographs are known to compress toward the training distribution on out-of-distribution inputs, which would produce exactly the pattern reported here: overestimated young prompts, underestimated old prompts, inflated MAE, and low correlation at age extremes. The authors themselves flag in Ethical Impact C that age estimation accuracy may decrease in older populations. The authors should add a synthetic-image calibration using human perceptual age ratings and/or a quantitative real-portrait baseline scored by the same estimators. Without such controls, the headline 2.99-5.26 year MAE range cannot be attributed to generator failure rather than evaluator domain shift.","section":"Section III.B and Section IV.A"},{"comment":"JAM is a proprietary age estimator developed by members of the author team (reference [12]), and the paper provides no architecture, training data description, or reported performance on synthetic images. Section IV.A uses JAM as one of two judges for every generator ranking, yet an independent reader cannot reproduce or scrutinize its behavior. Section IV.B claims that JAM 'broadly tracks age variations' on real-life production data but gives no quantitative results for that claim. The authors should either disclose sufficient technical detail about JAM, base the main claims on publicly available estimators with published error characteristics, or add a table comparing JAM's real-image MAE/correlation by age bucket so that estimator bias can be separated from generator error. This is a transparency and reproducibility issue, not a novelty issue.","section":"Section III.B and reference [12]"},{"comment":"Table IV reports large differences between generator-estimator combinations (e.g., Epic Realism MAE 2.99 versus Flux MAE 5.26 for MiVOLO-D1), but no confidence intervals, standard errors, or sample-size information are given for these metrics. The t-tests mentioned in Section IV.D are not reported with test statistics, p-values, or multiple-comparison corrections. Since the relative ranking of generators is a central conclusion, the authors should provide bootstrap confidence intervals for MAE/RMSE/correlation and a table of the t-test results. In addition, the age-bucket correlations in Section IV.B are computed within narrow 10-year ranges, which mechanically restricts the achievable correlation; the interpretation that generators 'struggle to differentiate subtle changes' should be checked against the estimator's own consistency on real faces over the same age spans.","section":"Section IV.A (Table IV) and Section IV.D"},{"comment":"The abstract and conclusion claim that text-to-image models 'can consistently generate faces reflecting different identities' and preserve 'identity-related cues,' but no identity-consistency metric appears in the methodology. The experiments vary age, nationality, and gender in prompts, yet there is no face-recognition embedding similarity measure, no repeated-generation test, and no controlled comparison linking the same identity across conditions. This positive claim is presented alongside the main negative result, so it should either be removed or supported with a concrete measurement; as written, it exceeds the evidence in the paper.","section":"Abstract and Section V"}],"minor_comments":[{"comment":"The text says '30 distinct ages from 10 to 78,' but 10 through 78 inclusive contains 69 integer ages; the authors should list the specific 30 ages or explain the sampling rule to remove ambiguity.","section":"Section III.A"},{"comment":"The age-bucket narrative discusses 10-19, 20-29, 30-39, 40-49, 50-59, and 70-79 but skips 60-69; the authors should clarify whether the 60-69 bucket was omitted, merged, or simply not highlighted.","section":"Section IV.B"},{"comment":"The outlier definition uses a deviation threshold of 'more than 11 years' without justification; a brief rationale or a sensitivity analysis for thresholds such as 8, 10, and 12 years would strengthen the robustness of the outlier discussion.","section":"Section IV.C"},{"comment":"The sentence beginning 'One aspect to assess from a risk perspective is potential bias...' is duplicated verbatim in the same section; the duplicate should be removed.","section":"Ethical Impact Statement, Section A"},{"comment":"The scatter plots would be easier to interpret if each panel included the MAE, sample size, and regression equation, and if the axis labels and point density (e.g., alpha or hex-bin) were specified.","section":"Figure 2"},{"comment":"No code or dataset is provided, and the dataset is only 'aimed' to be released in the future; given the paper's benchmarking character, releasing prompts, generation parameters, and per-image age estimates would materially improve reproducibility.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the journal's scope and the core question is worth publishing if the validator-validation issue is fixed. The proprietary, self-developed JAM estimator deserves particular editorial attention: the manuscript should clearly disclose its relation to the authors and provide enough information for an independent assessment. The manuscript would benefit from stricter statistical reporting (confidence intervals, test details) and from separating the central age-accuracy claim from the unsupported identity-consistency claim."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely new benchmark—12,960 prompts, 212 nationalities, 30 ages, three generators, two estimators—and the scale alone makes it useful to anyone working on synthetic data for biometrics. The authors also deserve credit for reporting demographic breakdowns and outlier rates rather than just a single MAE. The paper is readable and the cautionary direction is plausible.\n\nBut the load-bearing inference has a hole. The paper treats the age estimators as ground truth for apparent age and attributes all error to the generators. No one validates MiVOLO-D1 or JAM on synthetic faces. No human-rater study. No real-data baseline. The authors even state in Ethical Impact C that JAM's accuracy decreases in older populations, which is exactly the kind of estimator bias that could inflate MAE at the age extremes where they report the worst results. Without calibration, the headline 2.99–5.26 year MAE range is uninterpretable: it could be generator failure, estimator domain shift, or both. The stress-test note is right.\n\nThere are smaller problems too. JAM is proprietary and self-cited, so the second judge is a black box. No code, prompts, seeds, or dataset are released, and there are no confidence intervals on any metric. The paper also claims identity consistency is preserved, but I did not see an actual identity metric. The t-tests are mentioned but not shown.\n\nWhat holds up: the prompt-age-vs-predicted-age regressions, the age-bucket breakdowns, and the outlier analysis are internally consistent. The qualitative finding that different generators have different age-bias profiles is believable and worth pursuing. The central caution about synthetic data for high-stakes age verification is reasonable, but it needs to be reframed as a joint property of generators and estimators.\n\nBottom line: this paper deserves a serious referee, not a desk reject. The benchmark is too large and too relevant to ignore. But the referee should ask for estimator validation on synthetic images (human raters or a real-image control), confidence intervals, a released dataset or at least code and prompt list, and a reanalysis that separates generator error from estimator bias. I would not cite the headline numbers until that is done.","headline":"A genuinely large synthetic-portrait benchmark whose headline MAE numbers are undercut by unvalidated age estimators, so the paper's central attribution of error to generators does not yet hold.","tokens_in":7863,"tokens_out":1422,"would_cite":false,"duration_ms":14877,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that text-to-image models can generate plausible faces with varied identities, but their depictions of a specified age are off by three to five years on average, so synthetic portraits are not yet reliable for high-stakes…","keywords":["text-to-image generation","age estimation","synthetic portraits","demographic bias","generative models","MiVOLO-D1","JAM","portrait generation"],"falsifier":"Take a subset of the generated portraits to human raters and ask them to estimate each face's age: if human estimates match the prompt ages closely while MiVOLO-D1 and JAM still show MAEs of 3 to 5 years, the paper's conclusion that generators fail at age depiction would be weakened, because the estimators themselves would be the untested link.","tokens_in":6858,"feed_emoji":"🧑🦳","tokens_out":7655,"duration_ms":65801,"temperature":0.7,"pith_summary":"This paper tests whether text-to-image models can depict a specified age by generating 12,960 synthetic portraits from prompts that fix age, nationality, and gender, then feeding those portraits to two age-estimation models. The central finding is that the generators are better at producing plausible, identity-varying faces than at hitting the requested age: average errors range from about 3 to 5.3 years depending on the model and estimator, and the largest, most erratic errors occur for prompts over 60. The authors conclude that current synthetic portraits are not reliable enough for high-stakes age-dependent tasks such as identity verification or training age-estimation models, unless they are heavily filtered and curated. They argue that such data may still be useful in exploratory or low-precision applications where exact age is not critical.","feed_headline":"AI portraits miss prompted ages by up to 5 years","feed_subtitle":"A 12,960-prompt test of three generators shows synthetic faces are not dependable for age-critical tasks.","key_machinery":"The measurement pipeline is the prompt-to-estimator loop. A fixed prompt template ('Photorealistic selfie photo of a [age]-year-old [nationality] [gender] person, centered, high-resolution') is used to generate 12,960 portraits with three text-to-image models; two age estimators, MiVOLO-D1 and JAM, then produce a predicted age for each face. The argument is carried by the comparison between those predicted ages and the prompted age, summarized as MAE, RMSE, Pearson correlation, per-age-bucket and per-demographic breakdowns, and regression diagnostics. The load-bearing assumption is that the prompted age is the ground truth, so every reported error measures how far the generated face is from the age it was asked to show.","core_discovery":"On its own terms, the paper shows that text-to-image generators can produce faces with varied identities, but the age those faces appear to be is only loosely controlled by the prompt. Across 12,960 prompts spanning 30 ages from 10 to 78, 212 nationalities, and both genders, the two age estimators give mean absolute errors between 2.99 and 5.26 years depending on the generator and estimator, with Epic Realism the most accurate and FLUX.1-dev the least. Older prompts (60+) show the largest errors and the most erratic correlations, and both estimators reveal gender and regional offsets in synthetic faces that the authors say are larger than what the same estimator shows on real faces. The conclusion the authors draw is that age fidelity, not identity fidelity, is the weak point of current text-to-image models for biometric use.","pith_inferences":["A direct extension the authors leave open is a human-rater study on a sample of the same portraits; it would separate generator error from estimator miscalibration on synthetic faces.","The same prompt protocol could be rerun on newer text-to-image models to see whether age fidelity improves or whether the reported gender and regional offsets persist.","If the pattern generalizes, prompt-age error could serve as a cheap, image-only proxy for demographic fairness auditing of generative models, since systematic MAE offsets by gender and region are measurable without real-image data.","Adding known-age real portraits through the same estimators would provide a baseline the paper does not include, clarifying how much of the reported error is specific to synthetic faces."],"forward_implications":["Age-estimation practitioners should not use text-to-image portraits as ground-truth training data without per-image filtering or calibration.","Teams choosing a generator for age-sensitive work should prefer Epic Realism over FLUX.1-dev based on these error ranges, and should treat 60+ prompts as unreliable.","Age-verification systems that use synthetic data should only act on large age mismatches, since small-to-moderate errors are common.","The reported pattern motivates post-hoc correction or targeted prompt engineering for the older-age range, where MAE and correlation degrade most.","Synthetic portraits remain usable for exploratory tasks where approximate age distributions are enough."],"supporting_citations":[{"why":"FLUX.1-dev generator under test; its higher MAE is key evidence of variability.","marker":"[8]"},{"why":"Stable Diffusion 3.5 Large generator under test; produces mid-range errors.","marker":"[9]"},{"why":"SDXL Epic Realism generator under test; the best-performing generator in the comparisons.","marker":"[10]"},{"why":"MiVOLO-D1 age estimator; one of the two ground-truth measures of apparent age.","marker":"[11]"},{"why":"JAM age estimator; the other ground-truth measure, also used for demographic breakdowns.","marker":"[12]"},{"why":"Public evaluation of age estimators; cited to support using JAM as a reliable baseline on real data.","marker":"[13]"}],"fun_headline_variants":["AI portraits miss prompted age by up to 5 years","Age fidelity weak in text-to-image portraits: MAE up to 5 years","Synthetic faces misrepresent age by 3-5 years on average","Text-to-image models fumble age: errors up to 5 years","Prompted age vs. apparent age: synthetic portraits off by years"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire comparison treats the prompt age as ground truth and the two estimators' outputs as accurate measurements of the apparent age of synthetic faces, even though neither estimator was calibrated or validated on synthetic images.","fun_headline_variants_meta":{"raw":{"variants":["AI portraits miss prompted age by up to 5 years","Age fidelity weak in text-to-image portraits: MAE up to 5 years","Synthetic faces misrepresent age by 3-5 years on average","Text-to-image models fumble age: errors up to 5 years","Prompted age vs. apparent age: synthetic portraits off by years"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000522,"raw_usage":{"total_tokens":2510,"prompt_tokens":917,"completion_tokens":1593,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":533,"completion_tokens_details":{"reasoning_tokens":1499}},"tokens_in":533,"tokens_out":1593,"duration_ms":9525,"temperature":1.0,"reasoning_tokens":1499,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-09T04:48:04.972680+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a subset of the generated portraits to human raters and ask them to estimate each face's age: if human estimates match the prompt ages closely while MiVOLO-D1 and JAM still show MAEs of 3 to 5 years, the paper's conclusion that generators fail at age depiction would be weakened, because the estimators themselves would be the untested link.","supporting_citations":[{"cited_title":"Flux.1: Redefining text-to-image ai with superior visual fidelity","cited_arxiv_id":null,"evidence_quote":"FLUX.1-dev generator under test; its higher MAE is key evidence of variability."},{"cited_title":"Stable diffusion 3.5 large - the most advanced sd model","cited_arxiv_id":null,"evidence_quote":"Stable Diffusion 3.5 Large generator under test; produces mid-range errors."},{"cited_title":"Realism engine sdxl - v3.0 vae","cited_arxiv_id":null,"evidence_quote":"SDXL Epic Realism generator under test; the best-performing generator in the comparisons."},{"cited_title":"Mivolo: Multi-input trans- former for age and gender estimation, 2023","cited_arxiv_id":null,"evidence_quote":"MiVOLO-D1 age estimator; one of the two ground-truth measures of apparent age."},{"cited_title":"Novikov, Ruslan Parkhomenko, Artem V oronin, and Alix Melchy","cited_arxiv_id":null,"evidence_quote":"JAM age estimator; the other ground-truth measure, also used for demographic breakdowns."},{"cited_title":"Age estimation evaluation report, 2024","cited_arxiv_id":null,"evidence_quote":"Public evaluation of age estimators; cited to support using JAM as a reliable baseline on real data."}],"review_version":1}