{"id":"c3b8a0ea-064e-43a1-a792-a278be22c4f9","arxiv_id":"2508.17675","paper_version":3,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper reports that GPT-4o with advanced prompts can create synthetic normative responses for the Cookie Theft picture description task that distinguish diagnostic groups and demographic variations.","lead":"The abstract describes using GPT-4o models to generate synthetic text responses for standard cognitive test images, aiming to replace costly human-collected normative data. If the approach works, researchers could build new image-based cognitive tests without the traditional expense of gathering reference data.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The submitted full text is a different paper (adversarial advertisement embedding attacks, arXiv:2508.17674v2), so the abstract's claim about synthetic normative data has no supporting methods, data, or results in the manuscript.","rationale":"The reader's stated verdict, UNVERDICTED, is correct: the full text is a different paper, so the abstract's claims cannot be evaluated against the submitted manuscript. The reader's rationale identifies this mismatch directly. However, the reader's formal 'weakest_assumption' field points to the representativeness of embedding-based similarity measures—a substantive scientific concern that would matter if the study were actually present. In the submitted text, that concern is secondary; the first-order problem is that no portion of the full text describes the cognitive-assessment study. Thus I agree with the reader's overall conclusion but only partially with their stated weakest assumption. The recommended verdict remains UNVERDICTED rather than REJECT because the available evidence is insufficient to determine whether the abstract's claim is true or false; withholding judgment is the appropriate scholarly response to a document whose body is unrelated to its abstract. A simple metadata and keyword check would settle the document mismatch quickly and definitively.","tokens_in":8943,"tokens_out":3217,"duration_ms":32838,"concrete_test":"Retrieve the arXiv metadata for ID 2508.17675 and compare its title and abstract to the submitted full text. Then run a keyword search over the full text for 'Cookie Theft', 'GPT-4o', 'GPT-4o-mini', 'normative', 'BLEU', 'ROUGE', 'BERTScore', 'embedding', and 'diagnostic'. If the metadata for 2508.17675 matches the submitted abstract but the full text contains zero matches for these terms (or matches arXiv:2508.17674v2), the concern is confirmed: the manuscript body does not support the abstract's central claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract's central claim—that generative multimodal LLMs can feasibly generate robust synthetic normative data for cognitive tests—is supported only by the experiments described in the abstract. The provided full text, however, is a completely different paper: its header reads arXiv:2508.17674v2 [cs.CR] and it is titled 'Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models.' It contains no mention of Cookie Theft, GPT-4o, GPT-4o-mini, normative data, BLEU, ROUGE, BERTScore, diagnostic-group separability, or embedding-based analysis of cognitive responses. The manuscript body therefore supplies no methods, no dataset, no prompting details, no evaluation protocol, and no results that bear on the abstract's claim. This is not a subtle assumption failure or a debatable interpretation; the evidentiary basis for the central claim is entirely absent from the submitted document. The claim may be true, but the text under review cannot support it. Per the review rule to treat all manuscript passages as in-scope evidence, the mismatch itself is the decisive finding: the full text is not the study described by the abstract.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript (arXiv:2508.17675) is presented as a study of generating synthetic normative data for cognitive assessments using multimodal LLMs (GPT-4o and GPT-4o-mini), focusing on the Cookie Theft picture description task. The abstract reports qualitative findings about two prompting strategies, embedding-based separability of diagnostic groups, and standard text-generation metrics. However, the supplied full text is an unrelated security paper, 'Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models' (arXiv:2508.17674v2), which describes prompt- and model-distribution attacks on LLMs. The full text contains no methods, datasets, experiments, or results supporting the abstract's claims.","tokens_in":9125,"tokens_out":2517,"duration_ms":24096,"significance":"If the claimed study were present in proper form, it would address a real bottleneck in cognitive assessment—the cost of collecting normative data—by proposing a generative-model-based proxy, and it would introduce a concrete evaluation methodology (embeddings plus standard metrics). The potential significance of that idea is real, but the submitted document does not allow any assessment of it: the evidence base is entirely absent. The manuscript therefore cannot be considered a contribution to the cognitive-assessment literature in its current form.","major_comments":[{"comment":"The body of the manuscript is a different paper: its running head reads 'arXiv:2508.17674v2 [cs.CR]' and it is titled 'Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models.' None of the components named in the abstract—Cookie Theft, GPT-4o/GPT-4o-mini, naive versus advanced prompting, diagnostic-group separability in embedding space, BLEU/ROUGE/BERTScore, LLM-as-a-judge, or human normative samples—appears in the full text. The central claim of the abstract is therefore completely unsupported by the submitted document.","section":"Full text (entire document)"},{"comment":"The abstract's concluding statement that the study 'demonstrates' feasibility is not backed by any numeric result, sample size, effect size, error bar, or human validation in the abstract itself; the only evidence cited is qualitative ('more effectively distinguished,' 'higher realism and diversity'), which is insufficient to establish the claim.","section":"Abstract"},{"comment":"Even taken on its own terms, the evaluation protocol described in the abstract has a self-referential component: the LLM-as-a-judge evaluator and the generator come from the same model family (GPT-4o), and the diagnostic-group separability is computed from embeddings of synthetic responses without comparison to an external clinical ground truth. Without the actual methods and validation details, these design risks cannot be resolved.","section":"Abstract (evaluation design)"}],"minor_comments":[],"recommendation":"reject","confidential_remarks":"The submitted full text does not correspond to the abstract's paper; the arXiv identifier and title in the body indicate a different submission. This is not a correctable technical flaw in the experiments but a submission-integrity problem: the editor should verify that the correct manuscript was uploaded before any further review is considered. If the correct full text exists, a resubmission would be necessary."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know before you spend time on this. The full text under review is not the paper described in the abstract. The abstract is about generating synthetic normative data for cognitive assessments with GPT-4o; the actual manuscript is a security paper about advertisement embedding attacks against LLMs. There is no Cookie Theft, no BLEU/ROUGE/BERTScore, no normative data, no GPT-4o anywhere in the body. So every claim in the abstract is unsupported by the supplied document.\n\nTo be fair, the abstract poses a real research question. Using multimodal LLMs to produce provisional normative data for image-based cognitive tests is worth investigating, and the authors have articulated the motivation clearly: traditional norming is expensive, slow, and hard to update. If the actual paper matches the abstract, it could be a useful contribution to neuropsychological assessment methodology. I wanted to give credit for that.\n\nBut the soft spot is not a soft spot; it is the whole body. The submission appears to be a different arXiv paper (2508.17674v2). I cannot evaluate methods, sample sizes, analysis, or claims. Even the abstract alone is thin: no numeric results, no human validation, and the LLM-as-judge evaluation is partially circular because the judge comes from the same model family as the generator. Those would be revisable issues if the correct full text existed. The mismatch is not revisable; it makes the manuscript unverifiable.\n\nWho is this for? If the intended paper existed, it would be for clinical assessment researchers and maybe LLM evaluation people. As it stands, no one should cite it, and I would not bring it to reading group. The honest move is desk reject and ask the authors to resubmit the correct manuscript. If they do, treat it as a new submission and send it to a referee; the topic deserves a look. But this version does not deserve referee time.","headline":"The submitted full text is a different paper, so the abstract's claims about synthetic normative data are unverifiable from this submission.","tokens_in":9640,"tokens_out":2155,"would_cite":false,"duration_ms":21841,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Multimodal LLMs can synthesize normative data for cognitive tests, this paper argues.","keywords":["synthetic normative data","cognitive assessment","multimodal large language models","prompt engineering","Cookie Theft picture description","text embeddings","BERTScore","LLM-as-a-judge"],"falsifier":"Collect real human Cookie Theft descriptions with known diagnoses and demographic labels, generate synthetic responses under the advanced prompting strategy, and compare diagnostic-group separation in embedding space against the human data. If the synthetic groups separate in directions that do not match human group differences, or if apparent demographic variation in synthetic responses is only surface wording, the central feasibility claim fails.","tokens_in":8748,"feed_emoji":"🧠","tokens_out":4170,"duration_ms":46817,"temperature":0.7,"pith_summary":"The paper tries to establish that generative multimodal large language models can produce synthetic normative text for existing image-based cognitive assessments, removing a major bottleneck in test development. It argues that prompting strategy matters decisively: advanced prompts with contextual guidance yield synthetic responses that separate diagnostic groups and reflect demographic variation better than naive prompts. If the claim holds, new cognitive tests built on novel image stimuli could obtain provisional norms without expensive, slow human data collection, and existing tests could have their norms refreshed at low cost. The study is framed as a feasibility demonstration, supported by embedding-based group separability and several text-similarity metrics.","feed_headline":"LLMs can synthesize normative data for cognitive tests","feed_subtitle":"Advanced prompts make AI-generated responses separate diagnostic groups, opening a cheaper path to test norms.","key_machinery":"The central mechanism is the contrast between two prompting strategies applied to image-to-text generation: naive prompts containing basic instructions and advanced prompts enriched with contextual guidance about the task and its expected response patterns. The argument is carried by embedding-based analysis: generated responses are mapped into a vector space, and the degree to which those vectors separate diagnostic groups and demographic subgroups serves as the operational proxy for normative quality. Supporting machinery includes BLEU, ROUGE, and BERTScore for text similarity and an LLM-as-a-judge setup for holistic quality assessment.","core_discovery":"Using GPT-4o and GPT-4o-mini on the Cookie Theft picture description task, the paper claims that the quality of synthetic normative data is driven less by model scale than by how the model is prompted. Advanced prompting, enriched with contextual guidance about the assessment and expected response patterns, produced synthetic responses that more clearly separated diagnostic groups and captured more demographic diversity than naive prompting. The paper also reports that BERTScore tracked contextual similarity better than BLEU for these creative, open-ended responses, and that an LLM-as-a-judge evaluation offered preliminary but promising validation. The overarching claim is that generative multimodal LLMs, guided by refined prompting, can feasibly generate robust synthetic normative data for existing cognitive tests, thereby laying groundwork for developing novel image-based cognitive assessments without the traditional limitations of norm collection.","pith_inferences":["A reader should test whether embedding separability in synthetic responses actually tracks clinically meaningful differences by comparing against real human response distributions, since separability in vector space is not the same as clinical validity.","Demographic variation in synthetic norms may reflect stereotypes present in training data rather than true population structure, so subgroup calibration would need explicit handling before practical deployment.","A direct extension would be to measure distributional distance between synthetic and human embeddings on the same stimulus; if they diverge, improved prompting alone would not close the gap.","The same generate-and-embed pipeline could be applied to other norm-scarce open-ended clinical tasks, such as verbal fluency descriptions or autobiographical recall, where the cost of human norm collection is similarly prohibitive."],"forward_implications":["New image-based cognitive tests could generate provisional normative text within days instead of months, removing the recruitment bottleneck that currently limits test development.","Prompt design becomes a primary quality lever for synthetic norms, meaning careful contextual instructions matter at least as much as model choice.","BERTScore could serve as a dependable contextual-similarity metric for open-ended clinical language, while BLEU should not be trusted for creative outputs.","As generative models improve, normative data for existing tests could be regenerated and updated without repeating costly human collection.","A validation workflow for future tests would follow naturally: generate, embed, and check diagnostic-group separability before any clinical use."],"supporting_citations":[],"fun_headline_variants":["Prompting beats model size for synthetic cognitive norms","Better prompts, not bigger models, improve AI-generated test norms","Advanced prompting yields sharper diagnostic separation in synthetic norms","For AI-generated cognitive norms, prompt quality trumps model scale","GPT-4o shows prompting drives quality of synthetic normative data"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole result rests on assuming that measuring how similar generated responses are to each other in a mathematical embedding space tells us the same thing as measuring real patients' response patterns, so that when the generated responses of two diagnostic groups separate, that separation is clinically real.","fun_headline_variants_meta":{"raw":{"variants":["Prompting beats model size for synthetic cognitive norms","Better prompts, not bigger models, improve AI-generated test norms","Advanced prompting yields sharper diagnostic separation in synthetic norms","For AI-generated cognitive norms, prompt quality trumps model scale","GPT-4o shows prompting drives quality of synthetic normative data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000159,"raw_usage":{"total_tokens":1249,"prompt_tokens":983,"completion_tokens":266,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":599,"completion_tokens_details":{"reasoning_tokens":200}},"tokens_in":599,"tokens_out":266,"duration_ms":3020,"temperature":1.0,"reasoning_tokens":200,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T17:00:21.955691+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect real human Cookie Theft descriptions with known diagnoses and demographic labels, generate synthetic responses under the advanced prompting strategy, and compare diagnostic-group separation in embedding space against the human data. If the synthetic groups separate in directions that do not match human group differences, or if apparent demographic variation in synthetic responses is only surface wording, the central feasibility claim fails.","supporting_citations":[],"review_version":2}