{"id":"9c6397d0-83cf-4d48-9d57-dac601be77ff","arxiv_id":"2412.12445","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Persona-based prompting plus quality filtering yields more diverse suggested questions, and synthetic data from the pipeline trains a 360M model to near-GPT-4o levels.","lead":"Persona-SQ generates suggested questions for document apps by first inventing reader personas (profession and reading goal) and then writing questions tailored to each persona. The authors report more diverse, user-preferred questions, and show a 360M-parameter model trained on the pipeline's data rivals much larger models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No real-user validation: the synthetic persona proxy is never tested against actual reader profiles, so the 'personalized' claim may reduce to persona-conditioned diversity rather than personalization.","rationale":"The reader's weakest-assumption analysis identified exactly the load-bearing issue: synthetic personas are an unvalidated proxy for real user profiles. My reading of the full manuscript confirms this concern and finds no experiment that closes the gap. The human evaluation in Section 3.2 measures preference for one set of questions over another, but it never checks whether the 'preferred' questions match the preferences of a specific real user with a known profession and goal. The automatic alignment metrics are also generated and evaluated by the same LLM, making them consistency checks rather than evidence of genuine personalization. These concerns do not invalidate the engineering contribution: the pipeline produces more diverse and, by user preference, more attractive SQs than a non-personalized GPT4o baseline, and the fine-tuned small-model results are suggestive. But the central claim should be narrowed or subjected to a direct real-profile test before being accepted as a personalized SQ system. Since the reader's conditional verdict already reflects this limitation, I recommend no change to the verdict.","tokens_in":17976,"tokens_out":5574,"duration_ms":55995,"concrete_test":"Run a within-subject user study with at least 100 real readers, each reading documents from their own professional domain. For each participant, collect their actual profession and self-reported reading goal, then generate three question sets for the same document: (a) conditioned on the participant's real profile, (b) conditioned on a randomly chosen synthetic persona from the same domain, and (c) with no persona. Blind-rank and rate the three sets for preference, usefulness, and fit to the participant's needs. If (a) is not significantly preferred over (b), the synthetic persona proxy fails to support the personalization claim; if both (a) and (b) beat (c), the paper should be reframed as improving SQ diversity and appeal rather than as achieving personalization.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that conditioning SQ generation on personas improves personalization, but the paper never tests this against real users' actual backgrounds or reading goals. The only human study (Section 3.2, Tables 3 and 6) asks Prolific workers to rank generic questions for preference; it does not record each participant's profession or reading goal, nor does it compare questions generated from a participant's real profile against questions from a synthetic persona or from no persona. The automatic 'persona alignment' metric (Section C.2, Table 2) is also self-referential: GPT4o both generates/filters the questions and ranks which persona fits each question, so the high coverage ratios may reflect GPT4o's internal consistency rather than true alignment with user needs. The paper's own Limitations section concedes that Persona-SQ 'uses synthetically generated personas rather than actual user profiles' and 'does not yet achieve true personalization.' Therefore the measured gains in diversity and user preference support the weaker claim that persona-conditioned questions are more varied and broadly appealing, but not the stronger claim that they are actually personalized to the reader. The 'competitive with GPT4o' result is likewise only against non-personalized GPT4o, so it does not establish parity with a large personalized system.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Persona-SQ, a pipeline that generates suggested questions (SQs) for document-reading applications by first synthesizing reader personas (professions and reading goals) and then generating persona-conditioned questions, with LLM-based filtering for quality. The authors report two demonstrations: (1) an instantiation with GPT-4o that, compared with a non-persona GPT-4o baseline, produces more semantically diverse questions and higher coverage of the intended personas, and is preferred by human raters; and (2) a synthetic dataset generated with Llama-3.1-70B used to fine-tune a 360M-parameter SmolLM model whose SQs are competitive with or preferred over the non-persona GPT-4o baseline. The paper positions the framework as a drop-in upgrade for existing SQ systems and as a route to on-device personalized SQ generation.","tokens_in":18172,"tokens_out":6435,"duration_ms":57321,"significance":"If the central claims hold, the paper makes a useful engineering contribution: it shows that injecting synthetic profession/goal personas into SQ generation diversifies outputs and that the resulting synthetic data can train a very small on-device model whose outputs users rank favorably. The pipeline details, including prompts and quality-control steps, are described in unusual detail, and the human study with 400 raters is a serious attempt at validating user preference. However, the evidence supports the weaker claim that persona-conditioned SQs are more diverse and broadly appealing, not the stronger claim that they are personalized to real users. The paper's own Limitations section concedes that it 'does not yet achieve true personalization,' and the automatic persona-alignment metric relies on the same GPT-4o model that generated the questions. The small-model result is also compared only against non-personalized GPT-4o, not against a personalized large model. These gaps are load-bearing for the paper's framing and need to be addressed before the contribution can be accepted as stated.","major_comments":[{"comment":"The central claim of 'personalized SQs' is not supported by the experiments. Persona-SQ uses synthetically generated professions and goals rather than actual user profiles, and the Limitations section explicitly states that the approach 'does not yet achieve true personalization.' The human study (Section 3.2 and Appendix K) does not record each participant's profession or reading goal, nor does it compare questions generated from a participant's real profile against questions from a synthetic persona or from no persona. Consequently, the results demonstrate that persona-conditioned questions are more diverse and generally preferred, but not that they are tailored to the individual reader. Please either reframe the contribution as persona-conditioned SQ generation or add a study with real user profiles (for example, collecting each participant's self-reported profession and reading goal and evaluating whether persona-conditioned questions are preferred for matching profiles).","section":"Section 3.1, Appendix C.2, Appendix C.3"},{"comment":"The persona-alignment coverage ratio in Table 2 is computed by having GPT-4o rank personas for each generated question, but GPT-4o is also the model that generated the questions conditioned on those personas and filtered them in Section 2, Step 5. This creates a circularity: the high coverage ratios may reflect GPT-4o's ability to recover the conditioning persona from its own generated questions rather than genuine alignment with user needs. Additionally, Table 2 reports coverage ratios for the baseline, yet baseline questions have no intended persona; the paper does not specify how the 'intended' persona is assigned for baseline questions. To make this metric credible, use a different judge (a different model or human annotators) for the reverse ranking, and clearly define the baseline's intended-persona assignment.","section":"Abstract, Appendix I"},{"comment":"The abstract states that Persona-SQ is used to curate 'a large synthetic SQ dataset with 100k questions from thousands of diverse, real-world documents,' but Appendix I reports 'about 23k questions from around 1600 documents across a variety of professional documents,' and Table 10 sums to roughly 22k questions. This is a four-fold discrepancy in a headline number. Please correct the abstract and ensure all dataset statistics are consistent throughout the paper.","section":"Section 3.2, Tables 3 and 6"},{"comment":"The human preference results are reported as aggregate averages (e.g., Avg. Rank 2.88 vs. 4.12; Win Ratio 75.8%) with no confidence intervals, significance tests, or per-document variability, despite involving only 14 documents. Given that these results are the strongest evidence for user preference, please report standard errors or bootstrap confidence intervals and a paired statistical test across documents (or a mixed-effects model with document and participant as random effects). Also report how many questions were rated per document and whether any participant background information was collected.","section":"Section 4, Table 4"},{"comment":"The abstract and Section 4 claim that models fine-tuned on the Persona-SQ dataset 'outperform' GPT-4o, but Table 4 shows that on relevance, readability, and answerability, the GPT-4o baseline scores are higher than the Persona-SQ fine-tuned SmolLM (4.94 vs. 4.63, 5.00 vs. 4.77, and 4.86 vs. 4.17, respectively); only importance is higher for Persona-SQ. The human ranking in Table 6 is only against non-personalized GPT-4o and is not a comparison with a personalized large model. Please qualify the claim as 'competitive' or 'preferred in a human ranking' on specific metrics, and, if the claim is about outperforming GPT-4o on SQ generation, specify the metric and compare against a personalized large-model baseline.","section":"Section 4, Table 4"}],"minor_comments":[{"comment":"","section":"Section 2, Step 3"},{"comment":"The paper says it introduces 'five novel evaluation criteria' but only three are described in the main text (semantic diversity, persona alignment, and quality); the other two, persona distribution and coverage-ratio distribution skewness, are only in appendices. Please either introduce all five in the main text or adjust the wording.","section":"Section 3.1"},{"comment":"The text says 'Results in Table 4 show promising signal that users prefer the Persona-SQ fine-tuned small model over GPT4o baseline,' but Table 4 contains automatic quality scores, not user preferences; the user preference result is in Table 6. Please fix this cross-reference.","section":"Section 4, last paragraph"},{"comment":"The metric name 'Persona Distribution' in the appendix heading does not match the terminology in Section 3.1 ('Question Persona Alignment') or the later 'Coverage Ratio' in Appendix C.3. Please align the names to avoid confusion.","section":"Appendix C.2"},{"comment":"The description of the human evaluation procedure does not state whether the participants were screened for any reading-related background, whether the 14 documents were evenly distributed across the three domains, or whether the participants saw the document content beyond the title, summary, and URL. These details are important for interpreting the preference results.","section":"Appendix K"},{"comment":"There are several typos and formatting issues, including 'approahes' near Table 4, 'self-questions' in Appendix A, and an unclosed quote in the JSON example in Table 15 ('\"order 1\": \"persona3,'). A careful proofreading pass is needed.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper's core pipeline and on-device model results are interesting, but the 'personalization' framing is stronger than what the evidence supports, and the abstract overstates the dataset size by a factor of four. The authors should also be asked to address the circularity of the GPT-4o-based persona-alignment metric and the absence of statistical inference in the human study. These are fixable within a revision if the claims are reframed and appropriate analyses are added."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper is worth a serious look. It builds a pipeline that takes a document, generates synthetic reader personas (profession plus goals), uses an LLM to generate questions conditioned on each persona, filters them for quality and answerability, and shows that the resulting questions are more diverse and preferred by users than a non-persona baseline. The second half distills the synthetic data into a 360M-parameter model that holds its own against much larger models.\n\nWhat is genuinely new is the application: nobody else has looked at persona-conditioned suggested questions for document assistants, and the paper is honest that the individual techniques (prompting, filtering, distillation) are known. The evidence for the core claim is decent: embedding-based diversity drops, a 400-user ranking study shows a clear preference, and the small model's outputs are qualitatively sensible. The authors also explicitly acknowledge in the Limitations that they use synthetic personas and that this does not yet achieve true personalization. That is the right call.\n\nThe soft spots are real but not fatal. First, the evaluation loop is partly self-referential: GPT4o generates the questions and also judges quality and persona alignment. The coverage ratio compares each generated question against a list of personas including the one it was conditioned on, so high scores may partly reflect prompt adhesion rather than true alignment. Second, the human study asks Prolific workers to rank questions but does not record their actual profession or reading goal, so it supports 'more varied and broadly appealing' rather than 'personalized to me.' Third, there are no error bars or significance tests anywhere; the differences are large enough that significance is plausible, but the paper should say so. And the small-model 'competitive with GPT4o' claim is against non-personalized GPT4o, which is the right baseline for the drop-in claim but not parity with a personalized large model.\n\nThe stress-test note is right that the strongest reading of the paper is persona-conditioned diversity, not true personalization. But the authors already flag that limitation themselves, so it is not a hidden flaw.\n\nThe paper is for anyone working on question generation for reading assistants, on-device small models, or synthetic data pipelines. It deserves a real referee. My recommendation: send it to peer review, and ask the authors to add significance tests, a direct comparison against a personalized GPT4o baseline (even a small one), and ideally release code and data. The central contribution is sound.","headline":"A credible persona-conditioned suggested question pipeline with real human preference evidence; the true personalization claim is overstated, but the paper is honest about that and deserves a real referee.","tokens_in":18749,"tokens_out":1612,"would_cite":true,"duration_ms":14830,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that conditioning suggested questions on a synthetic reader profile—a profession and reading goals—makes them more diverse, better aligned, and more preferred, and that the same synthetic data trains a 360M-parameter…","keywords":["suggested question generation","persona","personalization","large language models","synthetic data","on-device model","reading assistants","question diversity"],"falsifier":"Give the same documents to readers whose real professions and reading goals are known, and generate SQs under three conditions: their own profile, a mismatched profile, and no profile. If readers do not reliably prefer questions generated for their own profile over the mismatched one, the personalization claim fails; if the no-profile condition ties the profile conditions, the gain is diversity without true personalization, which would tell against the paper's framing.","tokens_in":17743,"feed_emoji":"💬","tokens_out":10734,"duration_ms":88402,"temperature":0.7,"pith_summary":"Suggested questions are among the first things a user sees in an AI-powered reading app, but current generators condition only on the document, so different readers get similar questions. Persona-SQ adds a synthetic reader profile—a profession and a set of reading goals—to the generation prompt, then filters the candidate questions for persona relevance, document relevance, and answerability. The paper's central claim is that this conditioning alone makes generated questions more diverse, more aligned with the intended reader, and more preferred by users, as shown on finance, legal, and academic documents with GPT4o. The same pipeline can also generate synthetic training data: fine-tuning a 360M-parameter SmolLM model on Persona-SQ data outperforms fine-tuning on no-persona or public QA data and approaches the output quality of far larger models. If the claim holds, reading assistants can adopt the approach as a drop-in change and can deploy small local models that keep documents private.","feed_headline":"Persona-aware questions beat generic AI suggestions in user tests","feed_subtitle":"Profession and goal conditioning diversifies suggestions, and a 360M model trained on it rivals much larger ones.","key_machinery":"The machinery is a persona-goal conditioning variable built by generation plus filtering. In Steps 2–3, an LLM proposes professions and five reading goals per document, normalizes overlapping professions, and scores goals for relevance to the persona, keeping only scores of 4 or 5. In Steps 4–5, a second LLM generates questions conditioned on each profession–goal pair, then filters by length, by two relevance scores (question-to-persona and question-to-document), and by an answerability check that extracts an answer and supporting span from the document or discards the question. The persona–goal pair is the object that carries the argument: it forces the generator to spread questions across reader-interest axes instead of collapsing toward a generic, often domain-dominant persona such as 'lawyer' in the legal corpus.","core_discovery":"Persona-SQ's claim is that a reader profile, even a synthetic one, is the missing conditioning signal for suggested-question generation. Given a document, an LLM first proposes professions and reading goals, a filter keeps the high-quality persona–goal pairs, and a second LLM pass generates questions conditioned on each pair; further filters remove questions that are off-persona, off-document, or unanswerable. Across public finance, legal, and academic documents, GPT4o run through this pipeline produces questions with lower pairwise semantic similarity than GPT4o without persona information, higher coverage of the intended persona under an LLM-based reverse-ranking check, and higher user preference in a 400-participant ranking study (average rank 2.88 vs 4.12; win ratio 75.8% vs 24.2%). Fine-tuning SmolLM 360M on Persona-SQ synthetic data outperforms fine-tuning on no-persona or public QA data, and human raters prefer its questions over GPT4o no-persona baseline despite the model being far smaller. The author's conclusion is that persona-conditioned generation improves SQ quality and that synthetic data from the same pipeline transfers this capability to tiny deployable models.","pith_inferences":["The diversity gains may stem less from accurate personalization than from forcing the generator to spread questions over many invented interest axes; if so, the same pipeline could diversify questions along any user signal (reading level, language, task) even where accurate user modeling is unavailable.","A natural stress test is to compare real-profession profiles against synthetic ones on the same documents; this would separate the value of conditioning in general from the value of matching the actual reader.","Since the paper's alignment metric uses an LLM to rank personas for each question, an independent human-labeled persona-alignment set would show whether the coverage-ratio gains reflect true personalization rather than shared model bias."],"forward_implications":["Existing SQ systems can insert Persona-SQ in front of their current LLM and obtain more diverse, persona-aligned questions without changing the underlying model.","Synthetic persona-conditioned data can train 360M-parameter models whose SQ output is competitive with API models many times larger, enabling fully local and private generation.","The unoptimized 360M model takes about 760 MB in fp16, loads in about 0.5 seconds on a commercial CPU laptop, and generates a persona-plus-question in about 10 seconds; quantization could bring it near 200 MB.","Because the pipeline is extensible, the same persona-conditioning mechanism can incorporate other user signals—or later replace synthetic personas with real profiles collected from interaction logs—without architectural changes."],"supporting_citations":[{"why":"This paper supplies the GPT4o model used to instantiate Persona-SQ and also serves as the no-persona baseline and as the LLM judge in evaluations.","marker":"OpenAI, 2024"},{"why":"This paper supplies the open-source Llama-3.1-70B model used to generate the synthetic Persona-SQ fine-tuning dataset.","marker":"Dubey et al., 2024"},{"why":"This paper supplies the gte-Qwen2-1.5B-instruct embedding model used to compute the question semantic diversity metric.","marker":"Li et al., 2023"},{"why":"This paper grounds the use of an LLM-as-a-judge for question quality evaluation and win/tie comparisons.","marker":"Zheng et al., 2023"},{"why":"This paper supplies the FNS2020 finance documents used as one of the three evaluation domains.","marker":"El-Haj et al., 2020"},{"why":"This paper supplies the CUAD legal contract documents used as one of the three evaluation domains.","marker":"Hendrycks et al., 2021"},{"why":"This paper supplies the QASPER academic-paper documents used as one of the three evaluation domains.","marker":"Dasigi et al., 2021"}],"fun_headline_variants":["Persona-SQ: Personalized AI questions for better reading","Persona-aware AI suggestions outrank generic ones","Synthetic personas boost suggested question quality","Tiny model trained on persona data rivals big AI","Personalized AI questions beat generic ones in tests"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that synthetically invented profession-and-goal profiles are a faithful stand-in for real readers' information needs, so the gains measured under synthetic personas would survive contact with actual user profiles.","fun_headline_variants_meta":{"raw":{"variants":["Persona-SQ: Personalized AI questions for better reading","Persona-aware AI suggestions outrank generic ones","Synthetic personas boost suggested question quality","Tiny model trained on persona data rivals big AI","Personalized AI questions beat generic ones in tests"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000517,"raw_usage":{"total_tokens":2502,"prompt_tokens":938,"completion_tokens":1564,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":1492}},"tokens_in":554,"tokens_out":1564,"duration_ms":10945,"temperature":1.0,"reasoning_tokens":1492,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T14:04:13.275427+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give the same documents to readers whose real professions and reading goals are known, and generate SQs under three conditions: their own profile, a mismatched profile, and no profile. If readers do not reliably prefer questions generated for their own profile over the mismatched one, the personalization claim fails; if the no-profile condition ties the profile conditions, the gain is diversity without true personalization, which would tell against the paper's framing.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This paper supplies the FNS2020 finance documents used as one of the three evaluation domains."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"This paper supplies the CUAD legal contract documents used as one of the three evaluation domains."}],"review_version":1}