{"id":"8be8b0b9-466b-4190-811a-0f89744a5057","arxiv_id":"2607.11039","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Persona agents grounded in real career posts cut cognitive effort versus passive social-media browsing and redirect upward comparison toward constructive self-reframing while users still need authentic peer content for emotional grounding.","lead":"JobMate turns real social-media career posts into chatty AI personas so job seekers talk with peer stories instead of doomscrolling. A 24-person study found lower mental effort than native browsing and steered comparison toward next steps rather than “I’m not good enough.”","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"The comparison-redirection claim is not isolated from structure, ads, and LLM confounds; modality is confounded with the full JobMate package.","rationale":"The reader correctly flags the 30-minute lab task and small N as hard bounds and keeps CONDITIONAL. That ecological-validity concern is real (§5.2; §7.4) but secondary: even if the short task perfectly mirrored naturalistic comparison, the present design still cannot pin the headline effect on “shifting interaction modality” while holding content fixed, because the control confounds modality with ads, layout, challenge-tag framing, and LLM support. The paper’s own mechanism language in §7.1 (cognitive offloading + output-driven sensemaking + comparison reframing) bundles those factors; only an ablation or structured non-dialogue arm would settle which is load-bearing. I therefore agree with CONDITIONAL and with treating N/short-task limits as bounds, but the more precise load-bearing soft spot is causal isolation of modality, not only naturalism. No change to REJECT: methods are clear, both arms reduce CDDQ, and the system contribution remains useful if claims stay package-level rather than pure-modality.","tokens_in":21343,"tokens_out":657,"duration_ms":7020,"concrete_test":"Run a 2×2 or three-arm between-subjects study (same N/discipline strata and 30-min task): (A) native RedNote; (B) JobMate gallery + original posts only, no chat; (C) full JobMate with persona dialogue. Pre-register NASA-TLX Effort and a coded comparison-direction measure (upward threat vs. self-reframing/next-step). If Effort and reframing improve in B≈C ≫ A, structure/framing—not dialogue—carry the claim; if C ≫ B, modality is load-bearing.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim attributes lower NASA-TLX Effort (p_t=0.012) and redirection of social comparison from upward threat to constructive self-reframing to “AI-mediated dialogue” while “authentic peer content is held fixed” (Abstract; §1; §6.1–6.2; §7.1). The control is unconstrained native RedNote browsing (§5.2), not a structured-card or original-post-only condition. JobMate simultaneously removes ads/noise (DG1, Stage 1–2), compresses posts into challenge-foregrounded cards (DG3, Stage 3), shows the original post as an authenticity anchor, and adds dual-track RAG + SDT-framed persona chat (§4.2–4.4). Qualitative quotes credit “clean… cards save selection cost” and challenge tags as a “mutual-aid group” as much as dialogue (§6.1–6.2). §7.4 explicitly notes components were not ablated. Therefore the causal attribution to interaction modality (vs. de-noising, challenge framing, or LLM coaching) is not secured by the design, even if short-task ecological validity is set aside.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"The paper presents JobMate, a system that converts real RedNote career posts into challenge-foregrounded persona cards and SDT-informed, dual-track RAG conversational agents, aiming to keep authentic peer content while shifting users from passive feed browsing to active dialogue. A formative survey/interview study (N=64/8) motivates four design goals; a between-subjects lab study (N=24; CS, Psychology, Chinese Literature) compares JobMate to unconstrained native RedNote browsing on a 30-minute career-exploration task. Both conditions reduce CDDQ scores with no between-group difference on total reduction; JobMate shows lower NASA-TLX Effort (p_t=0.012) and marginally lower Frustration, with qualitative themes that comparison shifts from upward threat toward self-reframing and next-step sensemaking while authenticity of real posts remains the emotional anchor. The authors argue that interaction modality, not content alone, shapes the value–harm tension in peer-experience consumption, and they offer design implications for support modes, transparency, and coverage.","tokens_in":21760,"tokens_out":1215,"duration_ms":11073,"significance":"If the core claim holds—that redesigning interaction around authentic UGC can lower cognitive cost and redirect social comparison without discarding peer authenticity—the work is a useful contribution to HCI systems for career exploration and, more broadly, high-comparison UGC settings (study, fitness, parenting, health). Strengths include a complete end-to-end pipeline (cleaning, classification, structured extraction, dual-track RAG, SDT dialogue rules), explicit design goals tied to formative findings, multi-instrument evaluation (CDDQ, NASA-TLX, SDS, logs, interviews), and honest limitations on sample size, short task, single-platform coverage, and missing ablations. The contribution is primarily design-and-evaluation rather than a tightly isolated causal mechanism; its value for the field depends on how carefully claims about “modality” versus the full JobMate package are scoped.","major_comments":[{"comment":"Abstract, §1, §6.1–6.2, and §7.1 attribute lower Effort and redirection of social comparison primarily to “AI-mediated dialogue” / interaction modality while “authentic peer content is held fixed.” The control is unconstrained native RedNote browsing (§5.2), not a structured-card-only or original-post-only arm. JobMate simultaneously de-noises ads, compresses posts into challenge-first cards, anchors the original post, and adds dual-track RAG + SDT coaching (§4.1–4.4). Qualitative evidence credits clean cards and challenge tags as much as dialogue (§6.1–6.2), and §7.4 states components were not ablated. The causal claim for modality alone is therefore not secured; either add ablation/control conditions or reframe claims as effects of the full JobMate package versus native feed browsing.","section":null},{"comment":"§5.1–5.2 and §6: N=24 (n=4 per discipline×condition cell) with Mann–Whitney subgroup tests is underpowered for the boundary-condition claims in §6.3 and for treating disciplinary cognitive style as a robust moderator. Pre–post CDDQ improves in both arms with no between-group total difference (p_t=0.432); the quantitative headline rests on Effort (p_t=0.012) and a marginal Frustration result, while the comparison-redirection claim is almost entirely thematic. The manuscript should (i) center the primary confirmatory contrast, (ii) label discipline analyses as exploratory, and (iii) avoid overstating “redirected social comparison” relative to the mixed quantitative pattern.","section":null},{"comment":"§5.2 procedure and §7.4: the 30-minute fixed-objective lab task is a weak proxy for multi-session naturalistic job-seeking comparison and sensemaking. The paper’s own limitations note that anxiety rebound and action conversion were not observed. Given that the strongest claim concerns how comparison is experienced and converted into next steps, the ecological-validity gap is load-bearing; either strengthen the discussion of what short-task evidence can and cannot support, or plan/report a longer field deployment as central rather than future work.","section":null}],"minor_comments":[{"comment":"Figure 5 caption says plots label the control as “Baseline” for native RedNote; keep terminology consistent with “RedNote” throughout text and figures to avoid confusion with a true baseline condition.","section":null},{"comment":"Implementation (§4.5) states conversational agents use GPT-5.2 while supplementary materials list gpt-4o-mini defaults; reconcile model names and report the actual deployment model used in the user study.","section":null},{"comment":"Table 1 footnote uses † for both “Higher is better” and “Marginal (p<.10)”; disambiguate symbols.","section":null},{"comment":"§6.1 reports two exploration strategies (deep divers n=5, broad explorers n=4) from logs; a brief definition of how mixed users were classified would help reproducibility.","section":null},{"comment":"Related Work §2.3 and contributions: clarify more sharply what is novel relative to PlanHelper, DesignQuizzer, ComViewer, and SDT career chatbots beyond grounding personas in real others’ posts.","section":null},{"comment":"Supplementary participant tables and full prompts are valuable; ensure the camera-ready main text points to them and that any Chinese-to-English instrument rendering notes are explicit for CDDQ/SDS adaptations.","section":null}],"recommendation":"major_revision","confidential_remarks":"Fit for a solid HCI systems venue after revision. The main risk is over-claiming modality isolation; if authors reframe to “JobMate package vs native feed” and demote subgroup claims, the paper becomes much stronger. No integrity red flags; circularity is low. Sample and ablation gaps are the real barriers to a stronger recommendation."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: JobMate is a clean, usable pattern—real peer posts turned into challenge-first cards plus authenticity-anchored persona chat—and the study shows lower effort than native RedNote with similar short-term CDDQ gains. The headline that “AI-mediated dialogue” alone redirects social comparison is not isolated by the design.\n\nWhat is actually new is the combination, not any single piece. Prior work either structures UGC and keeps passive browsing, or builds career/persona agents on synthetic or self-configured content. Grounding conversational agents in real others’ RedNote posts, foregrounding challenges rather than wins, dual-track RAG (similar struggles + destination knowledge), and keeping the original post visible as an authenticity anchor is a coherent contribution. The formative work is tight, the pipeline is transparent (including prompts in the supplement), and the discussion is honest that users still lean on real posts for emotional grounding. That last point is useful: AI expands interactivity; trust still rides on authentic sources.\n\nSoft spots, in proportion. N=24 with n=4 per cell is thin for discipline claims; treat those as directional. The 30-minute lab task is a weak proxy for multi-session job-seeking anxiety and action—authors flag this. The stress-test lands: the control is unconstrained RedNote, not structured cards or original-post-only. JobMate simultaneously kills ads, compresses into challenge tags, shows the source post, and adds SDT-framed LLM chat. Quotes credit clean cards and “mutual-aid” tags as much as dialogue. Without ablation, attributing effects to modality versus de-noising or coaching is overstated. That is a real design gap, not a fatal one for a systems paper if claims stay bounded.\n\nMath and measures are standard (CDDQ, NASA-TLX, SDS); no circular redefinition. Citations cover social comparison, sensemaking, persona agents, and UGC tools fairly. No machine-checked proofs or shipped public corpus, but that is normal for this venue class.\n\nWho it is for: HCI and career-tech designers working on high-comparison UGC. Theory people will want stronger causal isolation. I would send it to peer review—important enough and clear enough to deserve referee time, with revision pressure on over-claiming modality and on N/longitudinal limits. Worth engaging if you care about AI that augments authentic peer content rather than replacing it.","headline":"Solid CHI system paper with a real design pattern; the modality claim is confounded with the full package, but the work still earns a serious referee.","tokens_in":22333,"tokens_out":591,"would_cite":true,"duration_ms":12358,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"When peer career posts stay fixed, shifting from passive scrolling to AI persona dialogue lowers cognitive effort and redirects upward comparison into self-reframing.","keywords":["Career exploration","Social comparison","Sensemaking","Persona-grounded agents","Social media UGC","Retrieval-augmented generation","Self-determination theory"],"falsifier":"A larger multi-week field study comparing JobMate-style dialogue with native browsing that found no NASA-TLX Effort difference, no qualitative shift from upward-comparison language to next-step self-reframing, and no continued reliance on real posts for emotional grounding would falsify the claim that modality, not content, drives the outcome.","tokens_in":22212,"feed_emoji":"💬","tokens_out":920,"duration_ms":15587,"temperature":0.7,"pith_summary":"Young job seekers use authentic peer posts for career grounding, but passive feeds also produce overload and upward comparison anxiety. This paper claims the tension is driven more by interaction modality than by the content itself. JobMate converts real social-media career posts into structured persona cards and Self-Determination-Theory-guided conversational agents, keeping the original posts visible as authenticity anchors. In a between-subjects study of 24 students across three disciplines, both native browsing and JobMate reduced career-decision difficulties, yet JobMate did so at significantly lower effort and shifted users from “I’m not as good as others” toward “what should I do next,” while still relying on real posts for emotional grounding. The result matters because it shows designers can keep the value of authentic peer stories while redesigning how people encounter them.","feed_headline":"AI chat turns peer career posts into lower-cost self-reframing","feed_subtitle":"Same authentic stories, less effort and less “I’m worse than them” when people talk to personas instead of scrolling.","key_machinery":"JobMate: a four-stage pipeline that cleans and classifies real career posts, extracts person-centric fields (background, outcome, challenges tags, summary), builds dual-track retrieval-augmented personas, and runs dialogue under Self-Determination Theory rules (relatedness via empathic self-disclosure, competence via reframing, autonomy via non-directive options). The mechanism converts passive feed consumption into active, grounded conversation while leaving the original post visible.","core_discovery":"Holding authentic peer career content fixed, AI-mediated persona dialogue reduces cognitive cost relative to native social-media browsing and redirects social comparison from potentially detrimental upward comparison toward constructive self-reframing and next-step sensemaking; users nevertheless continue to treat the real posts as the emotional and trust anchor.","pith_inferences":["The same modality shift—authentic posts kept, interaction changed to grounded dialogue—could reduce comparison harm in fitness, parenting, academic-grade, or chronic-illness peer content without removing the stories people trust.","Longitudinal deployment is needed to test whether short-term effort savings and self-reframing convert into actual career actions rather than rebound anxiety.","Coverage bias in the underlying posts (over-representation of tech or large-company paths) can make the method feel templated for underrepresented trajectories unless retrieval deliberately rebalances them.","Component ablations (cards alone vs dialogue alone vs SDT framing) would isolate which piece actually redirects comparison direction."],"forward_implications":["Interaction modality, not content alone, shapes whether authentic peer experiences produce anxiety or usable self-knowledge.","Foregrounding challenges rather than achievements on persona cards can steer comparison toward lateral normalization instead of upward threat.","AI systems that ground dialogue in real user-generated content can match the perceived support of human platforms while lowering screening cost.","Career and other high-comparison domains need adjustable density and emotion-versus-action modes for different cognitive styles.","Showing the real-post basis of each persona is required for users to trust AI-mediated peer experience."],"fun_headline_variants":["AI personas turn peer career posts into self-reframing talks","Chatting job-story personas cuts comparison cost vs scrolling","Dialogue with AI peers redirects envy into constructive reframing","Same posts, lower load: personas shift comparison to sensemaking","Persona chat grounds career sense while easing upward comparison"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"A single 30-minute laboratory exploration task with fixed objectives is assumed to capture the same social-comparison dynamics, sensemaking depth, and emotional grounding that arise in naturalistic multi-session job-seeking on social media.","fun_headline_variants_meta":{"raw":{"variants":["AI personas turn peer career posts into self-reframing talks","Chatting job-story personas cuts comparison cost vs scrolling","Dialogue with AI peers redirects envy into constructive reframing","Same posts, lower load: personas shift comparison to sensemaking","Persona chat grounds career sense while easing upward comparison"]},"model":"grok-4.5","effort":"low","cost_usd":0.004924,"raw_usage":{"total_tokens":1346,"prompt_tokens":740,"num_sources_used":0,"completion_tokens":63,"cost_in_usd_ticks":49240000,"prompt_tokens_details":{"text_tokens":740,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":543,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":740,"tokens_out":63,"duration_ms":4466,"temperature":1.0,"reasoning_tokens":543,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T07:28:04.059974+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A larger multi-week field study comparing JobMate-style dialogue with native browsing that found no NASA-TLX Effort difference, no qualitative shift from upward-comparison language to next-step self-reframing, and no continued reliance on real posts for emotional grounding would falsify the claim that modality, not content, drives the outcome.","supporting_citations":[],"review_version":1}