{"id":"ea1238ef-ef75-4ae9-b7f6-c1699705aa77","arxiv_id":"2606.26654","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"SocialPersona benchmark shows MLLMs detect broad user interests from multimodal timelines but drop sharply on fine-grained and recent interests when generating personalized responses.","lead":"The paper introduces SocialPersona, a benchmark built from 171 users' social-media timelines with text, images, and verified preference tags to test if multimodal LLMs can infer interests and personalize dialogue. A smart generalist might read it to see how far current AI is from understanding people from their natural online traces rather than explicit statements.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"The abstract-only review already located the single load-bearing precondition for the experimental claims. No additional technical flaw (e.g., in task formulation, metric definition, or model comparison) is visible that would independently threaten the headline result.","tokens_in":1720,"tokens_out":218,"duration_ms":13490,"concrete_test":"Re-annotate a random 20% subsample of the 171 timelines using the paper's stated annotation protocol and measure tag overlap (precision/recall) against the released set; if overlap falls below 85% on fine-grained or recent tags, the performance-drop claims require re-evaluation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption correctly isolates the central dependency: the benchmark's validity hinges on whether the 2,597 human-verified tags comprehensively and accurately reflect the interests latent in the 171 multimodal timelines. The abstract provides no further internal inconsistency or unsupported assumption that would undermine the reported performance patterns once this ground truth is granted.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces SocialPersona, a benchmark built from longitudinal multimodal timelines of 171 everyday social-media users containing text, images, timestamps, and 2,597 human-verified preference tags across seven domains (separating stable from recent interests). It defines two tasks—constructing structured user profiles from multimodal context and generating dialogue responses aligned with inferred profiles—and reports that MLLMs identify broad domains reasonably well but show clear drops on fine-grained/recent interests, with further degradation when inferred profiles are used for personalization; text and images are shown to provide complementary signals.","tokens_in":1754,"tokens_out":487,"duration_ms":20500,"significance":"If the ground-truth tags and timeline coverage are reliable, SocialPersona supplies a concrete, falsifiable testbed for cross-modal, long-horizon preference inference that goes beyond explicit memory recall. The separation of stable versus recent interests and the two-stage evaluation (profile construction then response generation) are useful design choices that could help the community quantify progress on revealed-preference modeling.","major_comments":[{"comment":"The central empirical claims rest on the claim that the 2,597 human-verified tags accurately and comprehensively reflect the stable and recent interests latent in the 171 multimodal timelines; the manuscript must supply a detailed account of the verification protocol, inter-annotator agreement, coverage statistics, and any filtering criteria (e.g., §3 or Dataset Construction) before the reported performance drops can be interpreted as evidence of model limitations rather than annotation artifacts.","section":"Dataset Construction"},{"comment":"The experiments section reports performance degradation when moving from broad domains to fine-grained/recent interests and from profile construction to dialogue generation, yet provides no statistical significance tests, confidence intervals, or error analysis that would establish these drops are robust rather than artifacts of prompt sensitivity or small per-user sample sizes.","section":"Experiments"}],"minor_comments":[{"comment":"Clarify the exact split between proprietary and open-weight models evaluated and report per-model numbers rather than aggregated trends only.","section":"Experiments"},{"comment":"Add a limitations paragraph discussing potential demographic or platform biases in the 171-user sample and the seven interest domains.","section":"Discussion"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their thoughtful review and constructive suggestions. We address each of the major comments below, and we plan to incorporate revisions to strengthen the manuscript.","responses":[{"response":"We agree that additional details on the dataset construction are essential for interpreting the results. The current manuscript provides an overview of the human verification process, but we acknowledge it lacks the requested granularity. In the revised version, we will expand the Dataset Construction section (likely §3) to include: (1) the full verification protocol, including annotator instructions and guidelines for identifying stable vs. recent interests; (2) inter-annotator agreement statistics, such as percentage agreement and Cohen's kappa where applicable; (3) coverage statistics, e.g., average tags per user, distribution across domains, and timeline length coverage; and (4) any filtering criteria applied to select users and tags. This will help demonstrate that the tags reliably capture the latent interests.","revision_made":"yes","referee_comment":"[Dataset Construction] The central empirical claims rest on the claim that the 2,597 human-verified tags accurately and comprehensively reflect the stable and recent interests latent in the 171 multimodal timelines; the manuscript must supply a detailed account of the verification protocol, inter-annotator agreement, coverage statistics, and any filtering criteria (e.g., §3 or Dataset Construction) before the reported performance drops can be interpreted as evidence of model limitations rather than annotation artifacts."},{"response":"We concur that statistical rigor would bolster the experimental claims. We will revise the Experiments section to include: bootstrap-derived confidence intervals for all reported metrics; statistical significance tests (e.g., paired t-tests or McNemar's test) for the performance differences between broad vs. fine-grained, stable vs. recent, and profile vs. response generation tasks; and an expanded error analysis categorizing model failures by interest granularity, recency, and input modality. To address prompt sensitivity, we will report results from at least two distinct prompt templates. While the per-user sample size is constrained by the 171 timelines, we will clarify that metrics are aggregated across users with appropriate variance estimates.","revision_made":"yes","referee_comment":"[Experiments] The experiments section reports performance degradation when moving from broad domains to fine-grained/recent interests and from profile construction to dialogue generation, yet provides no statistical significance tests, confidence intervals, or error analysis that would establish these drops are robust rather than artifacts of prompt sensitivity or small per-user sample sizes."}],"tokens_in":1391,"tokens_out":503,"duration_ms":28117,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The core contribution is a benchmark drawn from 171 everyday users' longitudinal timelines that include text, images, and timestamps, plus 2,597 human-verified tags split into stable and recent interests. It defines two tasks: building structured profiles from the multimodal context and then generating dialogue that matches those profiles.\n\nThe work shows models can pick up broad domains but drop on fine-grained and recent interests, with further loss when the inferred profile is used for response generation. The observation that text and images supply complementary signals is straightforward and worth having on record. The stable/recent split also adds a temporal dimension that most existing personalization tests lack.\n\nThe main limitation is that everything depends on the quality of those verified tags. The abstract supplies no details on how the tags were generated, what verification protocol was used, or any measure of agreement, so it is impossible to judge whether the reported performance gaps reflect model shortcomings or annotation artifacts. The sample of 171 users is modest, and without more on selection criteria or domain coverage the generalizability stays unclear.\n\nThis is aimed at groups working on long-horizon multimodal personalization who need something beyond synthetic or single-turn data. A reader looking for empirical baselines on current MLLM limits would find the results usable.\n\nI would send it to peer review. The benchmark idea is timely and the empirical patterns are worth checking once the construction details are filled in.","headline":"SocialPersona gives a practical new benchmark for inferring preferences from real multimodal timelines, but its claims rest heavily on unexamined tag verification.","tokens_in":2245,"tokens_out":356,"would_cite":true,"duration_ms":22420,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Multimodal models recover broad user interests from social-media timelines but lose accuracy on fine-grained and recent preferences when generating personalized responses.","keywords":["multimodal large language models","personalized profiling","social media timelines","preference inference","benchmark evaluation","dialogue personalization","user modeling"],"falsifier":"A model that matches or exceeds human accuracy on fine-grained and recent-interest tags while preserving that accuracy when its inferred profiles are used to generate dialogue responses would refute the reported performance gaps.","tokens_in":2621,"feed_emoji":"📱","tokens_out":619,"duration_ms":19158,"temperature":0.7,"pith_summary":"The paper creates SocialPersona to measure whether multimodal large language models can extract stable and recent preferences from users' real social-media timelines that include text, images, and timestamps, then apply those inferences in dialogue. A sympathetic reader would care because everyday personalization requires models to notice what people care about from traces they already leave behind rather than waiting for explicit statements in chat. Experiments across proprietary and open models show reliable detection of broad domains alongside clear drops on specific or time-sensitive interests, plus further losses when the inferred profile must shape responses. The results also indicate that text and images supply distinct but additive signals about preferences.","feed_headline":"Models spot broad interests in social timelines but miss fine details","feed_subtitle":"Benchmark shows MLLMs recover domain-level preferences yet drop on recent and specific ones, with further losses when profiles must shape re","key_machinery":"SocialPersona benchmark built from 171 users' multimodal timelines annotated with 2,597 human-verified preference tags across seven domains, supporting two tasks of structured profile construction from context and generation of profile-aligned responses.","core_discovery":"SocialPersona shows that current MLLMs identify broad interest domains from multimodal longitudinal timelines yet suffer measurable drops in accuracy on fine-grained and recent interests, with additional degradation when the resulting profiles are required to produce aligned dialogue responses; text and images supply complementary preference signals.","pith_inferences":["The benchmark could be reused to measure whether new architectures improve temporal tracking of preference shifts.","It highlights a possible need for explicit mechanisms that separate stable traits from transient signals when building user models.","Similar evaluation setups might apply to other public multimodal traces such as photo streams or forum histories."],"forward_implications":["Models achieve higher accuracy on broad interest domains than on fine-grained or recent ones.","Performance declines further when inferred profiles must drive response generation rather than profile construction alone.","Text and images supply distinct preference signals that together improve recovery.","Robust cross-modal modeling over long time horizons remains difficult for current systems."],"fun_headline_variants":["MLLMs spot broad interests but falter on fine and recent ones","Models recover domain level preferences from multimodal timelines","Further drops when profiles must shape dialogue responses","Text and images provide complementary signals for preferences"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The 2,597 human-verified preference tags accurately and comprehensively capture the stable and recent interests shown in the 171 users' social-media timelines.","fun_headline_variants_meta":{"raw":{"variants":["MLLMs spot broad interests but falter on fine and recent ones","Models recover domain level preferences from multimodal timelines","Further drops when profiles must shape dialogue responses","Text and images provide complementary signals for preferences"]},"model":"grok-4.3","cost_usd":0.008695,"raw_usage":{"total_tokens":3909,"prompt_tokens":647,"num_sources_used":0,"completion_tokens":59,"cost_in_usd_ticks":86949500,"prompt_tokens_details":{"text_tokens":647,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3203,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":647,"tokens_out":59,"duration_ms":25281,"temperature":1.0,"reasoning_tokens":3203,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T05:16:07.080000+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A model that matches or exceeds human accuracy on fine-grained and recent-interest tags while preserving that accuracy when its inferred profiles are used to generate dialogue responses would refute the reported performance gaps.","supporting_citations":[],"review_version":1}