{"id":"7b70778c-d275-4cac-8bd2-0db80aeed7cd","arxiv_id":"2608.06898","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A community evaluation framework for social-robot foundation models: five dimensions filtered through static benchmarks, simulated interactions, and robot-specific testing, with Pareto frontiers under latency constraints.","lead":"This paper proposes a three-tier evaluation funnel to help robotics researchers choose a foundation model for social robots, starting with cheap static benchmarks and ending with robot-specific testing. It maps five evaluation dimensions across the tiers and asks the community to build a shared leaderboard.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Funnel's cost saving depends on Tier-2 simulations ranking models like real users would; the paper cites evidence against this but provides no calibration, leaving the 'better informed' half of the central claim unsupported.","rationale":"The paper is a clearly written, well-referenced position paper, and I read it in good faith as a call to action rather than a completed empirical study. Its contributions—five evaluation dimensions, a three-tier funnel, and a coverage map—are useful structuring ideas. My stress-test focuses on the central value proposition: the funnel must make model selection both cheaper and better informed. Cheaper is plausible by construction, but better informed depends on the predictive validity of Tier 2's LLM-based simulations. That is exactly the weakest point the Reader identified. I agree with that assessment, and I did not find a separate concern that outweighs it. The paper deserves credit for explicitly citing the dangers of LLM-based social simulation and judge self-preference, but citing a known failure mode is not the same as demonstrating that the proposed funnel is robust to it. The concrete test is a small proof-of-concept that would directly measure Tier-2-to-Tier-3 rank correlation. Until such evidence exists, the appropriate verdict remains CONDITIONAL: the framework is plausible and worth pursuing, but its central claim is not yet supported. I therefore recommend no change to the Reader's verdict.","tokens_in":6615,"tokens_out":3502,"duration_ms":43609,"concrete_test":"Implement the Haru classroom case study (Section III): run Tier 1 and Tier 2 on 8–10 current open-weight models with the stated 2-second latency cutoff, then run a Tier 3 user study with approximately 30 Japanese high-school students scoring the top-3 and bottom-3 Tier-2 models on the five dimensions. Compute Kendall's tau between Tier-2 aggregate scores and Tier-3 human scores per dimension. If tau is below about 0.3, or if a bottom-3 Tier-2 model beats a top-3 model, the funnel's pruning is not predictive and the central claim needs qualification; if tau is strongly positive, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section II.B positions Tier 2 as the main pruning stage: thousands of simulated conversations per candidate, with only the survivors reaching expensive Tier 3. For the funnel to deliver 'cheaper and better informed' selection, Tier 2 rankings must predict Tier 3 outcomes—i.e., models that score well with LLM users and judges must also score well with the real target population on the real robot. The paper itself flags the risks: simulation without information asymmetry inflates social competence (Zhou et al., 2024), judges favor their own generations (Panickssery et al., 2024), and optimizing against a fixed judge drifts from human judgment (Wang et al., 2024). These are systematic biases, not random noise, and they may affect model families differently—e.g., a judge LLM may reward stylistic similarity to itself, or LLM-simulated students may not reproduce adolescent turn-taking patterns. If Tier 2 misorders candidates, scarce Tier 3 resources are spent validating the wrong models, and the final picks are no better than a guess. The Section III case study asserts the funnel will work but provides no evidence of predictive validity. The 'None of this is hypothetical' claim therefore overstates the current support for the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper addresses the problem of selecting a foundation model for social robotics applications, arguing that public leaderboards are misaligned with the needs of real-time embodied social interaction and that per-lab user studies are too costly. It proposes five evaluation dimensions (conversational competence, user safety, embodied character, target scene effectiveness, audience appropriateness) and a three-tiered evaluation funnel: Tier 1 curated static benchmarks, Tier 2 simulated interactions with LLM users and judges, and Tier 3 robot-specific evaluation, with a Pareto-frontier decision rule over performance versus deployment cost. The paper maps existing evaluation tools onto the resulting coverage matrix and closes with a call for a community-built shared leaderboard and platform-specific harnesses.","tokens_in":6956,"tokens_out":5467,"duration_ms":57541,"significance":"The funnel framework is a timely and useful organizing device for a field that increasingly needs cost-effective, reproducible model comparison. The paper is transparent about known limitations of LLM-based evaluation, correctly distinguishing priors from verdicts, and the coverage map provides a concrete starting point for community work. If the funnel's predictive validity can be established, it could genuinely lower the cost of model selection and improve cross-project comparability. However, the central claim that the funnel is 'cheaper and better informed' is not yet supported by evidence, and several load-bearing details of the proposal remain underspecified.","major_comments":[{"comment":"The claim that the funnel makes model selection 'cheaper and better informed' (Abstract) rests on the assumption that Tier 2's LLM-based simulated interactions rank candidates similarly to real users on the target robot in the target scene. The paper itself cites systematic biases: judges favor their own generations (ref. [29]), simulation without information asymmetry inflates social competence (ref. [30]), and optimizing against a fixed judge drifts from human judgment (ref. [31]). No evidence is provided that Tier 2 rankings correlate with Tier 3 outcomes, nor is there a calibration strategy. Without such evidence, the funnel may prune the wrong models and direct scarce Tier 3 resources to suboptimal candidates. Please either add a validation proposal (e.g., a retrospective study comparing funnel selections with full Tier 3 results for a set of models) or explicitly frame Tier 2 as a hypothesis-generation stage rather than a pruning stage, and temper the 'better informed' claim accordingly.","section":"Section II.B"},{"comment":"The statement 'None of this is hypothetical' overstates the paper's contributions. The case study in Section III is illustrative and does not provide evidence that the funnel's Tier 1 and Tier 2 stages would actually identify the 'one or two vetted models' that survive. No experiments, leaderboard implementations, or simulation results are presented. As a position paper, this is acceptable if the claims are framed as a research agenda; as written, the claim that the funnel is operational today is not supported. Please revise the wording to distinguish between what can be assembled from existing pieces and what has actually been demonstrated.","section":"Section II (introduction to the funnel)"},{"comment":"The decision rule 'pick your point on the Pareto frontier' is underspecified because the evaluation has five dimensions. The paper says the frontier is computed per dimension, but the conclusion presents a single 'Pareto frontier' as the answer. It is unclear how a practitioner should combine or trade off across dimensions: for example, a model on the frontier for safety may be off the frontier for conversational competence. To make the rule actionable, please specify either how per-dimension frontiers are aggregated into a single frontier, or describe the intended multi-objective decision procedure (e.g., lexicographic ordering by user priorities, or a weighted scalarization).","section":"Section II (Pareto frontier)"}],"minor_comments":[{"comment":"Several entries in the coverage map (SHREC, REPAIR-Bench, MinorBench, RoleLLM, TRAIT) are not cited in the text or reference list; please add references or remove the table until the entries can be verified.","section":"Table I"},{"comment":"The case study sets a 'sub-2 s replies' operating point, but the earlier discussion (Section I) notes that human turn transitions cluster around 200 ms and that a robot has at most a second to respond before disconnect is noticed. Please reconcile this apparent inconsistency or justify why the classroom scenario tolerates a longer response budget.","section":"Section III"},{"comment":"The abstract says Tier 1 'first filters with general metrics,' but the body describes Tier 1 as 'curated static evaluation.' Please align the terminology for consistency.","section":"Abstract"},{"comment":"The connection between the five dimensions and Markelius et al.'s desiderata [11] is asserted rather than explained; a short example of how one desideratum maps to a dimension would help readers verify the claimed 'practitioner-facing projection.'","section":"Section IV"}],"recommendation":"major_revision","confidential_remarks":"The paper is a position paper aimed at a workshop; for a journal, the absence of empirical validation is a concern, but the framework is a useful contribution if the claims are appropriately hedged. The taxonomy relies heavily on the authors' own unpublished work [11]; editors may want to ensure that reference is available or require the dimensions to be justified independently. The major revisions requested above are within the scope of the manuscript and should be addressable without new experiments, primarily by adding a validation plan and tightening the claims."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe paper is a position piece proposing a three-tier evaluation funnel for choosing foundation models for social robots: cheap static benchmarks first, then LLM-simulated interactions, then robot-specific user studies. The genuinely new bit is the assembly: the five-dimension taxonomy (competence, safety, character, scene effectiveness, audience fit), the coverage map that shows which existing benchmarks serve each dimension at each tier, and the Pareto-frontier framing with a latency cutoff. No one has put these pieces together for social robots before. The writing is clear and the citations are current, including the papers that undermine Tier 2's optimism—Zhou et al. on simulation inflation, Panickssery on judge self-preference, Wang on drift. That is honest scholarship.\n\nThe soft spot is exactly where the stress-test lands: the funnel's value proposition is 'cheaper and better informed,' but the 'better informed' half rests entirely on Tier 2's simulated rankings predicting Tier 3 outcomes. The paper flags the risks but provides no calibration and no proof of concept. 'None of this is hypothetical' is too strong if it means the funnel works; it is only true that the pieces can be assembled. A community leaderboard built tomorrow would give you numbers, but not evidence that those numbers correlate with real users on a real robot. That is an empirical question the paper leaves open, explicitly in its own discussion. For a position paper that is acceptable, but the claim should be hedged.\n\nTwo smaller issues. The Pareto frontier is defined per dimension, so the decision rule 'pick your point on the frontier' does not say how to trade off competence against safety or audience fit. A single model that wins on all dimensions at a given cost point is the easy case; the hard case is the multi-objective one. And the case study says the latency budget leaves 'at most a second' to respond, then later uses 'sub-2s replies' as the operating point; a slight inconsistency that should be tightened.\n\nOverall, this is a useful paper for anyone in social robotics who has to pick a model. It deserves a serious referee, and I would encourage a proof-of-concept study as follow-up. The coverage map alone is worth citing.","headline":"Clear, honest position paper that assembles a three-tier evaluation funnel for social-robot foundation models; the 'better informed' half of the claim is unvalidated until Tier 2 simulated rankings are shown to predict real-world outcomes.","tokens_in":7345,"tokens_out":3136,"would_cite":true,"duration_ms":30117,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that foundation-model selection for social robots should use a three-tier evaluation funnel, ending in a Pareto-frontier pick.","keywords":["social robots","foundation models","evaluation funnel","LLM-as-a-judge","Pareto frontier","embodied interaction","community leaderboard","model selection"],"falsifier":"Take a fixed social-robot scene (for instance, Haru's classroom English practice), run the same candidate pool through Tier 2 simulated interactions and through a Tier 3 live user study, and compare the rankings; if the Tier 2 winner is not among the best by Tier 3 in several scenes, the funnel's first filter is misleading.","tokens_in":6406,"feed_emoji":"🤖","tokens_out":6905,"duration_ms":63838,"temperature":0.7,"pith_summary":"Picking a foundation model for a social robot is currently done early and on thin evidence, because public leaderboards test skills that matter for software agents, not embodied social interaction, and because direct user studies are expensive. The paper proposes a three-tiered evaluation funnel that narrows candidate models cheaply before spending scarce participant time: curated static benchmarks first, then simulated interactions with LLM-based users and judges, and finally robot-specific evaluation. It also identifies five dimensions that any such evaluation should cover: conversational competence, user safety, embodied character, target scene effectiveness, and audience appropriateness. If the funnel works, a shared community leaderboard plus platform-specific final-phase harnesses can give every robotics lab defensible model shortlists without each lab running its own embodied comparison from scratch.","feed_headline":"Three tiers of testing pick better robot AI models for less","feed_subtitle":"A shared community leaderboard with cheap simulated chats can replace costly one-off embodied user studies.","key_machinery":"The central object is the evaluation funnel, a three-tier pipeline that narrows a candidate pool as evaluation cost and specificity grow. Tier 1 filters with curated static benchmarks (MMLU, Social IQa, TruthfulQA, XSTest, and similar) run under deployment conditions such as quantization and latency caps; Tier 2 stages simulated multi-turn conversations between each candidate and LLM-based users role-playing the target audience, scored by LLM judges; Tier 3 replays the surviving scenarios through a robot-specific harness and then runs live user studies only when needed. The decision rule is the Pareto frontier: for each evaluation dimension, candidates are plotted by performance against deployment cost (VRAM, latency, parameter count) under a real-time cutoff, and each researcher picks the strongest candidate at their operating point rather than a single global winner.","core_discovery":"The central claim is that the three-tiered evaluation funnel makes foundation-model selection for social robots cheaper and better informed, and that picking a point on the Pareto frontier over performance and deployment cost is the right decision rule. Per-dimension scores from Tier 1 and Tier 2 are robot-agnostic and can be assembled into a living community leaderboard; Tier 3 harnesses, which replay simulated scenarios through a platform's perception and behavior stack, are robot-specific and should be shared for popular platforms. If the claim is right, a shared leaderboard plus Tier 3 harnesses can substitute for per-lab, per-model embodied user studies as the primary comparison method, with live user time reserved for the research question rather than model comparison.","pith_inferences":["Editor's inference: the funnel's value hinges on Tier 2 rankings correlating with Tier 3 outcomes, so a community project to publish correlation scores between LLM-user-judge rankings and live user-study rankings would directly test the proposal.","Editor's inference: the Pareto-frontier decision rule generalizes beyond model choice to prompts, guardrails, and hardware configurations, since each has the same per-dimension, per-cost-point trade-off.","Editor's inference: the Haru classroom case study is a ready falsification experiment; running its Tier 2 simulation and comparing with real high-school student outcomes would show whether the funnel's cheap stages predict embodied social success."],"forward_implications":["Tier 1 and Tier 2 can be built today from existing benchmarks and multi-turn simulation tools, so the shared leaderboard is an immediate possibility, not a hypothetical.","Per-dimension evidence arrives before any hardware is touched, allowing labs to reject a model early without a robot in the loop.","For a deployment like Haru with 30–40 open-weight candidates, the funnel narrows the pool to one or two vetted models before any student participates in a study.","Shared Tier 3 harnesses for popular robot platforms lower the entry cost for new labs, and upstream failure detection spares vulnerable users from exposure to failing models.","Curation and community maintenance of the leaderboard resist benchmark decay better than each lab multiplying its own private suites."],"supporting_citations":[{"why":"Supplies the scenario-and-judge machinery for multi-turn social interaction simulation that Tier 2 builds on.","marker":"[9]"},{"why":"Supplies the community-distributed, crowd-judged head-to-head comparison model that embodied evaluation should adopt.","marker":"[10]"},{"why":"Provides the evaluation desiderata whose situated, relational, and knowledge layers expand the paper's five dimensions.","marker":"[11]"},{"why":"Introduces the holistic-evaluation perspective that motivates reporting performance and cost together as a Pareto frontier.","marker":"[12]"},{"why":"Applies Pareto-optimal trade-offs to language-model serving, grounding the deployment-cost axis of the frontier.","marker":"[13]"},{"why":"Makes Tier 1's curated static evaluation a configuration task rather than new code.","marker":"[23]"},{"why":"Provides the serving infrastructure that lets Tier 1 run under deployment conditions such as quantization and latency caps.","marker":"[24]"},{"why":"Supports the claim that LLM judges can reach human-level agreement on conversational quality, justifying Tier 2's judge-based scoring.","marker":"[27]"},{"why":"Documents that judges favor their own generations, which the paper uses to treat judge scores as priors rather than verdicts.","marker":"[29]"},{"why":"Shows simulation without information asymmetry inflates social competence, a caveat that defines Tier 2's limits.","marker":"[30]"}],"fun_headline_variants":["Three-tier funnel picks robot AI models that cost less","Community leaderboard makes robot model choice cheap and smart","Pick robot brains via funnel: cheap tests first, pricey last","Five dimensions, three tiers: community eval for robot AI","Shared robot model leaderboard: skip expensive user studies"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The funnel's cost savings depend on Tier 2's simulated interactions ranking candidate models the same way real users in the target scene would; if they do not, the cheap tiers can steer scarce Tier 3 resources to the wrong final picks.","fun_headline_variants_meta":{"raw":{"variants":["Three-tier funnel picks robot AI models that cost less","Community leaderboard makes robot model choice cheap and smart","Pick robot brains via funnel: cheap tests first, pricey last","Five dimensions, three tiers: community eval for robot AI","Shared robot model leaderboard: skip expensive user studies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1473,"prompt_tokens":860,"completion_tokens":613,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":476,"completion_tokens_details":{"reasoning_tokens":533}},"tokens_in":476,"tokens_out":613,"duration_ms":6269,"temperature":1.0,"reasoning_tokens":533,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:45:32.974873+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a fixed social-robot scene (for instance, Haru's classroom English practice), run the same candidate pool through Tier 2 simulated interactions and through a Tier 3 live user study, and compare the rankings; if the Tier 2 winner is not among the best by Tier 3 in several scenes, the funnel's first filter is misleading.","supporting_citations":[{"cited_title":"SOTOPIA: Interactive evaluation for social intelligence in language agents,","cited_arxiv_id":null,"evidence_quote":"Supplies the scenario-and-judge machinery for multi-turn social interaction simulation that Tier 2 builds on."},{"cited_title":"Desiderata for foundation models in social robots: Capturing embodied and social aspects for benchmarking,","cited_arxiv_id":null,"evidence_quote":"Provides the evaluation desiderata whose situated, relational, and knowledge layers expand the paper's five dimensions."},{"cited_title":"The language model evaluation harness,","cited_arxiv_id":null,"evidence_quote":"Makes Tier 1's curated static evaluation a configuration task rather than new code."},{"cited_title":"Is this the real life? is this just fantasy? the misleading success of simulating social interactions with LLMs,","cited_arxiv_id":null,"evidence_quote":"Shows simulation without information asymmetry inflates social competence, a caveat that defines Tier 2's limits."}],"review_version":1}