{"id":"7324d9cb-8236-4328-a0ea-005e09479e2c","arxiv_id":"1908.05341","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper proposes that empathy in artificial agents should be evaluated at both the whole-system level and the individual component level, using adapted human empathy questionnaires and behavioral tests.","lead":"This paper reviews how psychologists measure empathy in humans and proposes a two-level framework, system-level and feature-level, for evaluating empathy in artificial interactive agents. It is a position paper that urges the computational empathy field to adopt common evaluation standards rather than a study presenting new data.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"System-level evaluation relies on reworded self-report empathy scales that the paper itself admits are unvalidated; the framework's core metric is an untested hypothesis.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: human self-report empathy scales may not remain valid when reworded into perceived-empathy questionnaires for artificial agents, and Section III-A itself acknowledges the lack of validation. This is the central risk to the paper's proposal. However, the paper is explicitly a position piece aimed at initiating discussion, not a validated methodology. It does not overclaim validated results; it repeatedly frames the recommendations as needing adjustments and validation. Therefore the concern supports the reader's CONDITIONAL verdict: the framework is promising but untested, and the missing validation should be flagged prominently. No additional concern rises to the level of changing the verdict. The citation placeholders and stylistic issues are minor and do not affect the central argument.","tokens_in":11868,"tokens_out":1867,"duration_ms":20592,"concrete_test":"Conduct a validation study in which participants rate a human interaction partner using (a) the original first-person IRI and TEQ, (b) reworded perceived-empathy versions of the same scales, and (c) an established observer-rated measure such as the CARE questionnaire. Compute convergent validity between (b) and (c), and test measurement invariance of the reworded scales' factor structure relative to (a). If the reworded scales diverge substantially from CARE or show a different factor structure, Section III-A's proposed system-level metrics cannot be assumed to measure empathy in agents.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central proposal (Section III) is that empathy in artificial agents should be evaluated at system level via perceived-empathy questionnaires adapted from human self-report scales (IRI, EQ, TEQ), and at feature level via component metrics. The load-bearing assumption is that first-person self-report items remain valid when reworded into second- or third-person judgments about an agent. Section III-A explicitly states: 'Although these methods provide an evaluation that is aligned with the related research on empathy, they were not validated.' This is not a minor caveat: the system-level evaluation is the backbone of the proposed framework, and if the reworded scales do not measure perceived empathy—but instead capture likability, anthropomorphism, or social desirability—then the framework's central metric is invalid. The paper provides no psychometric evidence, no pilot data, and no argument for why the factor structure or construct validity would survive the rewording. The same issue recurs in feature-level evaluation: Section III-B2 notes that emotion regulation metrics 'have not been used by the empathic computing research and may require adjustments.' Thus the proposal is coherent as a roadmap, but its operational core rests on an acknowledged but untested assumption.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a position/review article that argues for a systematic approach to evaluating empathy in artificial agents. It surveys human empathy measures, organizing them by granularity (global vs. component) and by method (self-report, perceived, behavioral), and then maps these onto a two-level framework for artificial agents: system-level evaluation, based on perceived-empathy questionnaires adapted from human self-report scales (IRI, EQ, TEQ) along with user-, context-, and system-related factors, and feature-level evaluation, based on emotional communication, emotion regulation, and cognitive processes. The paper explicitly acknowledges that the proposed instruments are not validated and calls for collective effort to standardize evaluation in computational empathy.","tokens_in":12240,"tokens_out":3791,"duration_ms":35780,"significance":"If the proposed framework were operationalized and validated, it would address a real gap in an emerging field: computational empathy research currently lacks agreed-upon metrics, and this paper provides a useful synthesis of relevant psychology and HCI measures organized into a coherent taxonomy. The paper's strengths are its broad and mostly accurate literature review, its clear separation of system-level and feature-level concerns, and its honest acknowledgment of open problems. Its main significance, however, is as a roadmap rather than a validated method; the central recommendation depends on the untested assumption that first-person self-report empathy scales remain valid when reworded as perceived-empathy questionnaires about an artificial agent.","major_comments":[{"comment":"The system-level evaluation rests on adapting self-report empathy scales (IRI, TEQ, EQ) into perceived-empathy questionnaires, but the paper states on the same page that 'these methods provide an evaluation that is aligned with the related research on empathy, they were not validated.' This is load-bearing: if the reworded scales do not measure perceived empathy but instead capture likability, anthropomorphism, or social desirability, the framework's core metric is invalid. The authors should either provide preliminary psychometric evidence (e.g., internal consistency, convergent/discriminant validity from a pilot study) or explicitly reframe the proposal as a set of hypotheses to be tested rather than a ready-to-use evaluation method.","section":"Section III-A"},{"comment":"The feature-level evaluation is incomplete for two of its three proposed components: for emotion regulation the paper notes that existing metrics 'have not been used by the empathic computing research and may require adjustments,' and for cognitive processes it states 'there are no standardized method to evaluate these capabilities in artificial agents.' These admissions mean that the paper does not actually provide feature-level metrics for a majority of its own framework; it offers only a categorical placeholder. The authors should either supply concrete candidate instruments for these components or clearly mark them as open research problems requiring further development.","section":"Section III-B2 and III-B3"},{"comment":"The introduction promises that 'By providing a checklist of these factors' the paper will initiate a common ground, and Section III repeats that the authors will 'systematically list the factors that contribute to the evaluation of empathy.' However, the paper does not present an actual checklist: it provides taxonomy categories (user/context/system factors; emotional communication, emotion regulation, cognitive processes) and some examples, but no enumerable checklist with operational definitions, scoring, or administration guidance. If the checklist is part of the contribution, it should be included; otherwise the promised deliverable should be revised.","section":"Section I, final paragraph"}],"minor_comments":[{"comment":"There is an unresolved citation placeholder '[ ?]' in the sentence listing capabilities such as mimicry, affective matching, sympathy, and perspective taking.","section":"Section II, paragraph 1"},{"comment":"There is an unresolved citation placeholder '[ ?]' in the reference to the EMOTE project's use of the IRI questionnaire.","section":"Section III-A, paragraph 3"},{"comment":"The word 'sumulation' appears and should be 'simulation'.","section":"Section II-A, QCAE description"},{"comment":"The phrase 'Contrary this dual view' is missing 'to' and should read 'Contrary to this dual view'.","section":"Section II, paragraph 3"},{"comment":"The sentence 'it is useful to focus on a broader view of empathy for to arrive at a comprehensive framework' contains a redundant 'for' and should be revised.","section":"Section II, final paragraph of intro"},{"comment":"References [43] and [49] are the author's own prior work; this is acceptable, but the manuscript should clarify that [43] is an application example and not a validation of the perceived-empathy approach.","section":"References [43] and [49]"}],"recommendation":"major_revision","confidential_remarks":"The paper is better characterized as a position or vision paper than a conventional research article with empirical content. Its value is primarily organizational, and the load-bearing gap between the promised checklist and the actual content, together with the acknowledged lack of validation, makes the manuscript unsuitable for acceptance in its current form. The author should be encouraged to either add pilot data supporting the psychometric portability assumption or substantially modulate the claims to frame the contribution as an open research agenda. The unresolved citation placeholders are particularly problematic for a journal submission and should be fixed in any revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis paper is a clear, well-organized position piece. Its contribution is not new empirical results but a synthesis: it borrows the system-level / feature-level distinction from HCI evaluation and maps it onto psychological models of empathy, producing a checklist for evaluating empathic agents. That combination is genuinely useful for a field that lacks common standards. The review of human empathy measures is accurate and well-cited; I learned a few things about the CARE measure and the critiques of the IRI.\n\nThe soft spots are real but not fatal. The central system-level evaluation relies on rewording first-person self-report scales (IRI, EQ, TEQ) into perceived-empathy questionnaires for agents. The paper admits this is unvalidated (Section III-A), and the stress-test note is right that the construct validity could easily be contaminated by likability or anthropomorphism. That is not a minor caveat; it is the backbone of the proposal. Still, this is a position paper, not a claim of validated measurement. The same issue recurs in feature-level evaluation of emotion regulation, which the paper also concedes needs adjustment.\n\nMinor editorial issues: there are unresolved citation placeholders like '[ ?]' in Sections II and III-A, and a typo ('sumulation'). Those should be fixed before publication.\n\nReading it as a roadmap, the framework is sensible and the author is honest about what is missing. The paper would benefit from a more prominent caveat that the system-level metrics are untested hypotheses, and from a short discussion of what validation would look like (e.g., comparing perceived-empathy scores against behavioral outcomes or expert ratings). But the basic structure holds up as a starting point for discussion.\n\nMy recommendation: this deserves a serious referee. It is unlikely to become a landmark paper, but it is exactly the kind of provocation the field needs to converge on evaluation practices. I would accept it for peer review with the expectation of revision, and I would probably point a student toward it as a sensible taxonomy when designing an empathic-agent study.\n\nBest","headline":"A useful synthesis of empathy evaluation methods that proposes a sensible two-level framework but rests its system-level metric on an unvalidated adaptation of human self-report scales.","tokens_in":12590,"tokens_out":1156,"would_cite":true,"duration_ms":13665,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that empathy in artificial agents should be evaluated at two levels—system-level perceived empathy and feature-level components—with human questionnaires adapted into perceived-empathy surveys.","keywords":["empathy","artificial agents","affective computing","human-computer interaction","system-level evaluation","feature-level evaluation","perceived empathy","evaluation methodology"],"falsifier":"A validation study where participants interact with two agents that differ only in objective empathic functionality—one correctly recognizes and responds to emotional cues, one responds at random—would settle the central claim: if the adapted perceived-empathy questionnaires fail to discriminate the two, or if scores are driven by agent appearance or response speed, the portability assumption fails.","tokens_in":11691,"feed_emoji":"🤖","tokens_out":8789,"duration_ms":80648,"temperature":0.7,"pith_summary":"Computational empathy lacks agreed evaluation metrics, and this paper's aim is to start building them. It proposes that every artificial agent be evaluated twice: as a whole, through questionnaires that ask users or observers how empathic the agent seemed, and at the level of individual features—emotion recognition and expression, emotion regulation, and cognitive processes such as appraisal and perspective-taking. For the system-level questionnaires, it recommends rewording established human self-report empathy scales—the Interpersonal Reactivity Index, the Empathy Quotient, and the Toronto Empathy Questionnaire—from first-person statements into second- or third-person perceived-empathy items. It also catalogs user-related, context-related, and system-related factors such as mood, role expectations, aesthetics, and response fluency that should be measured as controls. The paper acknowledges that the reworded scales have not been validated, and positions the checklist as the starting point for a collective effort rather than a finished metric.","feed_headline":"AI empathy needs a two-level test of whole and parts","feed_subtitle":"Reword human empathy questionnaires into perceived-empathy surveys, then check components like emotion recognition to compare agents fairly.","key_machinery":"The load-bearing mechanism is the two-tier evaluation scheme. At the system level, adapted perceived-empathy questionnaires derived from the Interpersonal Reactivity Index, the Empathy Quotient, and the Toronto Empathy Questionnaire measure the agent's overall empathic impression, while standard HCI metrics control for anthropomorphism, animacy, likeability, perceived intelligence, and perceived safety. At the feature level, the Russian Doll Model—a layered account in which basic emotion-sharing mechanisms support higher cognitive empathy—supplies the categories emotional communication, emotion regulation, and cognitive processes that define what to test in isolation. The split between levels is what converts an abstract construct into a testable checklist.","core_discovery":"The central claim is that the translation of empathy evaluation from humans to artificial agents is tractable if evaluation is split into two tiers. System-level evaluation treats the agent as a whole and measures global perceived empathy, using human empathy questionnaires reworded into perceived-empathy items and HCI control metrics for factors like anthropomorphism and likeability. Feature-level evaluation inspects the components that the Russian Doll Model of Empathy places beneath empathic behavior: emotional communication via recognition and expression, emotion regulation, and cognitive processes such as appraisal, re-appraisal, and perspective-taking. The paper argues these components must be tested separately because errors propagate—an agent with a weak emotion recognizer cannot appraise or respond appropriately—and no single global score can localize the failure. It does not claim there is one universal metric; it claims a systematic checklist of levels, features, and context factors from which application-specific evaluations can be constructed.","pith_inferences":["A direct testable consequence is that system-level and feature-level scores can diverge, and the gap itself would be a diagnostic signal—a likable but functionally shallow agent could score high on perceived empathy while failing component tests.","The checklist could evolve into a shared benchmark if researchers agree on a fixed battery: one perceived-empathy questionnaire plus tests of emotion recognition, expression, regulation, and perspective-taking, with context factors as metadata.","An extension the paper leaves implicit is that regressing perceived-empathy scores on the listed user, context, and system factors would turn the checklist into an empirical model of which variables actually drive perceived empathy.","A natural next experiment would manipulate one system factor such as response fluency while holding empathic functionality fixed, to measure how much of the perceived-empathy score is attributable to aesthetics rather than empathy."],"forward_implications":["Different agents—chatbots, social robots, and virtual assistants—could be compared on a common two-level template instead of ad hoc empathy labels.","Feature-level testing will expose where an agent's empathic chain breaks, such as recognizing an emotion but failing to express an appropriate response.","User, context, and system variables would be reported as standard covariates, making empathy scores interpretable across studies.","The human scales would need validation studies in agent contexts before their system-level scores can be trusted.","Application-specific empathy evaluations can be assembled from the checklist rather than forcing all agents into one metric."],"supporting_citations":[{"why":"This survey establishes computational empathy as a field and suggests adapting the IRI into perceived-empathy questionnaires.","marker":"[8]"},{"why":"This reference supplies the Interpersonal Reactivity Index, the multidimensional self-report scale recommended for adaptation.","marker":"[9]"},{"why":"This reference supplies the Empathy Quotient, a self-report scale treated as one of the most accepted measures and recommended for translation.","marker":"[10]"},{"why":"This reference supplies the Toronto Empathy Questionnaire, whose items were used to build a perceived-empathy version.","marker":"[22]"},{"why":"This reference provides the hierarchical view of affective and cognitive empathy that motivates the feature-level categories.","marker":"[12]"},{"why":"This reference introduces the Russian Doll Model, the layered account of empathy mechanisms used to organize component evaluation.","marker":"[17]"},{"why":"This reference argues for separating system-level from feature-level evaluation in multimodal dialogue systems.","marker":"[37]"},{"why":"This reference argues for component-level and system-level assessment of embodied conversational agents and lists HCI quality factors.","marker":"[38]"},{"why":"This reference supplies standardized HCI control metrics such as anthropomorphism, animacy, likeability, perceived intelligence, and perceived safety.","marker":"[48]"},{"why":"This reference provides the Reading-the-Mind-in-the-Eyes test, a behavioral feature-level measure of emotion recognition that the paper adapts in principle to agents.","marker":"[31]"}],"fun_headline_variants":["Two-tier empathy test: whole agent, then its parts","AI empathy: evaluate holistic score, then component skills","Empathy in agents: check the system, then the features","For AI empathy, test global perception and inner components"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"Everything rests on the assumption that a human empathy questionnaire still measures empathy after being reworded from 'I feel...' to 'The agent seemed...' and answered about an artificial agent—an assumption the paper concedes is not yet validated.","fun_headline_variants_meta":{"raw":{"variants":["Two-tier empathy test: whole agent, then its parts","AI empathy: evaluate holistic score, then component skills","Empathy in agents: check the system, then the features","For AI empathy, test global perception and inner components"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000177,"raw_usage":{"total_tokens":1243,"prompt_tokens":848,"completion_tokens":395,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":464,"completion_tokens_details":{"reasoning_tokens":329}},"tokens_in":464,"tokens_out":395,"duration_ms":4372,"temperature":1.0,"reasoning_tokens":329,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T13:15:29.732646+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A validation study where participants interact with two agents that differ only in objective empathic functionality—one correctly recognizes and responds to emotional cues, one responds at random—would settle the central claim: if the adapted perceived-empathy questionnaires fail to discriminate the two, or if scores are driven by agent appearance or response speed, the portability assumption fails.","supporting_citations":[{"cited_title":"Empa thy in virtual agents and robots: a survey,","cited_arxiv_id":null,"evidence_quote":"This survey establishes computational empathy as a field and suggests adapting the IRI into perceived-empathy questionnaires."},{"cited_title":"Measuring individual differences in empat hy: Evidence for a multidimensional approach","cited_arxiv_id":null,"evidence_quote":"This reference supplies the Interpersonal Reactivity Index, the multidimensional self-report scale recommended for adaptation."},{"cited_title":"The empathy quotie nt: an investi- gation of adults with asperger syndrome or high functioning autism, and normal sex differences,","cited_arxiv_id":null,"evidence_quote":"This reference supplies the Empathy Quotient, a self-report scale treated as one of the most accepted measures and recommended for translation."},{"cited_title":"The toronto empathy questionnaire: Scale development and init ial validation of a factor-analytic solution to multiple empathy measures ,","cited_arxiv_id":null,"evidence_quote":"This reference supplies the Toronto Empathy Questionnaire, whose items were used to build a perceived-empathy version."},{"cited_title":"Mammalian empathy: beha vioural manifestations and neural basis,","cited_arxiv_id":null,"evidence_quote":"This reference provides the hierarchical view of affective and cognitive empathy that motivates the feature-level categories."},{"cited_title":"The russian dollmodel of empathy and imit ation,","cited_arxiv_id":null,"evidence_quote":"This reference introduces the Russian Doll Model, the layered account of empathy mechanisms used to organize component evaluation."},{"cited_title":"Evaluation a nd usability of multimodal spoken language dialogue systems,","cited_arxiv_id":null,"evidence_quote":"This reference argues for separating system-level from feature-level evaluation in multimodal dialogue systems."},{"cited_title":"Evaluating ecas-w hat, how and why?","cited_arxiv_id":null,"evidence_quote":"This reference argues for component-level and system-level assessment of embodied conversational agents and lists HCI quality factors."},{"cited_title":"Measur ement in- struments for the anthropomorphism, animacy, likeability , perceived intelligence, and perceived safety of robots,","cited_arxiv_id":null,"evidence_quote":"This reference supplies standardized HCI control metrics such as anthropomorphism, animacy, likeability, perceived intelligence, and perceived safety."},{"cited_title":"The reading the mind in the eyes test revised version: a study wit h normal adults, and adults with asperger syndrome or high-function ing autism,","cited_arxiv_id":null,"evidence_quote":"This reference provides the Reading-the-Mind-in-the-Eyes test, a behavioral feature-level measure of emotion recognition that the paper adapts in principle to agents."}],"review_version":1}