{"id":"c7c4d3f4-6bcf-4901-aefb-0be6143bb90a","arxiv_id":"2507.14084","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Third-party continuous group affect annotations show no statistically reliable temporal alignment with group memorability annotations in online conversations.","lead":"This paper tests whether emotion labels assigned by outside viewers predict which parts of a conversation a group will remember. Using continuous annotations from online meetings, it finds no reliable temporal link between observed group emotion and group memorability, cautioning against using affect as a proxy for memory in AI systems.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Experiment 3's null may be an artifact of low statistical power or preserved block structure, not evidence of no temporal alignment.","rationale":"The reader's weakest assumption identifies the key concern: the sensitivity of the Experiment 3 permutation procedure. I agree with that identification. The paper's argument is that because none of the metrics reach significance across all three experiments, affect annotations are not reliable proxies for memorability. This is a null result, and the paper does not provide a power analysis, effect size estimate, or confidence interval. The permutation test is also susceptible to a subtle issue: shuffling 15-second blocks preserves each block's affect values, so the null includes autocorrelated traces; this can inflate the null distribution's variance relative to a test that shuffles individual seconds, making the test conservative and prone to false negatives. Moreover, the paper's decision rule requiring all three experiments to be significant is arbitrary and stacks the deck toward the null, since Experiments 1 and 2 test different hypotheses (distributional match) than Experiment 3 (temporal alignment), and the first two experiments already show many significant results. The central claim as stated (cannot be reliably distinguished from random chance) is a claim about the absence of a relationship, and absence-of-evidence is being treated as evidence-of-absence. However, the paper is honest about the limitation that these are group-level third-party annotations, and the permutation framework is appropriate if properly powered. The concern is therefore not that the result is wrong, but that the strength of the conclusion exceeds what the statistical procedure can support without a power analysis or positive control. The verdict CONDITIONAL is appropriate: the paper should either soften the claim or add a power analysis/positive control before the strongest interpretation is accepted. I do not recommend REJECT because the reported null is plausible given the third-party group-level measurement, and the reader's proposed condition (verify sensitivity) is a reasonable bar rather than a demonstration of failure.","tokens_in":18505,"tokens_out":1907,"duration_ms":470274,"concrete_test":"Re-run Experiment 3 with a positive-control dataset: take the real affect traces, then add a known time-localized alignment by boosting arousal/valence within a subset of time windows that overlap with high-memorability intervals (e.g., +2 units on the annotation scale for 10-20% of memorable episodes), and run the exact same 10,000-iteration shuffled-block permutation procedure. If the permutation test fails to reject the null for this synthetic ground-truth signal, the test is underpowered for the effect sizes the paper claims to rule out. Additionally, report the mean and 95% CI of the observed DTW/Euclidean/PATE statistics across bootstrap resamples of the ~30 sessions, and a power curve for Experiment 3 as a function of number of sessions and effect size.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The paper's central claim rests on Experiment 3's all-insignificant result, but the inference from that null to the conclusion is insecure for power reasons. The permutation null shuffles 15-second blocks of three affect dimensions separately per video, and the observed comparison statistics are averaged across sessions. The test's sensitivity depends on (a) the number of sessions (the paper states 35 in Section 5.2 and the Appendix, but says 30 in Section 4.2), (b) the temporal structure of the memory index within sessions, (c) the autocorrelation and block-constant nature of the affect annotations, and (d) the choice of metrics. In particular, DTW and Euclidean distance on memory-vs-affect traces give more weight to the marginal distributions of each trace and to within-block variation; shuffling 15-second blocks preserves within-block affect values and the marginal distribution exactly, so the null distribution's spread may be large relative to the observed statistic if only a few sessions carry any true temporal alignment. With only ~30 sessions, a genuine effect concentrated in a subset of sessions could easily fall within the null envelope. The paper's own decision rule (significant across all three experiments) also ignores that Experiments 1 and 2 primarily test distributional similarity, not temporal alignment; they are not power-analyses for Experiment 3. Therefore the headline statement that the affect-memory relationship 'cannot be reliably distinguished from random chance' is not established: the study lacks a reported effect size, confidence interval, or a power analysis for Experiment 3. A false negative remains a live possibility, so the claim as stated is stronger than the evidence supports.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper empirically tests whether time-continuous third-party group affect annotations (valence, arousal, and intensity) align with group memorability annotations in multi-party conversations, using the MeMo corpus and the affect annotations of Raj Prabhu et al. Three permutation experiments generate synthetic affect data under increasingly restrictive null assumptions: uniform random values, values sampled from each session's observed range, and temporally shuffled 15-second blocks of the real affect annotations. Four metrics (PATE F1, PATE, Euclidean distance, and DTW) compare observed affect-memory alignment against these null distributions. The authors report that no metric is significant across all three experiments and conclude that the observed relationship between affect and memorability annotations cannot be reliably distinguished from random chance, implying that such affect annotations are not reliable proxies for conversational memorability.","tokens_in":18709,"tokens_out":4089,"duration_ms":44351,"significance":"The question is timely and the study fills a real gap: it operationalizes both affect and memorability as time-continuous, group-level constructs in an ecologically valid conversational setting, which is closer to how affective computing systems are actually built than the static, individual-level paradigms in prior work. The permutation-based comparison framework is a principled way to interpret otherwise uninterpretable similarity metrics, and the authors are honest about reporting null results. If the findings withstand scrutiny, they serve as an important caution against assuming that third-party affect annotations capture memory-relevant information. The strength of the contribution, however, depends critically on whether the all-insignificant Experiment 3 is a true absence of temporal alignment or an artifact of low power and an arbitrary decision rule.","major_comments":[{"comment":"The number of sessions is reported inconsistently: §4.2 states the subset consists of '30 conversational sessions' and '12 groups with 42 participants', while §5.2, §6.1.1, and Appendix 1 repeatedly refer to '35 videos' or '35 sessions'. This is load-bearing because the width of the permutation null distribution and the statistical power of the tests depend directly on the number of sessions. The authors should correct the count throughout and state the exact number of sessions, groups, and participants actually used in each experiment.","section":"§4.2 vs. §5.2, §6.1.1, Appendix 1"},{"comment":"The decision rule that the null hypothesis can be rejected only if the p-value is significant across all three experiments is ad hoc and conflates three different null hypotheses (random values, range-matched random values, and temporally shuffled values). Requiring significance in all three is neither a standard overall test nor a test of any single scientific claim, and it makes the conclusion 'cannot be reliably distinguished from random chance' much stronger than the data support. The authors should either justify this conjunction rule from a pre-specified hypothesis or, preferably, designate Experiment 3 as the primary test of temporal alignment and report the corresponding effect sizes and confidence intervals.","section":"§5.2"},{"comment":"The all-insignificant Experiment 3 may be a false negative caused by low statistical power rather than evidence of no temporally aligned affect-memory relationship. Shuffling 15-second blocks preserves the block-constant structure and the marginal distribution of the affect annotations, so the null distribution can be wide; with only about 30 sessions, a genuine effect concentrated in a subset of sessions could easily fall inside the null envelope. The paper does not report a power analysis, effect sizes, or confidence intervals for the observed statistics. The authors should provide these, and should discuss the sensitivity of the conclusion to the choice of 15-second blocks and to the four metrics.","section":"§6.3, Table 1"},{"comment":"The abstract's phrasing that the affect-memory relationship 'cannot be reliably distinguished from what might be expected under random chance' overstates what the analysis shows. Experiments 1 and 2 mostly produced significant differences from uniform and range-matched random data; only the temporal-shuffle null of Experiment 3 was not rejected. A more accurate statement would be that the authors found no evidence of temporal alignment beyond what block-level shuffling of the affect annotations produces. The conclusion in §8 already contains the more careful qualifier 'within the scope of this dataset and methodology', and the abstract should be aligned with that qualifier.","section":"Abstract, §7, §8"}],"minor_comments":[{"comment":"The sentence 'This selection resulted in 3 groups from the original MeMo corpus being exploded and the timestamps...' is unclear; 'exploded' appears to be a typo for 'excluded' or a similar intended word, and the subsequent group count of 12 should be reconciled with the mention of 3 groups.","section":"§4.2"},{"comment":"Section 4.3 states that intensity is binarized at a threshold of 4, which is not the mathematical midpoint of the intensity range, yet §6.1.2 describes the binarization as using 'the middle of Likert scales for each dimension of affect'; this should be clarified to avoid implying a uniform midpoint rule.","section":"§4.3.2 / §6.1.2"},{"comment":"The appendix contains typos such as 'Eucledian distance' in the figure captions and 'continuos' in §8; these should be corrected.","section":"Appendix 2"},{"comment":"The table caption says 'Green cells indicate significant p-values (p<0.004, Bonferroni correction), while uncolored cells are insignificant (p≤0.004)' but the two conditions in the caption are inconsistent; the second should read 'p>0.004'.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a genuinely important question and the null finding would be valuable if it is real. My main concern is that the central conclusion rests on an underpowered and arbitrarily defined decision rule, together with an internal inconsistency in the reported number of sessions. These issues are fixable with additional analysis and careful rewording, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the first test I know of for whether third-party, time-continuous group affect annotations—the kind MER systems actually use—line up with group memorability annotations in conversations, and it runs the right permutation design. But the abstract oversells the result, and the paper's own decision rule makes the null harder to reject than the evidence alone would.\n\nThe novelty is real. Prior emotion-memory work relies on first-person self-report or physiology; here affect is operationalized the way affective computing does it, from observers watching behavior. Using MeMo with the group affect annotations from [22] is a reasonable testbed. Experiment 3, the temporal shuffle, is the critical experiment: shuffling 15-second blocks preserves the marginal distribution of affect and destroys only temporal alignment. That is the correct null for the claim that affect annotations are not reliable proxies.\n\nThe results, however, are more nuanced than the abstract. Experiments 1 and 2 mostly show real affect traces differ from random and range-matched random traces. Experiment 3 shows that once temporal order is destroyed, observed metric values sit near the center of the null distribution. So the evidence is for absence of temporal alignment, not absence of relationship. Saying the link 'cannot be reliably distinguished from random chance' is wrong twice: experiments 1 and 2 do distinguish it from random data, and the relevant null is a shuffle, not chance.\n\nThe main soft spot is the decision rule: rejecting only if all three experiments are significant is arbitrary and mixes different null hypotheses. There is no power analysis or effect size. With roughly 30 sessions and block-constant affect, a genuine effect in a few sessions could be missed. Still, observing metrics at the shuffle mean is weak but real evidence against a strong temporal alignment; it is not just an underpowered null. Fixes: report effect sizes and confidence intervals for Experiment 3, add a sensitivity analysis, and reconcile the session count (35 in Section 5.2 and Appendix vs 30 in Section 4.2; also '3 groups' in 4.2 should be 12). The intensity binarization threshold is a free parameter, but the continuous metrics reduce that concern. Citation pattern is fine; reusing their own dataset and [22]'s annotations is not circular because the null is an empirical outcome.\n\nThis is a useful candidate negative result for affective computing and meeting support. It deserves a serious referee. I would send it out, but ask for the reporting fixes, a tempered abstract, and a clearer justification for the combined decision rule before accepting.","headline":"A genuinely new negative result about third-party group affect annotations and conversational memorability, but the abstract overstates 'random chance' and the all-three-experiments rule is arbitrary; worth refereeing after revision.","tokens_in":19301,"tokens_out":3714,"would_cite":true,"duration_ms":41378,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Observer-rated group emotion shows no reliable link to which conversation moments people remember.","keywords":["conversational memory","group affect","affective computing","memorability annotation","third-party annotation","time-continuous annotation","null hypothesis simulation","meeting support"],"falsifier":"The central claim would be overturned by a temporal-shuffle test in which a real affect-memory metric—for any of arousal, valence, or intensity—falls outside the shuffled null distribution at the Bonferroni-corrected threshold; a natural first check is to re-run Experiment 3 with 1-second or 5-second shuffle windows, since the 15-second block shuffle may preserve within-block structure that could mask a true relationship.","tokens_in":18278,"feed_emoji":"🧠","tokens_out":11112,"duration_ms":111641,"temperature":0.7,"pith_summary":"The paper sets out to test whether emotional annotations produced by outside observers—specifically continuous ratings of a group's pleasure, arousal, and intensity—can serve as proxies for what the group will remember from a conversation. Using recorded multiparty meeting sessions with third-party affect labels and participant-reported memorable moments, the authors compare the affect-memory alignment with three synthetic null baselines, including temporally shuffled affect traces. Since no metric passes all three comparison experiments, they conclude that the observed relationship cannot be reliably distinguished from random chance. If correct, this means systems that use observed group emotion as a stand-in for memorability are building on a shaky assumption.","feed_headline":"Emotion labels don't predict what groups remember","feed_subtitle":"Across three simulations, observer-rated affect matched remembered moments only as well as shuffled data.","key_machinery":"The comparison machinery is a three-part null-hypothesis simulation. For each session, synthetic affect time series are generated under three different null assumptions: uniform random values over the full scale, random values restricted to the range observed in that session, and the real affect sequence shuffled in 15-second blocks, which destroys temporal alignment while keeping the exact distribution. The alignment of each synthetic series with the real memory labels is measured with four metrics: PATE F1 and PATE (proximity-aware evaluation that tolerates small timing shifts), Euclidean distance, and dynamic time warping distance (DTW), which allows stretches and shifts in time. The real data's metric values are then compared against the 10,000-iteration null distributions. The load-bearing rule is that the null is rejected only if the same metric shows significance in all three experiments; because Experiment 3 produced no significant results, the claim of a reliable affect-memory relationship fails.","core_discovery":"The paper's central claim is negative: the relationship between perceived group emotions (valence, arousal, and intensity) and group memorability, measured through continuous time-based annotations, is not reliable enough to distinguish from chance. The evidence is three simulation experiments. Experiments 1 and 2, which generate random affect data with minimal assumptions or with the observed range of values, mostly show that the real affect-memory alignment is unlikely to be random. Experiment 3, which shuffles the actual affect annotations in 15-second windows and thereby preserves every property of the data except temporal alignment with memory, produces no significant metric for any affect dimension. Under the paper's pre-specified decision rule—reject the null only if a metric is significant across all three experiments—the null hypothesis cannot be rejected. The authors therefore conclude that the significant effects in the first two experiments were artifacts of distributional differences, not evidence of affect tracking memory, and that third-party group affect annotations are not dependable proxies for conversational memorability.","pith_inferences":["This null result may be specific to the group-level, third-party operationalization: individual-level first-person affect annotations could still predict individual memorability, and the paper itself flags this as future work; if so, the failure is about aggregation and annotation perspective, not about the emotion-memory link.","Because the affect labels were collected in 15-second blocks and Experiment 3 shuffles whole blocks, any genuine relationship at sub-15-second timescales would be invisible to this analysis; a finer-grained shuffle test could change the outcome.","The group memorability index aggregates individual recall reports, which may dilute the signal: if only one participant remembers a moment, the group index treats it as memorable for the whole group, adding noise that could weaken an existing affect-memory association.","The paper's rejection rule—requiring significance in all three experiments—is conservative; a study designed around the temporal-shuffle null alone might be more decisive for the temporal-alignment question."],"forward_implications":["Meeting support, summarization, and memory-augmentation systems should stop treating observed group emotion as a stand-in for what users will remember; direct memorability signals such as recall-based annotations are needed.","Emotion-recognition pipelines that use third-party continuous affect annotations as relevance labels should re-validate their ground truth against first-person or self-report measures.","The well-documented emotion-memory link from cognitive science does not automatically transfer to the annotation practices used in affect-recognition technology—third-party, continuous, group-level labels capture something different from experienced emotion.","Time-continuous affect features may still be useful for other purposes, but using them as a proxy for long-term event relevance in conversational AI cannot be justified by this relationship.","The non-significant temporal-shuffle experiment implies that what made Experiments 1 and 2 look significant was the distribution of affect values, not their alignment in time with memorable moments."],"supporting_citations":[{"why":"Supplies the conversational memory corpus: session recordings, participant free-recall memorable moments, and the group-level memory annotations used as the outcome.","marker":"[27]"},{"why":"Provides the third-party time-continuous group affect annotations (arousal and valence) on the same sessions that this paper tests against memorability.","marker":"[22]"},{"why":"Defines the PATE and PATE F1 proximity-aware time series metrics used to measure affect-memory alignment with tolerance for temporal shifts.","marker":"[52]"},{"why":"Defines dynamic time warping, the metric used to measure similarity between affect and memory traces when timing may be stretched or displaced.","marker":"[55]"},{"why":"The circumplex model of affect underlies the pleasure-arousal annotation scheme the affect labels are based on.","marker":"[48]"},{"why":"The component model of emotion motivates the distinction between experienced emotion and observer-rated emotion that frames the central research question.","marker":"[19]"},{"why":"Establishes from behavioural science that emotional arousal modulates memory storage, the link that would justify using affect as a proxy.","marker":"[6]"},{"why":"Identifies relevance appraisal as a driver of memory encoding, another pillar of the proxy assumption being evaluated.","marker":"[12]"}],"fun_headline_variants":["Group emotions don't flag memorable moments","Affect annotations fail memory test","Observer emotions don't track recall","Emotion-memory link is statistical noise","Third-party emotions miss memorable bits"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that shuffling affect annotations in 15-second blocks, together with the PATE, Euclidean, and DTW metrics, is powerful enough to expose a genuine temporally aligned affect-memory relationship; if the shuffle keeps too much structure or the metrics are too weak, the all-insignificant Experiment 3 could be a false negative.","fun_headline_variants_meta":{"raw":{"variants":["Group emotions don't flag memorable moments","Affect annotations fail memory test","Observer emotions don't track recall","Emotion-memory link is statistical noise","Third-party emotions miss memorable bits"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000173,"raw_usage":{"total_tokens":1309,"prompt_tokens":1004,"completion_tokens":305,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":620,"completion_tokens_details":{"reasoning_tokens":246}},"tokens_in":620,"tokens_out":305,"duration_ms":4216,"temperature":1.0,"reasoning_tokens":246,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T16:01:25.794021+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"The central claim would be overturned by a temporal-shuffle test in which a real affect-memory metric—for any of arousal, valence, or intensity—falls outside the shuffled null distribution at the Bonferroni-corrected threshold; a natural first check is to re-run Experiment 3 with 1-second or 5-second shuffle windows, since the 15-second block shuffle may preserve within-block structure that could mask a true relationship.","supporting_citations":[{"cited_title":"Introducing MeMo: A Multimodal Dataset for Memory Modelling in Multiparty Conversations","cited_arxiv_id":"2409.13715","evidence_quote":"Supplies the conversational memory corpus: session recordings, participant free-recall memorable moments, and the group-level memory annotations used as the outcome."},{"cited_title":"Dynamics of Collective Group Affect: Group-level Annotations and the Multimodal Modeling of Convergence and Divergence","cited_arxiv_id":"2409.08578","evidence_quote":"Provides the third-party time-continuous group affect annotations (arousal and valence) on the same sessions that this paper tests against memorability."},{"cited_title":"Pate: Proximity- aware time series anomaly evaluation,","cited_arxiv_id":null,"evidence_quote":"Defines the PATE and PATE F1 proximity-aware time series metrics used to measure affect-memory alignment with tolerance for temporal shifts."},{"cited_title":"A circumplex model of affect,","cited_arxiv_id":null,"evidence_quote":"The circumplex model of affect underlies the pleasure-arousal annotation scheme the affect labels are based on."},{"cited_title":"Emotion as a multicomponent process: A model and some cross-cultural data,","cited_arxiv_id":null,"evidence_quote":"The component model of emotion motivates the distinction between experienced emotion and observer-rated emotion that frames the central research question."},{"cited_title":"Modulation of memory storage,","cited_arxiv_id":null,"evidence_quote":"Establishes from behavioural science that emotional arousal modulates memory storage, the link that would justify using affect as a proxy."},{"cited_title":"Dopamine and adaptive memory,","cited_arxiv_id":null,"evidence_quote":"Identifies relevance appraisal as a driver of memory encoding, another pillar of the proxy assumption being evaluated."}],"review_version":1}