{"id":"b165fe55-9fe6-4333-80fb-5b410f640e21","arxiv_id":"2506.04132","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":2,"one_line_summary":"A ResNet50 trained on mel-spectrograms of basil plant voltages reportedly reaches 97% accuracy in a seven-emotion classification, yet the same table shows near-zero recall for three emotions and a macro F1 of 0.56.","lead":"The paper reports that a deep learning model can read human emotions from plant voltage signals with 97% accuracy, but the confusion matrix shows the model never recognizes three of the seven emotions. The real performance is a macro F1 of 0.56, driven mostly by the overrepresented 'neutral' and 'sad' classes.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 97% emotion-classification claim is not protected against activity confounds: induction sessions leave participant and plant together, and the confusion matrix shows three emotions are never predicted, so the headline number does not establish seven-way emotion decoding.","rationale":"The reader's weakest assumption is the same point I identify: without physical isolation or blinding, the plant signal may encode activity rather than emotion. I agree, and I add that the confusion matrix makes the '97%' internally misleading, while the shuffled-label control is not a valid null because it sits below the majority-class baseline. These are correctness risks, not merely disagreements with consensus. The paper does have some positive features: open-source code is provided, preprocessing choices are specified, and prior related work exists. But the decisive experiment for the central claim—clean separation of emotion from correlated human activity—is not reported here. The paper's own §6 lists mechanism and replication limits but does not flag the confound that matters most. Since the reader's REJECT verdict is consistent with this analysis, no verdict adjustment is needed.","tokens_in":6592,"tokens_out":3584,"duration_ms":36040,"concrete_test":"Preregister and run an isolation replication of §3.1: induce each of the seven emotions while the participant is in a separate, electrically shielded room (or inside a Faraday cage) with no visual, acoustic, or vibrational contact to the plant; train the same ResNet50 pipeline on this isolated dataset and report a session-independent (leave-one-session-out) test accuracy. Add a matched-motor control in which participants perform the same facial, speech, and body routines for each label without feeling the emotion. If accuracy in the isolated condition falls to chance or is matched by the motor-only control, the original 97% reflects activity artifacts, not emotion.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim depends on §3.1–§3.3 and §4.1: participants undergo emotion induction while a plant electrode records voltage, and a ResNet50 then reports 97% accuracy. For this result to mean plants encode emotional state, two conditions must hold. First, the labels must be caused by internal emotion rather than by correlated motor, acoustic, or electrical activity. The methods describe face-api.js video validation of emotion, but no Faraday cage, physical separation, audio masking, or blinding is reported for the induction sessions in this study. The §4.4 'artifact control' is a citation to previous work, not data shown here. Human speech, facial movement, breathing, and body shifts generate electrostatic and low-frequency electrical signals that a 0.1–50 Hz bioelectric amplifier can readily record from nearby electrodes; with the participant in the same room as the plant, the classifier may simply be detecting 'emotion-induction activity.' Second, the reported 97% is not a seven-way accuracy in any meaningful sense: Table 1 assigns zero precision/recall to Disgusted, Fearful, and Surprised, and macro F1 is 0.56. The number is dominated by Neutral (462/898) and Sad (261/898). The shuffled-label control in Table 2 is also not a clean null: its 30% accuracy lies below the 51.4% majority-class baseline, indicating the control model is not the same decision rule, so the '97% versus 30%' contrast overstates significance. If condition (1) is false, the classification result, and with it the central early-warning hypothesis, loses its evidential base.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a five-year research program claiming that plants generate bioelectric signals encoding human proximity, identity, movement, voice, stress, and emotion. The central new result is a ResNet50 classifier applied to mel-spectrograms of basil voltage recordings that allegedly achieves 97% accuracy in distinguishing seven human emotional states, with a shuffled-label control reported at 30% accuracy. The authors interpret these results as evidence for an evolved anti-herbivory early warning system and propose applications in agriculture, healthcare, and human-plant interaction. The methods describe an ESP32/INA128 sensor, 20-second spectrogram windows, and transfer learning with ResNet50, with code and data links provided.","tokens_in":6906,"tokens_out":4914,"duration_ms":47370,"significance":"If the 97% emotion-classification result were valid and causally attributable to the participant's internal emotional state, it would be a remarkable finding with implications for plant sensory biology and human-plant interaction. The paper deserves credit for releasing open-source code, for documenting a multi-year research program, and for attempting machine-learning-based analysis of plant bioelectric signals. However, as presented, the central claim is not supported: the confusion matrix shows that three of seven emotion classes are never predicted, macro F1 is 0.56, the shuffled-label control is not a valid chance baseline, and the experimental design does not exclude the most salient confound of concurrent human activity. The significance of the claimed phenomenon is therefore not established by this manuscript.","major_comments":[{"comment":"The headline 97% accuracy does not represent seven-way emotion classification. In Table 1, Disgusted, Fearful, and Surprised all have precision, recall, and F1 of 0.00, and the macro F1 is 0.56. The two majority classes, Neutral (462) and Sad (261), account for 723 of 898 test samples, so a trivial classifier that always predicts these two classes would achieve 80.5% accuracy. The 97% figure is therefore dominated by majority classes, and the abstract's claim of 'classifying human emotional states' is misleading without reporting per-class metrics and a majority-class baseline.","section":"§4.1, Table 1"},{"comment":"The experimental design does not control for the most plausible confound: the participant and the plant are in the same room during emotion induction, and the electrode chain is sensitive to low-frequency electrical activity in the 0.1-50 Hz band. No Faraday cage, physical separation, audio masking, or blinded electrode setup is reported for this study; the artifact controls cited in §4.4 are attributed to 'our previous artifact elimination work' and no data are shown here. Under these conditions, the classifier may be detecting speech, facial movement, breathing, body shifts, or electrode drift correlated with the emotion-induction procedure rather than with the participant's internal emotional state. This is a load-bearing internal-validity gap, not a presentation issue.","section":"§3.1, §4.4"},{"comment":"The shuffled-label control is not a valid null baseline as interpreted. With the reported class frequencies, random guessing according to empirical priors gives an expected accuracy of about 36.5%, and always predicting the majority class gives 51.4%. The shuffled model's 30% accuracy is below both, indicating that the shuffled training run did not implement the same decision rule as the valid model. The contrast '97% versus 30%' therefore does not establish that plant voltage spectrograms contain emotion information; a proper permutation control using the identical pipeline and decision rule is required.","section":"§4.1, Table 2"},{"comment":"The manuscript asserts 'geographic replication confirmed in laboratories across Europe and North America,' 'seasonal consistency,' and 'rigorous artifact elimination' but provides no data, protocols, or effect sizes for these claims in this paper. Since the Discussion and Conclusions rely on these validations, the unsupported claims cannot be accepted as evidence. In addition, the Limitations section (6.1) acknowledges that the mechanism is unknown but does not acknowledge the confound-control or class-imbalance problems that directly bear on the main claim, making the limitations statement incomplete.","section":"§4.4, §6.1"}],"minor_comments":[{"comment":"The Methods do not report the number of participants, number of sessions, demographic information, or emotion-induction procedure details; these are needed for reproducibility and for assessing the generalizability of the classifier.","section":"§3.1"},{"comment":"The 80/20 train/test split is described as stratified, but it is not stated whether the split was performed at the level of individual windows, sessions, or participants. Because segmentation uses 20-second windows with 10-second overlap, a window-level split could leak information from the same session into both training and test sets; this should be clarified.","section":"§3.3"},{"comment":"The text states that the primary discriminative signal is in the 0.5-10 Hz range, while the preprocessing section reports a 0.1-50 Hz bandpass filter; the paper should clarify whether classification used the full filtered band or a narrower sub-band.","section":"§4.3"},{"comment":"The caption states that the system 'achieved 97% confidence in detecting happiness,' which is a face-api.js confidence score for facial expression, not the plant classifier's confidence; the wording should distinguish these two very different quantities.","section":"Figure 3"},{"comment":"The Codariocalyx motorius acoustic-sensing study is cited as a supporting evidence source, but it is a Master's thesis; the text should describe its methods and controls rather than relying solely on the citation.","section":"§5.2, Reference [14]"},{"comment":"There are several typographical issues, including 'Phase2from2020-2022' in §1.1, 'T raining Protocol' and 'V alidation Strategy' in §3.3, and inconsistent capitalization of 'Codariocalyx motorius' versus 'Codariocalyx Motorius' in §5.2 and reference [14].","section":"Throughout"}],"recommendation":"reject","confidential_remarks":"The manuscript is essentially a synthesis of the authors' own research program, with the new classification analysis extending reference [11] and no independent replication. Combined with the class-imbalance problem in Table 1 and the invalid shuffled-label baseline in Table 2, the central claim cannot be accepted in its current form. A revision would require new data collection with physical isolation of the participant from the plant, a proper permutation baseline, and per-class reporting; this is beyond the scope of a standard revision and would amount to a new study. I also note that several strong validation claims (Faraday-cage controls, geographic replication) are asserted without supporting data in this manuscript, which creates an unusually high burden of proof for the extraordinary conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know before reading: the abstract says 97% accuracy on seven emotions, but Table 1 shows three of the seven classes are never predicted and macro F1 is 0.56. The number is just accuracy on the majority classes (Neutral and Sad are 81% of the test set). That alone sinks the central claim as stated.\n\nWhat is genuinely new here is the evolutionary wrapper, the \"early warning hypothesis\" for anti-herbivory detection. The empirical core, however, is not new: the ResNet50 emotion classification from plant spectrograms is essentially a repeat of the authors' own 2024 Sensors paper (ref [11]). This manuscript adds context and a narrative, not new measurements.\n\nCredit where it is due: the methods section is concrete, the code is on GitHub, and the limitation section openly says the mechanism is unknown. The species-response hierarchy (basil responding more than defended species) is at least a coherent, testable pattern. Those are real features.\n\nThe soft spots are load-bearing. First, no blinding or physical separation is reported for the induction sessions here; the participant and plant are in the same room, so the model may be detecting body movement, speech, or electrode drift rather than emotion. The Faraday-cage artifact control is only cited from prior work, not shown. Second, the shuffled-label control hits 30%, which is below the 51% majority-class baseline; that means the control model isn't the same decision rule, so \"97% versus 30%\" is not a clean significance contrast. Third, three emotion classes have zero samples correctly classified, so the \"seven-way decoding\" language is simply inaccurate.\n\nWho is this for? Anyone tracking the plant-electrophysiology fringe, or interested in how ML can produce optimistic numbers from small, unbalanced datasets. It is a useful teaching case, but not a reliable scientific claim. A serious editor could send it to review because the topic is high-impact-if-true and the code is available for auditors, but my own verdict is that it fails on its own reported numbers. I would not cite it for any substantive claim.","headline":"The paper's own Table 1 disproves its headline 97% seven-way emotion decoding, and the rest is a re-analysis of the authors' earlier work with a speculative wrapper.","tokens_in":7459,"tokens_out":1347,"would_cite":false,"duration_ms":15289,"reading_group":"maybe","serious_thinker":"no","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A five-year study claims basil plants generate voltage signals that let a deep-learning classifier read a nearby human's emotional state at 97% accuracy, with shuffled-label controls at 30%.","keywords":["plant electrophysiology","bioelectric signaling","emotion recognition","machine learning","human-plant interaction","early warning systems","Ocimum basilicum","mel-spectrogram classification"],"falsifier":"A decisive control would place the participant in an electrically shielded, sound-isolated room with no sightline, no shared air, and no audio or video link to the plant while the same electrode rig records. If classification accuracy stays near 97%, emotion-specific biophysical coupling is supported; if accuracy falls to chance, the classifier was reading concurrent human activity rather than emotional state.","tokens_in":6350,"feed_emoji":"🌿","tokens_out":5455,"duration_ms":54458,"temperature":0.7,"pith_summary":"The paper claims that basil plants produce bioelectric voltage patterns that carry decodable information about a nearby human's emotional state, and that this reflects an evolved early-warning system against herbivores. Using electrode recordings transcribed into mel-spectrograms, a transfer-learning neural network reportedly reached 97% accuracy on seven emotion categories, while shuffled labels yielded 30% accuracy, which the authors take as evidence the signal is genuine. The paper's own per-class table shows that the 97% figure is driven by the three dominant classes (neutral, sad, happy); the rare classes (disgusted, fearful, surprised) are never recognized, and the macro F1 is 0.56. The authors integrate five years of related experiments—individual recognition, gesture detection, voice response, stress prediction, and sleep staging—into the hypothesis that high-palatability plant species evolved pre-contact detection of approaching animals.","feed_headline":"Plant voltage signals reportedly read human emotions at 97%","feed_subtitle":"Deep learning on basil-plant voltage spectrograms claims a live link to human emotion—if it survives controls.","key_machinery":"The machinery is a transfer-learning classifier on mel-spectrogram images of plant voltage. A differential amplifier (ESP32 with INA128) records leaf-to-soil voltage at 400 Hz; the signal is bandpass filtered at 0.1–50 Hz, segmented into 20-second windows, and converted into 64-bin mel-spectrograms. A ResNet50 network pretrained on a large general image corpus reuses its visual feature extractors to classify these spectrogram images into seven emotion classes. The mel-spectrogram is the pivotal object: it turns a one-dimensional voltage trace into a two-dimensional image whose time-frequency structure is what the network learns to associate with emotions.","core_discovery":"The central claim is that Ocimum basilicum plants generate bioelectric potential fluctuations that differ systematically with the emotional state of a nearby human, and that these differences are large enough for a machine-learning classifier to read them. The paper reports 97% overall accuracy on a held-out set of 898 spectrogram windows, versus 30% for shuffled labels, and interprets this contrast as evidence that plant voltage patterns contain genuine information about internal human state. It generalizes this into an early-warning hypothesis: plants under strong herbivore pressure evolved the ability to detect approaching animals through bioelectric or other pre-contact cues, allowing defense mobilization before damage occurs. The reported accuracy is not uniform across the seven classes: neutral, sad, and happy dominate the test set and are classified well, while disgusted, fearful, and surprised are never correctly classified, producing a macro F1 of 0.56.","pith_inferences":["Editorial inference: The reported dataset cannot rule out that the network is detecting correlated physical activity rather than felt emotion; a blinded, physically separated protocol would separate the two.","Editorial inference: Taking the per-class results at face value, the practical signal is a coarse calm-versus-aroused or pleasant-versus-unpleasant axis, not fine-grained seven-way emotion recognition.","Editorial inference: The evolutionary story predicts a testable gradient: strongly defended plant species should show weak or absent responses and should fail the same classifier.","Editorial inference: The 'electromagnetic' framing is one of several possible channels; voice, breathing, or temperature cues could produce the same data, so targeted occlusion experiments (soundproofing, filtered airflow, blocked sightlines) could identify the actual mechanism."],"forward_implications":["If the 97% result replicates, plant bioelectric monitoring could offer a non-invasive, continuous window onto human emotional state that works at a distance.","The early-warning hypothesis predicts that high-palatability species (basil, lettuce, tomato) respond most strongly while defended species (zucchini, corn, orchids) respond weakly; the paper reports exactly this species pattern.","The same pipeline is claimed to extend to stress prediction above 90% accuracy and to sleep staging, pointing toward health-monitoring applications.","The paper releases its full code pipeline as open-source software, allowing independent labs to test the method on new data.","The per-class results imply that real deployments would need many more examples of rare emotions before a seven-way emotion classifier becomes trustworthy."],"supporting_citations":[{"why":"Earlier phase of the same program: plants distinguish individuals at 66% accuracy and happy vs sad at 85%, which this paper extends to seven emotions.","marker":"[13]"},{"why":"The most direct precedent for deep-learning emotion recognition from plant sensitivity; the ResNet50 spectrogram pipeline builds on it.","marker":"[11]"},{"why":"Eurythmic-gesture detection at 74.9% accuracy with machine learning, supporting plant sensitivity to human movement.","marker":"[3]"},{"why":"Initial eurythmic-dancing-with-plants experiments that the paper treats as Phase 2 evidence for movement responses.","marker":"[4]"},{"why":"AI tracking of eurythmic human-plant interaction from the same research line, cited for gesture recognition results.","marker":"[10]"},{"why":"Tomato plants reacting to human voices, cited as evidence for non-linguistic detection mechanisms.","marker":"[6]"},{"why":"Stress and exam-performance prediction above 90% accuracy via basil potentials, cited as physiological-state evidence.","marker":"[15]"},{"why":"Telegraph-plant acoustic discrimination, cited to extend the early-warning hypothesis from human emotion to broader sensory modalities.","marker":"[14]"}],"fun_headline_variants":["Plants read human emotions via bioelectric signals, study claims","97% accuracy: basil plants classify human emotions from voltage","Plant bioelectric signals may be early warnings of approaching animals","Deep learning on basil plants reads human emotion, but only for a few feelings","Plant voltage spectrograms decode human emotions with 97% accuracy"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The plant-voltage differences during emotion sessions are caused by the participant's internal emotional state rather than by correlated body movement, breathing, voice, temperature changes, or electrode drift.","fun_headline_variants_meta":{"raw":{"variants":["Plants read human emotions via bioelectric signals, study claims","97% accuracy: basil plants classify human emotions from voltage","Plant bioelectric signals may be early warnings of approaching animals","Deep learning on basil plants reads human emotion, but only for a few feelings","Plant voltage spectrograms decode human emotions with 97% accuracy"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000793,"raw_usage":{"total_tokens":3465,"prompt_tokens":890,"completion_tokens":2575,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":506,"completion_tokens_details":{"reasoning_tokens":2488}},"tokens_in":506,"tokens_out":2575,"duration_ms":19215,"temperature":1.0,"reasoning_tokens":2488,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T10:46:21.251854+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive control would place the participant in an electrically shielded, sound-isolated room with no sightline, no shared air, and no audio or video link to the plant while the same electrode rig records. If classification accuracy stays near 97%, emotion-specific biophysical coupling is supported; if accuracy falls to chance, the classifier was reading concurrent human activity rather than emotional state.","supporting_citations":[{"cited_title":"Recognizing Individuals and Their Emotions Using Plants as Bio-Sensors through Electro-static Discharge","cited_arxiv_id":"2005.04591","evidence_quote":"Earlier phase of the same program: plants distinguish individuals at 66% accuracy and happy vs sad at 85%, which this paper extends to seven emotions."},{"cited_title":"A., Ciechanowski, L., Dupuis, A., Vazquez, I., & Gloor, P","cited_arxiv_id":null,"evidence_quote":"The most direct precedent for deep-learning emotion recognition from plant sensitivity; the ResNet50 spectrogram pipeline builds on it."},{"cited_title":"A., & Weinbeer, M","cited_arxiv_id":null,"evidence_quote":"Eurythmic-gesture detection at 74.9% accuracy with machine learning, supporting plant sensitivity to human movement."},{"cited_title":"Eurythmic Dancing with Plants -- Measuring Plant Response to Human Body Movement in an Anthroposophic Environment","cited_arxiv_id":"2012.12978","evidence_quote":"Initial eurythmic-dancing-with-plants experiments that the paper treats as Phase 2 evidence for movement responses."},{"cited_title":"F., Weinbeer, M., & Gloor, P","cited_arxiv_id":null,"evidence_quote":"AI tracking of eurythmic human-plant interaction from the same research line, cited for gesture recognition results."},{"cited_title":"I., & Gloor, P","cited_arxiv_id":null,"evidence_quote":"Tomato plants reacting to human voices, cited as evidence for non-linguistic detection mechanisms."},{"cited_title":"O., & Gloor, P","cited_arxiv_id":null,"evidence_quote":"Stress and exam-performance prediction above 90% accuracy via basil potentials, cited as physiological-state evidence."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Telegraph-plant acoustic discrimination, cited to extend the early-warning hypothesis from human emotion to broader sensory modalities."}],"review_version":1}