REVIEW 3 major objections 5 minor 26 references
Active Electrosensing and Communication in MARL-trained Weakly Electric Fish Collectives
T0 review · 3 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read Training foraging agents with only individual-fitness rewards reproduces electric signaling patterns of weakly electric fish, including heavy-tailed pulse intervals, context-dependent modulation, and freeloading.
desk verdict A promising MARL testbed for electric-fish electrocommunication where the EOD statistics genuinely emerge, but the 'no social rewards' claim is undercut by dominance-dependent penalties in the reward, and the fish comparisons need stats and code. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing mechanism is a recurrent neural network controller trained by multi-agent proximal policy optimization, operating in a 2D physics simulator that models electric field generation and propagation. Each agent receives three egocentric sensor channels: distortions of its own EODs (short-range active sensing), low-frequency background fields (passive sensing), and sharp EOD pulses from conspecifics (long-range social sensing). The same EOD emission serves as both an active-sensing probe and a communication signal; rewards are individual foraging success with asymmetric penalties for aggression between different dominance levels. This setup lets EOD modulation and social strategie
What would settle it
Retrain the two-fish foraging assay with a reward function that has no dominance-dependent asymmetry (symmetric aggression penalties or none at all) and compare dominance-related foraging success; if the dominance effect disappears, the emergence claim for social communication is undercut, while if it persists, the claim is supported.
Extended reading notes
Core claim
The central claim is that evolution-inspired individual-fitness rewards, combined with biophysically motivated electrosensory input, are sufficient to produce communication-like collective behavior. Specifically, the paper reports that MARL-trained agents reproduce heavy-tailed EOD interval statistics of real fish; shift their EOD rates with environmental context (higher rates under competition, lower rates when collective sensing is available); exhibit freeloading; and show dominance-dependent social foraging in a minimal two-fish assay. The paper also demonstrates that these behaviors depend causally on the electrosensory channels (e.g., ablating collective sensing reduces freeloading) and
Load-bearing premise
The paper's claim that social behaviors like dominance asymmetries emerge from individual fitness alone is weakened because the reward function already contains asymmetric penalties for aggression between differently ranked fish, so the two-fish assay may reflect the experimenter-defined dominance parameter rather than spontaneously evolved social communication.
Editorial extensions
If this is right
- If correct, the framework offers a way to generate concrete, testable predictions about EOD signaling in weakly electric fish, including which signal features are functionally relevant.
- It suggests that heavy-tailed EOD interval distributions and social 'freeloading' do not require dedicated social reward circuitry but can arise from individual foraging efficiency in a shared environment.
- Complete access to RNN dynamics enables causal intervention studies—silencing EODs, ablating sensor channels—to identify what drives social foraging, which is difficult in live animals.
- The synthetic communication corpora could be aligned with real recordings using unsupervised translation methods, offering a path to decode EOD 'meaning' without multi-brain recordings.
Reading between the lines
- The paper's two-fish assay may conflate emergent social communication with experimenter-defined dominance, since the reward function already contains asymmetric aggression penalties; retraining with symmetric penalties would clarify whether dominance-dependent foraging is truly emergent.
- The collective-sensing result suggests a general design principle for distributed sensing in multi-agent systems: eavesdropping on neighbors' active sensing can reduce individual energy expenditure, which might generalize beyond electric fish to any active-sensing collective (e.g., sonar or lidar swarms).
- If agent EOD statistics can be matched to real fish under a specified ecology, the same training pipeline could be used inversely to infer ecological pressures from field EOD recordings.
- The encoding analyses hint that social context is represented in recurrent dynamics, which could support targeted 'steering' of agent behavior—essentially programmable electro-communication for testing hypotheses about signal meaning.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces a multi-agent reinforcement learning (MARL) framework in which recurrent-neural-network agents, modeled on weakly electric fish, forage in a 2D physics simulator with biomimetic electrosensory input, EOD emission, and social sensing. The central claims are that, from individual-fitness rewards and no explicit social reward, agents reproduce heavy-tailed EOD interval statistics, context-dependent EOD rate shifts, 'freeloading' under collective sensing, and dominance-dependent social foraging asymmetries; additional ablations and RNN analyses are presented as causal and mechanistic support. The paper argues that communication-like collective behavior can emerge without explicit social incentives.
Significance. If the emergence claim survives the reward-confound concern, this is a valuable contribution: it provides an interpretable, intervention-friendly testbed for weakly electric fish communication, with full access to RNN internal dynamics and the ability to generate synthetic communication corpora. The biomimetic electrosensory grounding and the minimal two-fish assay are strengths, and the heavy-tailed EOD statistics and freeloading are not directly optimized, so those particular results are non-trivial. However, the claim that dominance-related social behavior is emergent is currently unproven because the reward function includes an explicit dominance-dependent social penalty. The paper also needs substantially more statistical rigor before 'consistent with real fish' can be accepted.
major comments (3)
- [§2 Methods; Abstract] The central emergence claim — 'no explicit collective behaviors are rewarded' and 'rather than through rewarding agents explicitly for social interactions' — is contradicted by the reward description in §2: rewards 'provide asymmetric penalties during aggressive encounters between fish of different dominance levels.' This is an explicit social interaction term coupling agents through a predefined dominance hierarchy. The two-fish assay (Fig. 3) then varies relative dominance and reports dominance-dependent foraging success (Fig. 3c,d) as an emergent social phenomenon. As stated, the result may reflect the experimenter-imposed asymmetry rather than emergent communication. Please train with a symmetric or absent aggression penalty and re-run the Fig. 3 analysis, or treat dominance as a learned/emergent variable rather than a reward input.
- [§3, Figs. 1d and 2a] The claim that trained agents reproduce real-fish hallmarks is supported only by qualitative distributional comparisons. Fig. 1d overlays SPI distributions without error bars, number of seeds, or statistical tests; Fig. 2a reports EOD probabilities as point estimates without confidence intervals. Heavy-tailed distributions can arise from many trivial stochastic processes, so a quantitative comparison (e.g., KS distances, tail-exponent estimates) and variance across training seeds, including null/ablated controls, are necessary. Without this, the headline 'consistent with real fish collectives' claim is not supported.
- [§2 Methods, Reproducibility] The Methods omit key training details needed to evaluate the robustness of the results: reward coefficients for foraging vs. aggression penalty, dominance-level assignment, PPO hyperparameters, number of independent training seeds, number of episodes, arena dimensions, and food replenishment parameters. Since the results include several conditional comparisons (competitive vs. non-competitive, with/without Knollenorgan), these details are necessary to determine whether differences are robust and reproducible. Please provide a complete specification and, ideally, release code and data.
minor comments (5)
- [Abstract] Typos: 'likeGnathonemus petersii' should be 'like Gnathonemus petersii'; 'collective behavior, Experimental' needs a period and capitalization.
- [Fig. 2a caption] The caption says 'Left to Right' but the panels are arranged in a 2x2 grid. Please refer to panels as (a1)–(a4) consistently and describe the layout explicitly.
- [Abstract/§3] The abstract mentions 'EOD silencing' as an intervention, but no EOD-silencing ablation appears in the main figures or analyses. Either add the result or remove the mention.
- [Abstract; §2] The phrase 'evolution-inspired rewards' in the abstract is inconsistent with the Methods, which use PPO (a policy-gradient method, not evolutionary search). Please reconcile the terminology.
- [General] The manuscript does not include a code or data availability statement. For a computational modeling paper, please state whether code and trained agents will be released.
Circularity Check
Dominance-dependent foraging is pre-specified by the reward function, so the claim that social behaviors emerge without explicit social rewards is only partially supported.
-
self definitional
[Section 2 (Methods, reward specification) and Fig. 3 caption]
"Agents are trained using Multi-Agent Proximal Policy Optimization [13, 14, 15] with rewards that encourage successful foraging and provide asymmetric penalties during aggressive encounters between fish of different dominance levels. ... Importantly, no explicit collective behaviors are rewarded, coordination and communication emerge solely from individual fitness optimization in a shared environment."
The asymmetric penalty is an explicit social/interaction reward parametrized by dominance. The two-fish assay then 'varies the relative dominance levels' and reports that 'B performs better when it is more dominant.' That directional result is installed in the reward, not discovered; the dominance-dependent component of foraging success is an input called an emergent finding. The heavy-tailed EOD statistics and freeloading are not directly rewarded, so the circularity is partial.
full rationale
Most quantitative predictions are not fitted: heavy-tailed SPI distributions, context-dependent EOD shifts, and freeloading emerge from an individual foraging reward and are compared with external real-fish data (Ref. [4]), so those results are self-contained and non-circular. The exception is dominance: the Methods reward includes 'asymmetric penalties during aggressive encounters between fish of different dominance levels,' and Fig. 3 then reports dominance-dependent foraging success as a finding. Because relative dominance is a reward parameter, the dominance result reduces to the reward design and undercuts the abstract claim that behaviors emerge 'rather than through rewarding agents explicitly for social interactions.' This is a partial, not total, circularity: the paper's main EOD-statistics claims remain independent.
Assumptions & free parameters
free parameters (5)
- Dominance level assignment / asymmetric aggression penalty =
not reported
- Reward weights for foraging vs. aggression penalty =
not reported
- Electrosensory sensor parameters (e.g., Knollenorgan range, receptor gains) =
not reported
- Electric field propagation and reflection constants =
not reported
- Food patch replenishment rates / arena dimensions =
not reported
assumptions (4)
- domain assumption The custom 2D electric-field simulator faithfully represents behaviorally relevant electrosensory physics.
- domain assumption Reward for individual foraging plus asymmetric aggression penalties approximates evolutionary fitness.
- domain assumption PPO with recurrent networks converges to representative, not degenerate, strategies.
- domain assumption Heavy-tailed SPI distribution shape is a sufficient hallmark for biological correspondence.
Cite this review
Pith. "Pith review of Active Electrosensing and Communication in MARL-trained Weakly Electric Fish Collectives." pith.science (2026). https://pith.science/paper/TBEISVFY
@misc{pith2026251108436,
author = {Pith},
title = {Pith review of: Active Electrosensing and Communication in MARL-trained Weakly Electric Fish Collectives},
year = {2026},
howpublished = {\url{https://pith.science/paper/TBEISVFY}},
note = {Machine review of arXiv:2511.08436}
}
read the original abstract
How complex collective behavior emerges from individual interactions is a fundamental scientific question, but experimental cost and difficulty of simultaneous multi-brain recordings limit direct study in animals. Here we introduce a novel computational framework modeling weakly electric fish-like agents with biophysically inspired electrosensing and actuation, trained to forage collectively via multi-agent reinforcement learning (MARL). Trained agents reproduce hallmarks of real fish, including curvilinear homing trajectories and heavy-tailed electric organ discharge (EOD) interval statistics, while exhibiting emergent active sensing, social foraging, dominance-like asymmetries, and aggression. We perform in silico interventions including sensor ablations, EOD silencing, and food distribution changes to identify causal drivers of social foraging. Analyses of recurrent neural dynamics further show robust encoding of task-relevant variables and social context. Our work has broad implications for the neuroethology of weakly electric fish and other social animals where extensive multi-individual neural recordings, and thus traditional data-driven modeling, remain challenging.
Figures
Reference graph
Works this paper leans on
-
[1]
Electrolocation of capacitive objects in four species of pulse-type weakly electric fish: Ii
Gerhard von der Emde. Electrolocation of capacitive objects in four species of pulse-type weakly electric fish: Ii. electric signalling behaviour.Ethology, 92(3):177–192, 1992
1992
-
[2]
Active electrolocation of objects in weakly electric fish.Journal of experimental biology, 202(10):1205–1215, 1999
Gerhard V on der Emde. Active electrolocation of objects in weakly electric fish.Journal of experimental biology, 202(10):1205–1215, 1999
1999
-
[3]
Avner Wallach and Nathaniel B. Sawtell. An internal model for canceling self-generated sensory input in freely behaving electric fish.Neuron, 111(16):2570–2582.e5, August 2023
2023
-
[4]
Federico Pedraja and Nathaniel B. Sawtell. Collective sensing in electric fish.Nature, 628(8006):139–144, April 2024
2024
-
[5]
A conceptual modeling of flocking-regulated multi-agent reinforcement learning
CS Chen, Yaqing Hou, and Yew-Soon Ong. A conceptual modeling of flocking-regulated multi-agent reinforcement learning. In2016 International Joint Conference on Neural Networks (IJCNN), pages 5256–5262. IEEE, 2016
2016
-
[6]
Multiagent planning and control for swarm herding in 2-d obstacle environments under bounded inputs.IEEE Transactions on Robotics, 37(6):1956–1972, 2021
Vishnu S Chipade and Dimitra Panagou. Multiagent planning and control for swarm herding in 2-d obstacle environments under bounded inputs.IEEE Transactions on Robotics, 37(6):1956–1972, 2021
1956
-
[7]
Collaborative hunting in artificial agents with deep reinforcement learning.Elife, 13:e85694, 2024
Kazushi Tsutsui, Ryoya Tanaka, Kazuya Takeda, and Keisuke Fujii. Collaborative hunting in artificial agents with deep reinforcement learning.Elife, 13:e85694, 2024
2024
-
[8]
C. C. Bell, C. D. Hopkins, K. Grant, and T. Natoli. Contributions of electrosensory systems to neurobiology and neuroethology: Proceedings of a conference in honor of the scientific career of Thomas Szabo.Journal of Comparative Physiology A, 173(6):657–763, December 1993
1993
Show all 26 references
-
[9]
The animal translators.The New York Times, Aug 2022
Emily Anthes. The animal translators.The New York Times, Aug 2022
2022
-
[10]
Bronstein, Roee Diamant, Denley Delaney, Shane Gero, Shafi Goldwasser, David F
Jacob Andreas, Gašper Beguš, Michael M. Bronstein, Roee Diamant, Denley Delaney, Shane Gero, Shafi Goldwasser, David F. Gruber, Sarah de Haas, Peter Malkin, Nikolay Pavlov, Roger Payne, Giovanni Petri, Daniela Rus, Pratyusha Sharma, Dan Tchernov, Pernille Tønnesen, Antonio Tor...
2022
-
[11]
The young person’s guide to the theil index: Suggesting intuitive interpretations and exploring analytical applications
Pedro Conceição and Pedro Ferreira. The young person’s guide to the theil index: Suggesting intuitive interpretations and exploring analytical applications. 2000
2000
-
[12]
House, Rudiger Krahe, and Mark E
Ling Chen, Jonathan L. House, Rudiger Krahe, and Mark E. Nelson. Modeling signal and background components of electrosensory scenes.Journal of Comparative Physiology A, 191(4):331–345, April 2005
2005
-
[13]
Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347, 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347, 2017
2017 arXiv
-
[14]
Recurrent model-free rl is a strong baseline for many POMDPs.arXiv preprint arXiv:2110.05038, 2021
Tianwei Ni, Benjamin Eysenbach, and Ruslan Salakhutdinov. Recurrent model-free rl is a strong baseline for many POMDPs.arXiv preprint arXiv:2110.05038, 2021
2021 arXiv
-
[15]
The surprising effectiveness of PPO in cooperative multi-agent games.Advances in Neural Information Processing Systems, 35:24611–24624, 2022
Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. The surprising effectiveness of PPO in cooperative multi-agent games.Advances in Neural Information Processing Systems, 35:24611–24624, 2022
2022
-
[16]
Electric organ discharge patterns during group hunting by a mormyrid fish.Proceedings of the Royal Society B: Biological Sciences, 272(1570):1305–1314, July 2005
Matthew E Arnegard and Bruce A Carlson. Electric organ discharge patterns during group hunting by a mormyrid fish.Proceedings of the Royal Society B: Biological Sciences, 272(1570):1305–1314, July 2005
2005
-
[17]
Carlson and Carl D
Bruce A. Carlson and Carl D. Hopkins. Stereotyped temporal patterns in electrical communication.Animal Behaviour, 68(4):867–878, October 2004. 5
2004
-
[18]
Proposal: Deciphering electrocommunication with marl and unsupervised machine translation
Satpreet Harcharan Singh, Sonja Johnson-Yu, Zhouyang Lu, Aaron Walsman, Federico Pedraja, Denis Turcu, Pratyusha Sharma, Naomi Saphra, Nathaniel Sawtell, and Kanaka Rajan. Proposal: Deciphering electrocommunication with marl and unsupervised machine translation. InThe Thirty-N...
2025
-
[19]
A theory of unsupervised translation motivated by understanding animal communication.Advances in Neural Information Processing Systems, 36:37286–37320, 2023
Shafi Goldwasser, David Gruber, Adam Tauman Kalai, and Orr Paradise. A theory of unsupervised translation motivated by understanding animal communication.Advances in Neural Information Processing Systems, 36:37286–37320, 2023
2023
-
[20]
Unsupervised translation of emergent communication
Ido Levy, Orr Paradise, Boaz Carmeli, Ron Meir, Shafi Goldwasser, and Yonatan Belinkov. Unsupervised translation of emergent communication. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 23231–23239, 2025
2025
-
[21]
Dissecting larval zebrafish hunting using deep reinforcement learning trained rnn agents.arXiv preprint arXiv:2510.03699, 2025
Raaghav Malik, Satpreet H Singh, Sonja Johnson-Yu, Nathan Wu, Roy Harpaz, Florian Engert, and Kanaka Rajan. Dissecting larval zebrafish hunting using deep reinforcement learning trained rnn agents.arXiv preprint arXiv:2510.03699, 2025
2025 arXiv
-
[22]
Inputdsa: Demixing then comparing recurrent and externally driven dynamics.arXiv preprint arXiv:2510.25943, 2025
Ann Huang, Mitchell Ostrow, Satpreet H Singh, Leo Kozachkov, Ila Fiete, and Kanaka Rajan. Inputdsa: Demixing then comparing recurrent and externally driven dynamics.arXiv preprint arXiv:2510.25943, 2025
2025
-
[23]
Learning dynamics and the geometry of neural dynamics in recurrent neural controllers
Ann Huang, Satpreet Harcharan Singh, and Kanaka Rajan. Learning dynamics and the geometry of neural dynamics in recurrent neural controllers. InWorkshop on Interpretable Policies in Reinforcement Learning RLC-2024, 2024
2024
-
[24]
Measuring and controlling solution degeneracy across task-trained recurrent neural networks.ArXiv, pages arXiv–2410, 2025
Ann Huang, Satpreet H Singh, Flavio Martinelli, and Kanaka Rajan. Measuring and controlling solution degeneracy across task-trained recurrent neural networks.ArXiv, pages arXiv–2410, 2025
2025
-
[25]
Keep it real: rethinking the primacy of experimental control in cognitive neuroscience.NeuroImage, 222:117254, 2020
Samuel A Nastase, Ariel Goldstein, and Uri Hasson. Keep it real: rethinking the primacy of experimental control in cognitive neuroscience.NeuroImage, 222:117254, 2020
2020
-
[26]
Keypoint annotation for electrocommu- nication source separation with pikachu and raichu
Kaden Zheng, Sonja Johnson-Yu, Satpreet Harcharan Singh, Denis Turcu, Federico Pedraja, Pratyusha Sharma, Naomi Saphra, Nathaniel Sawtell, and Kanaka Rajan. Keypoint annotation for electrocommu- nication source separation with pikachu and raichu. InThe Thirty-Ninth Annual Conf...
2025
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.