Pith. sign in

REVIEW 3 major objections 5 minor 26 references

Active Electrosensing and Communication in MARL-trained Weakly Electric Fish Collectives

T0 review · 3 major / 5 minor · reviewed 2026-08-03 · deepseek-v4-flash

Pith's one-line read Training foraging agents with only individual-fitness rewards reproduces electric signaling patterns of weakly electric fish, including heavy-tailed pulse intervals, context-dependent modulation, and freeloading.

desk verdict A promising MARL testbed for electric-fish electrocommunication where the EOD statistics genuinely emerge, but the 'no social rewards' claim is undercut by dominance-dependent penalties in the reward, and the fish comparisons need stats and code. read the letter →

arxiv 2511.08436 v2 pith:TBEISVFY submitted 2025-11-11 cs.NE cs.AIcs.MAcs.SYeess.SYq-bio.NC

classification cs.NEcs.AIcs.MAcs.SYeess.SYq-bio.NC
keywords weaklyelectricfishorgandischarge(EOD)multi-agentreinforcementlearningcollectivesensingsocialforagingemergentcommunicationactiveelectrosensingdominance
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper asks whether complex collective electro-communication can arise purely from individual fitness optimization. It trains recurrent-neural-network agents in a physics simulator with biomimetic electrosensing, rewarding only successful foraging (plus asymmetric aggression penalties), and finds that the agents spontaneously develop electric organ discharge patterns matching real fish: heavy-tailed interval distributions, shifts with food competition, and 'freeloading' where one agent reduces its own discharges while benefiting from neighbors' active sensing. A minimal two-fish assay shows that access to a conspecific's EODs and relative dominance jointly shape foraging success, and analyses of the RNN internals show robust encoding of task-relevant variables and social context. The value is a controllable in silico testbed for neuroethology, where multi-animal neural recordings are difficult.

What carries the argument

The load-bearing mechanism is a recurrent neural network controller trained by multi-agent proximal policy optimization, operating in a 2D physics simulator that models electric field generation and propagation. Each agent receives three egocentric sensor channels: distortions of its own EODs (short-range active sensing), low-frequency background fields (passive sensing), and sharp EOD pulses from conspecifics (long-range social sensing). The same EOD emission serves as both an active-sensing probe and a communication signal; rewards are individual foraging success with asymmetric penalties for aggression between different dominance levels. This setup lets EOD modulation and social strategie

What would settle it

Retrain the two-fish foraging assay with a reward function that has no dominance-dependent asymmetry (symmetric aggression penalties or none at all) and compare dominance-related foraging success; if the dominance effect disappears, the emergence claim for social communication is undercut, while if it persists, the claim is supported.

Watch

Extended reading notes

Core claim

The central claim is that evolution-inspired individual-fitness rewards, combined with biophysically motivated electrosensory input, are sufficient to produce communication-like collective behavior. Specifically, the paper reports that MARL-trained agents reproduce heavy-tailed EOD interval statistics of real fish; shift their EOD rates with environmental context (higher rates under competition, lower rates when collective sensing is available); exhibit freeloading; and show dominance-dependent social foraging in a minimal two-fish assay. The paper also demonstrates that these behaviors depend causally on the electrosensory channels (e.g., ablating collective sensing reduces freeloading) and

Load-bearing premise

The paper's claim that social behaviors like dominance asymmetries emerge from individual fitness alone is weakened because the reward function already contains asymmetric penalties for aggression between differently ranked fish, so the two-fish assay may reflect the experimenter-defined dominance parameter rather than spontaneously evolved social communication.

Editorial extensions

If this is right

  • If correct, the framework offers a way to generate concrete, testable predictions about EOD signaling in weakly electric fish, including which signal features are functionally relevant.
  • It suggests that heavy-tailed EOD interval distributions and social 'freeloading' do not require dedicated social reward circuitry but can arise from individual foraging efficiency in a shared environment.
  • Complete access to RNN dynamics enables causal intervention studies—silencing EODs, ablating sensor channels—to identify what drives social foraging, which is difficult in live animals.
  • The synthetic communication corpora could be aligned with real recordings using unsupervised translation methods, offering a path to decode EOD 'meaning' without multi-brain recordings.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's two-fish assay may conflate emergent social communication with experimenter-defined dominance, since the reward function already contains asymmetric aggression penalties; retraining with symmetric penalties would clarify whether dominance-dependent foraging is truly emergent.
  • The collective-sensing result suggests a general design principle for distributed sensing in multi-agent systems: eavesdropping on neighbors' active sensing can reduce individual energy expenditure, which might generalize beyond electric fish to any active-sensing collective (e.g., sonar or lidar swarms).
  • If agent EOD statistics can be matched to real fish under a specified ecology, the same training pipeline could be used inversely to infer ecological pressures from field EOD recordings.
  • The encoding analyses hint that social context is represented in recurrent dynamics, which could support targeted 'steering' of agent behavior—essentially programmable electro-communication for testing hypotheses about signal meaning.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces a multi-agent reinforcement learning (MARL) framework in which recurrent-neural-network agents, modeled on weakly electric fish, forage in a 2D physics simulator with biomimetic electrosensory input, EOD emission, and social sensing. The central claims are that, from individual-fitness rewards and no explicit social reward, agents reproduce heavy-tailed EOD interval statistics, context-dependent EOD rate shifts, 'freeloading' under collective sensing, and dominance-dependent social foraging asymmetries; additional ablations and RNN analyses are presented as causal and mechanistic support. The paper argues that communication-like collective behavior can emerge without explicit social incentives.

Significance. If the emergence claim survives the reward-confound concern, this is a valuable contribution: it provides an interpretable, intervention-friendly testbed for weakly electric fish communication, with full access to RNN internal dynamics and the ability to generate synthetic communication corpora. The biomimetic electrosensory grounding and the minimal two-fish assay are strengths, and the heavy-tailed EOD statistics and freeloading are not directly optimized, so those particular results are non-trivial. However, the claim that dominance-related social behavior is emergent is currently unproven because the reward function includes an explicit dominance-dependent social penalty. The paper also needs substantially more statistical rigor before 'consistent with real fish' can be accepted.

major comments (3)
  1. [§2 Methods; Abstract] The central emergence claim — 'no explicit collective behaviors are rewarded' and 'rather than through rewarding agents explicitly for social interactions' — is contradicted by the reward description in §2: rewards 'provide asymmetric penalties during aggressive encounters between fish of different dominance levels.' This is an explicit social interaction term coupling agents through a predefined dominance hierarchy. The two-fish assay (Fig. 3) then varies relative dominance and reports dominance-dependent foraging success (Fig. 3c,d) as an emergent social phenomenon. As stated, the result may reflect the experimenter-imposed asymmetry rather than emergent communication. Please train with a symmetric or absent aggression penalty and re-run the Fig. 3 analysis, or treat dominance as a learned/emergent variable rather than a reward input.
  2. [§3, Figs. 1d and 2a] The claim that trained agents reproduce real-fish hallmarks is supported only by qualitative distributional comparisons. Fig. 1d overlays SPI distributions without error bars, number of seeds, or statistical tests; Fig. 2a reports EOD probabilities as point estimates without confidence intervals. Heavy-tailed distributions can arise from many trivial stochastic processes, so a quantitative comparison (e.g., KS distances, tail-exponent estimates) and variance across training seeds, including null/ablated controls, are necessary. Without this, the headline 'consistent with real fish collectives' claim is not supported.
  3. [§2 Methods, Reproducibility] The Methods omit key training details needed to evaluate the robustness of the results: reward coefficients for foraging vs. aggression penalty, dominance-level assignment, PPO hyperparameters, number of independent training seeds, number of episodes, arena dimensions, and food replenishment parameters. Since the results include several conditional comparisons (competitive vs. non-competitive, with/without Knollenorgan), these details are necessary to determine whether differences are robust and reproducible. Please provide a complete specification and, ideally, release code and data.
minor comments (5)
  1. [Abstract] Typos: 'likeGnathonemus petersii' should be 'like Gnathonemus petersii'; 'collective behavior, Experimental' needs a period and capitalization.
  2. [Fig. 2a caption] The caption says 'Left to Right' but the panels are arranged in a 2x2 grid. Please refer to panels as (a1)–(a4) consistently and describe the layout explicitly.
  3. [Abstract/§3] The abstract mentions 'EOD silencing' as an intervention, but no EOD-silencing ablation appears in the main figures or analyses. Either add the result or remove the mention.
  4. [Abstract; §2] The phrase 'evolution-inspired rewards' in the abstract is inconsistent with the Methods, which use PPO (a policy-gradient method, not evolutionary search). Please reconcile the terminology.
  5. [General] The manuscript does not include a code or data availability statement. For a computational modeling paper, please state whether code and trained agents will be released.

Circularity Check

1 steps flagged · score 6.0 of 10

Dominance-dependent foraging is pre-specified by the reward function, so the claim that social behaviors emerge without explicit social rewards is only partially supported.

  1. self definitional [Section 2 (Methods, reward specification) and Fig. 3 caption]
    "Agents are trained using Multi-Agent Proximal Policy Optimization [13, 14, 15] with rewards that encourage successful foraging and provide asymmetric penalties during aggressive encounters between fish of different dominance levels. ... Importantly, no explicit collective behaviors are rewarded, coordination and communication emerge solely from individual fitness optimization in a shared environment."

    The asymmetric penalty is an explicit social/interaction reward parametrized by dominance. The two-fish assay then 'varies the relative dominance levels' and reports that 'B performs better when it is more dominant.' That directional result is installed in the reward, not discovered; the dominance-dependent component of foraging success is an input called an emergent finding. The heavy-tailed EOD statistics and freeloading are not directly rewarded, so the circularity is partial.

full rationale

Most quantitative predictions are not fitted: heavy-tailed SPI distributions, context-dependent EOD shifts, and freeloading emerge from an individual foraging reward and are compared with external real-fish data (Ref. [4]), so those results are self-contained and non-circular. The exception is dominance: the Methods reward includes 'asymmetric penalties during aggressive encounters between fish of different dominance levels,' and Fig. 3 then reports dominance-dependent foraging success as a finding. Because relative dominance is a reward parameter, the dominance result reduces to the reward design and undercuts the abstract claim that behaviors emerge 'rather than through rewarding agents explicitly for social interactions.' This is a partial, not total, circularity: the paper's main EOD-statistics claims remain independent.

Assumptions & free parameters 5 free parameters · 4 assumptions · 0 invented entities

No new physical entities are postulated; the simulated agents are computational objects, not hypothesized biological mechanisms. The central claims rest on multiple unreported quantitative parameters and on domain assumptions about the fidelity of the simulator and rewards.

free parameters (5)
  • Dominance level assignment / asymmetric aggression penalty = not reported
    Relative dominance is a manipulated variable in the reward and in the two-fish assay; values not specified.
  • Reward weights for foraging vs. aggression penalty = not reported
    Shapes all emergent behavior; no values or sweep reported.
  • Electrosensory sensor parameters (e.g., Knollenorgan range, receptor gains) = not reported
    Controls collective sensing and freeloading; cited as biomimetic but quantitative values absent.
  • Electric field propagation and reflection constants = not reported
    Underlies active sensing; model details not given.
  • Food patch replenishment rates / arena dimensions = not reported
    Competitive vs non-competitive conditions depend on these values.
assumptions (4)
  • domain assumption The custom 2D electric-field simulator faithfully represents behaviorally relevant electrosensory physics.
    All results depend on this; no validation against measured fields shown (Section 2).
  • domain assumption Reward for individual foraging plus asymmetric aggression penalties approximates evolutionary fitness.
    Used to justify emergence claims; not derived from biological data (Section 2).
  • domain assumption PPO with recurrent networks converges to representative, not degenerate, strategies.
    Assumed by use of MARL; no multiple-seed analysis shown (Section 2).
  • domain assumption Heavy-tailed SPI distribution shape is a sufficient hallmark for biological correspondence.
    Fig. 1d compares shapes only; no statistical test (Section 3).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Active Electrosensing and Communication in MARL-trained Weakly Electric Fish Collectives." pith.science (2026). https://pith.science/paper/TBEISVFY

@misc{pith2026251108436,
  author       = {Pith},
  title        = {Pith review of: Active Electrosensing and Communication in MARL-trained Weakly Electric Fish Collectives},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/TBEISVFY}},
  note         = {Machine review of arXiv:2511.08436}
}
read the original abstract

How complex collective behavior emerges from individual interactions is a fundamental scientific question, but experimental cost and difficulty of simultaneous multi-brain recordings limit direct study in animals. Here we introduce a novel computational framework modeling weakly electric fish-like agents with biophysically inspired electrosensing and actuation, trained to forage collectively via multi-agent reinforcement learning (MARL). Trained agents reproduce hallmarks of real fish, including curvilinear homing trajectories and heavy-tailed electric organ discharge (EOD) interval statistics, while exhibiting emergent active sensing, social foraging, dominance-like asymmetries, and aggression. We perform in silico interventions including sensor ablations, EOD silencing, and food distribution changes to identify causal drivers of social foraging. Analyses of recurrent neural dynamics further show robust encoding of task-relevant variables and social context. Our work has broad implications for the neuroethology of weakly electric fish and other social animals where extensive multi-individual neural recordings, and thus traditional data-driven modeling, remain challenging.

Figures

Figures reproduced from arXiv: 2511.08436 by the authors.

Figure 1
Figure 1. Overview of our MARL framework for modeling weakly electric fish communication. [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. (a) Comparison of EOD probabilities under various conditions: (Left to Right) (a1) Effect of the Knollenorgan in competitive environments. The presence of the Knollenorgan (which provides long-range information about other agents) increases EOD rates in competitive scenarios only, suggesting the importance of social information in limited-resource regimes. (a2) Effect of the Knollenorgan in non-competitive environme… view at source ↗
Figure 3
Figure 3. (a) Minimal social foraging assay with two agents, A and B. A is initialized within a fully-replenishing food patch, while B is randomly initialized within communication radius to A. (b) Example trajectories in different A/B relative dominance scenarios. (c) We vary the relative dominance levels of A/B, then compare the percentage of trials where B reaches the patch (100 runs). B performs better when it is more domi… view at source ↗

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

26 extracted references · 3 linked inside Pith

  1. [1]

    Electrolocation of capacitive objects in four species of pulse-type weakly electric fish: Ii

    Gerhard von der Emde. Electrolocation of capacitive objects in four species of pulse-type weakly electric fish: Ii. electric signalling behaviour.Ethology, 92(3):177–192, 1992

  2. [2]

    Active electrolocation of objects in weakly electric fish.Journal of experimental biology, 202(10):1205–1215, 1999

    Gerhard V on der Emde. Active electrolocation of objects in weakly electric fish.Journal of experimental biology, 202(10):1205–1215, 1999

  3. [3]

    Avner Wallach and Nathaniel B. Sawtell. An internal model for canceling self-generated sensory input in freely behaving electric fish.Neuron, 111(16):2570–2582.e5, August 2023

  4. [4]

    Federico Pedraja and Nathaniel B. Sawtell. Collective sensing in electric fish.Nature, 628(8006):139–144, April 2024

  5. [5]

    A conceptual modeling of flocking-regulated multi-agent reinforcement learning

    CS Chen, Yaqing Hou, and Yew-Soon Ong. A conceptual modeling of flocking-regulated multi-agent reinforcement learning. In2016 International Joint Conference on Neural Networks (IJCNN), pages 5256–5262. IEEE, 2016

  6. [6]

    Multiagent planning and control for swarm herding in 2-d obstacle environments under bounded inputs.IEEE Transactions on Robotics, 37(6):1956–1972, 2021

    Vishnu S Chipade and Dimitra Panagou. Multiagent planning and control for swarm herding in 2-d obstacle environments under bounded inputs.IEEE Transactions on Robotics, 37(6):1956–1972, 2021

  7. [7]

    Collaborative hunting in artificial agents with deep reinforcement learning.Elife, 13:e85694, 2024

    Kazushi Tsutsui, Ryoya Tanaka, Kazuya Takeda, and Keisuke Fujii. Collaborative hunting in artificial agents with deep reinforcement learning.Elife, 13:e85694, 2024

  8. [8]

    C. C. Bell, C. D. Hopkins, K. Grant, and T. Natoli. Contributions of electrosensory systems to neurobiology and neuroethology: Proceedings of a conference in honor of the scientific career of Thomas Szabo.Journal of Comparative Physiology A, 173(6):657–763, December 1993

Show all 26 references
  1. [9]

    The animal translators.The New York Times, Aug 2022

    Emily Anthes. The animal translators.The New York Times, Aug 2022

  2. [10]

    Bronstein, Roee Diamant, Denley Delaney, Shane Gero, Shafi Goldwasser, David F

    Jacob Andreas, Gašper Beguš, Michael M. Bronstein, Roee Diamant, Denley Delaney, Shane Gero, Shafi Goldwasser, David F. Gruber, Sarah de Haas, Peter Malkin, Nikolay Pavlov, Roger Payne, Giovanni Petri, Daniela Rus, Pratyusha Sharma, Dan Tchernov, Pernille Tønnesen, Antonio Tor...

  3. [11]

    The young person’s guide to the theil index: Suggesting intuitive interpretations and exploring analytical applications

    Pedro Conceição and Pedro Ferreira. The young person’s guide to the theil index: Suggesting intuitive interpretations and exploring analytical applications. 2000

  4. [12]

    House, Rudiger Krahe, and Mark E

    Ling Chen, Jonathan L. House, Rudiger Krahe, and Mark E. Nelson. Modeling signal and background components of electrosensory scenes.Journal of Comparative Physiology A, 191(4):331–345, April 2005

  5. [13]

    Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347, 2017

    John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347, 2017

  6. [14]

    Recurrent model-free rl is a strong baseline for many POMDPs.arXiv preprint arXiv:2110.05038, 2021

    Tianwei Ni, Benjamin Eysenbach, and Ruslan Salakhutdinov. Recurrent model-free rl is a strong baseline for many POMDPs.arXiv preprint arXiv:2110.05038, 2021

  7. [15]

    The surprising effectiveness of PPO in cooperative multi-agent games.Advances in Neural Information Processing Systems, 35:24611–24624, 2022

    Chao Yu, Akash Velu, Eugene Vinitsky, Jiaxuan Gao, Yu Wang, Alexandre Bayen, and Yi Wu. The surprising effectiveness of PPO in cooperative multi-agent games.Advances in Neural Information Processing Systems, 35:24611–24624, 2022

  8. [16]

    Electric organ discharge patterns during group hunting by a mormyrid fish.Proceedings of the Royal Society B: Biological Sciences, 272(1570):1305–1314, July 2005

    Matthew E Arnegard and Bruce A Carlson. Electric organ discharge patterns during group hunting by a mormyrid fish.Proceedings of the Royal Society B: Biological Sciences, 272(1570):1305–1314, July 2005

  9. [17]

    Carlson and Carl D

    Bruce A. Carlson and Carl D. Hopkins. Stereotyped temporal patterns in electrical communication.Animal Behaviour, 68(4):867–878, October 2004. 5

  10. [18]

    Proposal: Deciphering electrocommunication with marl and unsupervised machine translation

    Satpreet Harcharan Singh, Sonja Johnson-Yu, Zhouyang Lu, Aaron Walsman, Federico Pedraja, Denis Turcu, Pratyusha Sharma, Naomi Saphra, Nathaniel Sawtell, and Kanaka Rajan. Proposal: Deciphering electrocommunication with marl and unsupervised machine translation. InThe Thirty-N...

  11. [19]

    A theory of unsupervised translation motivated by understanding animal communication.Advances in Neural Information Processing Systems, 36:37286–37320, 2023

    Shafi Goldwasser, David Gruber, Adam Tauman Kalai, and Orr Paradise. A theory of unsupervised translation motivated by understanding animal communication.Advances in Neural Information Processing Systems, 36:37286–37320, 2023

  12. [20]

    Unsupervised translation of emergent communication

    Ido Levy, Orr Paradise, Boaz Carmeli, Ron Meir, Shafi Goldwasser, and Yonatan Belinkov. Unsupervised translation of emergent communication. InProceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 23231–23239, 2025

  13. [21]

    Dissecting larval zebrafish hunting using deep reinforcement learning trained rnn agents.arXiv preprint arXiv:2510.03699, 2025

    Raaghav Malik, Satpreet H Singh, Sonja Johnson-Yu, Nathan Wu, Roy Harpaz, Florian Engert, and Kanaka Rajan. Dissecting larval zebrafish hunting using deep reinforcement learning trained rnn agents.arXiv preprint arXiv:2510.03699, 2025

  14. [22]

    Inputdsa: Demixing then comparing recurrent and externally driven dynamics.arXiv preprint arXiv:2510.25943, 2025

    Ann Huang, Mitchell Ostrow, Satpreet H Singh, Leo Kozachkov, Ila Fiete, and Kanaka Rajan. Inputdsa: Demixing then comparing recurrent and externally driven dynamics.arXiv preprint arXiv:2510.25943, 2025

  15. [23]

    Learning dynamics and the geometry of neural dynamics in recurrent neural controllers

    Ann Huang, Satpreet Harcharan Singh, and Kanaka Rajan. Learning dynamics and the geometry of neural dynamics in recurrent neural controllers. InWorkshop on Interpretable Policies in Reinforcement Learning RLC-2024, 2024

  16. [24]

    Measuring and controlling solution degeneracy across task-trained recurrent neural networks.ArXiv, pages arXiv–2410, 2025

    Ann Huang, Satpreet H Singh, Flavio Martinelli, and Kanaka Rajan. Measuring and controlling solution degeneracy across task-trained recurrent neural networks.ArXiv, pages arXiv–2410, 2025

  17. [25]

    Keep it real: rethinking the primacy of experimental control in cognitive neuroscience.NeuroImage, 222:117254, 2020

    Samuel A Nastase, Ariel Goldstein, and Uri Hasson. Keep it real: rethinking the primacy of experimental control in cognitive neuroscience.NeuroImage, 222:117254, 2020

  18. [26]

    Keypoint annotation for electrocommu- nication source separation with pikachu and raichu

    Kaden Zheng, Sonja Johnson-Yu, Satpreet Harcharan Singh, Denis Turcu, Federico Pedraja, Pratyusha Sharma, Naomi Saphra, Nathaniel Sawtell, and Kanaka Rajan. Keypoint annotation for electrocommu- nication source separation with pikachu and raichu. InThe Thirty-Ninth Annual Conf...

Pith tools

Reviewed August 3, 2026 · model on record in the stance chip above.