{"id":"beaa3f9c-5955-4f40-abc8-7114f96b8a24","arxiv_id":"2501.15001","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Simulated evolution of embodied agents shows visual tasks drive eye morphology, lenses emerge to balance acuity and light throughput, and poor acuity bottlenecks neural scaling.","lead":"Researchers evolved virtual agents with simulated eyes, letting selection pressure shape both eye design and learned behavior. The simulations reproduce known vision-evolution patterns, including compound eyes for navigation, camera eyes for object discrimination, and lenses emerging to balance sharpness with light collection.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 2.4's navigation/detection bifurcation is confounded: navigation is restricted to 1D photoreceptor mutations, so the 'task-only' comparison also changes the mutation operator.","rationale":"I read the paper as making a strong causal claim: task-specific selection, not experimental convenience, drives the three evolutionary outcomes. The lens and scaling results are interesting but secondary to the headline bifurcation. The reader's weakest assumption was broader external validity of rewards. I partially agree, but I think the more decisive issue is internal: Section 2.4's two arms differ in mutation operator. This is not merely a missing replicate; it means the reported bifurcation may be written into the search space. The concrete test above would settle it. If the test is clean, remaining concerns (no code, few replicates, overbroad 'fundamental bottleneck' language) are addressable via release and tempering; if the test fails, the primary claim should be downgraded. Therefore I keep the reader's CONDITIONAL verdict but with a sharper condition.","tokens_in":21688,"tokens_out":7078,"duration_ms":72374,"concrete_test":"Run the Section 2.4 Navigation evolution with the same full two-dimensional photoreceptor mutation operator used in the Detection arm (allow adding/removing photoreceptors in both horizontal and vertical dimensions), keeping the Navigation reward, environment, initialization, and evolutionary hyperparameters fixed. Repeat with at least five independent CMA-ES seeds and report the population-mean morphology at generation 50. If the population still converges to distributed 1x4 compound-type eyes, the bifurcation claim survives this specific confound; if it converges to a high-resolution forward-facing camera-type eye (e.g., 8x8 or 15x15), then the reported bifurcation was caused by the asymmetric mutation operator, not by the visual task alone.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing concern is an internal confound in the Section 2.4 bifurcation experiment. The paper states that in the Navigation task 'only photoreceptors along the horizontal width of the eye are added/removed' (Section 2.4), whereas Detection agents are free to add photoreceptors in two dimensions and converge to 15x15 camera-type eyes. At the same time, the authors claim to have observed the bifurcation 'by strictly changing the visual task' (Section 1). This is not a task-only comparison: the allowed mutation operator, the reward function, and the environment layout all change between the two arms. In particular, a 2D high-acuity camera eye cannot be reached by the Navigation mutation operator at all, so the reported absence of camera-type eyes in Navigation may be an artifact of the search space rather than of task-specific selection. The concern is testable and internal: the fitness function imposes no penalty for additional photoreceptors, so if 2D receptor mutations were allowed in Navigation and the population still converged to distributed 1x4 compound-type eyes, the claim would be supported. Without that control, the central 'what if' conclusion is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a computational framework that co-evolves the physical eye (photoreceptor layout, pupil, phase mask, refractive index), eye placement, and neural controllers of embodied agents trained by reinforcement learning, and uses it to revisit hypotheses about eye evolution. Three claims are made: (i) changing only the visual task produces a bifurcation, with navigation yielding distributed compound-type eyes and object discrimination yielding high-acuity camera-type eyes; (ii) when optical elements are allowed to mutate, lens-like structures evolve to resolve the trade-off between light collection and spatial precision; and (iii) task error follows power-law scaling with neural capacity that is bounded by visual acuity, so that poor acuity cannot be compensated by larger brains alone. The results are presented as computational counterfactual evidence for the evolutionary role of task-specific selective pressures.","tokens_in":21940,"tokens_out":3439,"duration_ms":36332,"significance":"If the central claims hold, the framework would be a genuinely useful instrument for testing what-if questions in vision evolution, because it couples physically motivated image formation with embodied reinforcement learning and evolutionary search. The paper has real strengths: the genotype spans morphology, optics, and neural capacity in a unified encoding; the imaging model in Eqs. (4)-(8) is concrete and physically grounded; the CPD and MTF analyses give quantitative, falsifiable characterizations of the evolved eyes; and the reported outcomes—compound versus camera morphologies, lens-like phase masks, and acuity-bounded scaling—are the kind of concrete predictions that computational evolution should provide. The significance is tempered, however, by the fact that the central bifurcation experiment changes more than the task, the lens result is shown on a single trajectory, and the scaling-law claim rests on finite sweeps whose uncertainty is not reported.","major_comments":[{"comment":"The headline claim that the camera/compound bifurcation is caused 'by strictly changing the visual task' (Section 1 and Section 2.4) is not supported as stated, because the two arms differ in the allowed mutation operator. The text says 'As the Navigation task is essentially two-dimensional, only photoreceptors along the horizontal width of the eye are added/removed,' while Detection agents add photoreceptors in two dimensions and converge to 15x15 camera eyes. Under the Navigation mutation operator, a 15x15 camera-type eye is unreachable by construction, so the absence of camera-type eyes in Navigation could be an artifact of the search space rather than of task-specific selection. The confound is testable and internal: run Navigation with the same two-dimensional photoreceptor mutation operator and report whether the population still converges to distributed 1x4 compound-type eyes. Without this control, the paper's central 'what-if' conclusion is not established.","section":"Section 2.4, Figure 3, Section 4 (Mutation operators)"},{"comment":"The lens-emergence result is reported on a single evolutionary trajectory: Phase I (pupil-only, 30 generations) is followed by Phase II (optics enabled) and the narrative describes a specific progression through cup, pinhole, unfocused, and focused lens eyes. CMA-ES with RL fitness evaluation is stochastic, and Section 4 reports no independent replicates for this experiment. To support the claim that evolution reliably discovers lens-like optics to resolve the light-throughput/spatial-precision trade-off, the authors should report multiple independent evolutionary runs with different seeds, showing the distribution of final PSFs, fitness gains, and the fraction of runs that produce well-focused lens-like phase masks. A single trajectory also cannot justify the statement that pinhole design 'plateaus' as a general evolutionary outcome.","section":"Section 2.5, Figure 4, Figure A2"},{"comment":"The scaling-law claims are load-bearing for the paper's third contribution, but the supporting evidence is under-specified. The paper reports fitted power laws such as L = (9.50e-3) N^0.69 for navigation as if they are deterministic, with no confidence intervals, number of seeds, or per-acuity-level fit quality. The claim that 'poor visual acuity creates a fundamental bottleneck that cannot be overcome by simply scaling neural capacity' is an extrapolation from finite sweeps of CPD and parameter counts; the data in Figure A5 show scattered errors, and the ceiling behavior needs to be quantified (e.g., by fitting saturating curves and reporting where the plateau begins) rather than asserted from the plotted envelopes. Relatedly, since CPD and input resolution jointly determine the network's input dimension, the independence of the 'acuity' and 'parameter' axes should be made explicit.","section":"Section 2.6, Figure 5, Appendix C"},{"comment":"The 'Image Quality' metric in Section 2.5 and Appendix A, defined as the product of MTF area above a noise floor and light throughput (which decreases quadratically with pupil radius), embeds exactly the trade-off that the paper claims evolution discovers. If this metric is used as the primary evidence that lens eyes are better, the argument is in part circular: the metric prefers lenses by construction. The independent evidence is behavioral—higher fitness in Phase II (Figure A2)—and the paper should make clear that the Image Quality metric is a post-hoc diagnostic rather than the objective being optimized, and should present the fitness comparison as the main test. If the authors intend the Image Quality metric as a causal driver, they need to show that evolution would behave differently under an alternative analysis metric.","section":"Section 2.5, Appendix A (Image Quality metric)"}],"minor_comments":[{"comment":"There are several typos in Appendix B: 'orientaotin' should be 'orientation', 'differnet' should be 'different', and 'genptype' should be 'genotype'. The same appendix also spells 'stomatopods' inconsistently and should be proofread.","section":"Appendix B"},{"comment":"In the Figure A3 caption, 'photreceptor' should be 'photoreceptor', and the sentence explaining the single-receptor case is syntax-heavy and should be rewritten for clarity.","section":"Figure A3 caption"},{"comment":"The phrase 'with with larger input arrays' appears in the paragraph following Figure 3; the duplicated 'with' should be removed.","section":"Section 2.4"},{"comment":"The definition of U_in(x,y) is given twice, once as Eq. (6) and once as an unnumbered equation before Eq. (8), and the sentence 'pupil function pupil function' in the paragraph after Eq. (4) is duplicated. These should be consolidated and edited.","section":"Section 4, Eqs. (4)-(7)"},{"comment":"The claim of 'approximately 10^20 unique agent vision types' is stated twice but no derivation or counting argument is given; a brief enumeration in the methods would make the claim checkable.","section":"Section 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is an ambitious crossover of evolutionary robotics, computational optics, and vision biology, and it clearly targets a broad high-profile audience. In my view the vision is sound, but the current evidence does not yet carry the weight of the three headline claims. The Section 2.4 confound is the most serious issue because it directly affects the paper's central 'what-if' narrative; it is, fortunately, fixable with an additional control condition. The single-trajectory lens result and the unreported uncertainty in the scaling fits are also fixable with additional experiments and analysis. I would not recommend rejection, because the framework and the specific counterfactual controls are exactly what would be needed to make the claims credible. I would also encourage the editor to ask the authors for a code/data availability statement, since reproducibility of evolutionary-robotics results depends on releasing environments and seeds."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I read the paper and I largely agree with the reader: conditional, moderate confidence, real novelty. The framework is the actual contribution — a unified genetic encoding over eye morphology, optics, and neural capacity, wrapped in a CMA-ES outer loop and PPO inner loop. That is new for vision evolution, and the imaging model with depth-independent PSF convolution is physically respectable. The lens-emergence experiment is a clean proof of concept: given a fixed morphology and task pressure, optical elements evolve to resolve the aperture/resolution trade-off. The scaling-law sweeps are systematic, even if calling the resulting ceilings a 'fundamental bottleneck' is too strong for finite sweeps.\n\nThe load-bearing soft spot is the Section 2.4 bifurcation experiment. The paper claims to show task-specific selection by 'strictly changing the visual task,' but it also changes the mutation operator: Navigation agents can only add/remove photoreceptors along the horizontal width, while Detection agents mutate freely in 2D. So a camera-type eye like the 15x15 array in Detection is simply unreachable in the Navigation arm. The reported absence of camera eyes in Navigation may be an artifact of the search space, not of selection. This is testable and internal — allow 2D receptor mutations in the Navigation task and see whether the population still converges to distributed 1x4 compound eyes. Without that control, the central 'what if' claim is not established.\n\nOther issues: no independent evolutionary replicates are reported (the lens trajectory looks like a single run), and no code or data are released, which hurts reproducibility for a computational evolution paper. The authors' own limitation paragraph is honest, but the abstract and discussion overclaim relative to the evidence.\n\nWho is this for? Evolutionary biologists interested in counterfactual vision, and computational researchers working on embodied sensor evolution. Engineers may borrow the genetic encoding for camera design. It is a solid paper to send to peer review, but the referee should demand the Navigation control, multiple seeds, and tempered scaling-law language. If the control holds, this could be a strong paper; if not, the bifurcation becomes a statement about the simulator's mutation space rather than about vision evolution.","headline":"A genuinely new evolutionary vision framework whose headline result is undercut by a mutation-operator confound; worth serious refereeing, but needs a control experiment.","tokens_in":22452,"tokens_out":3150,"would_cite":true,"duration_ms":32404,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Task-specific selection alone can recreate the evolution of compound eyes, camera eyes, and lenses in simulation.","keywords":["computational evolution","eye evolution","compound eyes","camera eyes","lens emergence","embodied agents","reinforcement learning","visual acuity scaling"],"falsifier":"Run the same evolution with the Navigation reward but make the maze exit identifiable only by a fine striped pattern, or with the Detection reward but make food and poison differ only in size rather than high-frequency texture. If the eye morphologies do not switch accordingly, the bifurcation is an artifact of task geometry rather than a general property of task-driven eye evolution. A second decisive check is to allow optical mutations in the Navigation task: if high-acuity lensed eyes also emerge there, the claim that lenses arise specifically from object-discrimination demand would be weakened.","tokens_in":21513,"feed_emoji":"👁️","tokens_out":7728,"duration_ms":65628,"temperature":0.7,"pith_summary":"Vision in nature spans simple light-sensitive patches to lensed camera eyes, but the role of specific environmental tasks in driving that diversity is hard to test because evolution cannot be rerun. This paper builds a computational evolution framework in which embodied agents co-evolve eye morphology, optics, and neural processing while learning a single visual task, and claims that the task alone determines the outcome. Agents evolved for maze navigation converge on distributed compound-type eyes, while agents evolved for object discrimination converge on two forward-facing, high-acuity camera-type eyes. When optical elements are allowed to mutate, lens-like focusing structures emerge to resolve the trade-off between light collection and spatial precision. The paper also reports power-law scaling between visual acuity and neural capacity, concluding that poor acuity is a bottleneck that more neurons cannot overcome.","feed_headline":"Simulated evolution rediscovers nature's two eye designs","feed_subtitle":"Navigation grows compound eyes; object discrimination grows camera eyes; lenses evolve on their own.","key_machinery":"The load-bearing mechanism is a unified genetic encoding with three independently mutating gene clusters: morphological genes (number, placement, and field of view of eyes), optical genes (pupil size, a programmable phase mask acting as a diffractive optical element, and refractive index), and neural genes (hidden-layer size and temporal memory). The outer evolutionary loop mutates and selects these genes over generations, while the inner loop trains each agent's policy with a reinforcement-learning algorithm during its lifetime, mirroring the Baldwin effect. Image formation is modeled by convolving the scene with a physically based point-spread function, so closing the aperture sharpens the image but reduces light throughput in proportion to the square of the pupil radius and increases noise. This light-versus-precision trade-off is the pressure that lens evolution is claimed to resolve, and acuity is quantified in cycles per degree (CPD) so that sensory quality can be compared with network parameter count.","core_discovery":"At its center, the paper claims that task-specific selection is sufficient to reproduce the major branch points of eye evolution. Agents are initialized as a single photoreceptor and then selected by an outer evolutionary loop while each genotype is trained by an inner reinforcement-learning loop. Only the reward function differs between experiments, and the eyes that emerge diverge accordingly: orientation and obstacle avoidance produce many distributed, low-resolution eyes with broad field of view, while discrimination between visually similar objects produces two frontal, high-resolution eyes. In a separate experiment, evolution first narrows the pupil toward pinhole-like eyes and then, once optical mutations are allowed, discovers lens-like phase profiles that keep images sharp while admitting more light, escaping the pinhole's performance ceiling. Finally, task error is reported to obey power laws in both neural parameters and acuity, with acuity setting a floor: increasing network size alone cannot compensate for poor vision. The authors present these results as a method for turning 'what-if' questions about natural vision into counterfactual computational experiments.","pith_inferences":["A natural extension is to evolve agents on multiple tasks at once; the paper itself notes that natural vision solves many tasks simultaneously, so multi-task rewards might produce intermediate or mosaic eyes rather than the clean bifurcation seen under isolated tasks.","The acuity-bottleneck result suggests a concrete design rule for artificial vision systems: co-design sensor resolution and network capacity, because improving either alone will hit a ceiling.","The lens-emergence result could be tested further by varying ambient light intensity rather than fixing noise, which would directly probe whether lighting conditions alone can trigger lens evolution.","The methodology implies a general counterfactual engine for vision: any environmental change, such as a world without movement or a monochromatic world, can be instantiated as a reward and environment and then evolved."],"forward_implications":["Comparative hypotheses about a species' visual ecology can be tested by recreating its presumed task in simulation and checking whether the evolved morphology matches.","High-acuity vision under dim conditions requires an optical innovation, not simply more photoreceptors, because pinhole-like apertures trade away light.","Embodied agents cannot escape sensory limits by adding parameters alone; acuity and network capacity must improve together to reduce task error.","Starting from a single primitive photoreceptor, the same genome can diverge into compound and camera designs depending only on behavioral demand, which offers a computational explanation for convergent eye types in nature.","Because the genotype spans roughly 10^20 configurations, the same machinery can serve as a generative design tool for manufacturable bio-inspired imaging systems."],"supporting_citations":[{"why":"supplies the thesis that eye evolution is driven by visually guided behavior, which the three tasks operationalize","marker":"[4]"},{"why":"supplies the distinction between orientation/navigation vision and object vision that the task split is designed to test","marker":"[31]"},{"why":"provides the simple-eye starting point and the light-collection versus acuity trade-off that lens evolution is claimed to resolve","marker":"[5]"},{"why":"gives the honeybee compound-eye navigation example that the Navigation result is compared against","marker":"[34]"},{"why":"supplies cycles per degree as the metric used to measure and compare evolved visual acuity","marker":"[33]"},{"why":"establishes the scaling-law framework that the acuity-versus-neural-capacity results are patterned on","marker":"[32]"},{"why":"provides the policy-optimization algorithm used in the inner learning loop","marker":"[52]"},{"why":"provides the physics-based simulation environment in which agents move, sense, and learn","marker":"[35]"}],"fun_headline_variants":["Task pressure alone drives eye evolution in simulated agents","Simulated evolution reveals why eyes split into two designs","Lenses emerge naturally in evolved eyes, solving a tradeoff","Artificial evolution tests what-if questions about real eyes","Virtual creatures evolve camera and compound eyes on demand"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim depends on the assumption that the three hand-designed reward functions and their environments capture the natural selective pressures that shaped real vision; if different reward formulations or scene layouts produced different eye morphologies, the results would describe the simulator rather than vision evolution.","fun_headline_variants_meta":{"raw":{"variants":["Task pressure alone drives eye evolution in simulated agents","Simulated evolution reveals why eyes split into two designs","Lenses emerge naturally in evolved eyes, solving a tradeoff","Artificial evolution tests what-if questions about real eyes","Virtual creatures evolve camera and compound eyes on demand"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000603,"raw_usage":{"total_tokens":2836,"prompt_tokens":986,"completion_tokens":1850,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":602,"completion_tokens_details":{"reasoning_tokens":1775}},"tokens_in":602,"tokens_out":1850,"duration_ms":12592,"temperature":1.0,"reasoning_tokens":1775,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T14:43:20.661380+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same evolution with the Navigation reward but make the maze exit identifiable only by a fine striped pattern, or with the Detection reward but make food and poison differ only in size rather than high-frequency texture. If the eye morphologies do not switch accordingly, the bifurcation is an artifact of task geometry rather than a general property of task-driven eye evolution. A second decisive check is to allow optical mutations in the Navigation task: if high-acuity lensed eyes also emerge there, the claim that lenses arise specifically from object-discrimination demand would be weakened.","supporting_citations":[{"cited_title":"Frontiers in Neuroanatomy 16 (2022) https://doi.org/10.3389/fnana.2022.789375","cited_arxiv_id":null,"evidence_quote":"supplies the distinction between orientation/navigation vision and object vision that the task split is designed to test"},{"cited_title":"Journal of Experimental Biology 199(1), 237–244 26 (1996)","cited_arxiv_id":null,"evidence_quote":"gives the honeybee compound-eye navigation example that the Navigation result is compared against"},{"cited_title":"Trends in ecology & evolution 33(5), 358–372 (2018)","cited_arxiv_id":null,"evidence_quote":"supplies cycles per degree as the metric used to measure and compare evolved visual acuity"},{"cited_title":"In: IROS, pp","cited_arxiv_id":null,"evidence_quote":"provides the physics-based simulation environment in which agents move, sense, and learn"}],"review_version":1}