{"id":"9a20ad8f-ec85-4572-bdae-1b721f31c0f7","arxiv_id":"2505.13615","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"From 53 million tracked mosquito positions, Bayesian inference learns Langevin force fields for visual, CO2, and combined cues, and predicts crowding around a human head.","lead":"This paper uses infrared 3D tracking and Bayesian inference to learn equations of motion that describe how female Aedes aegypti mosquitoes fly toward visual and carbon dioxide targets. The learned model reproduces observed flight patterns and predicts mosquito density around a human head, a step toward designing better mosquito traps.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Human-head prediction relies on an untested linear rescaling of cue length scales; CO2 plume range is set by flow rate, not sphere radius, so transferring the 8-inch combined-cue model to a 12-inch head is not justified.","rationale":"The reader's weakest_assumption and my concern are the same: the human-head prediction is the headline claim and it rests on a size-rescaling rule that is stated but not demonstrated. I agree with the CONDITIONAL verdict. The paper has real strengths: the dataset is large and publicly available, the inference pipeline is validated on synthetic ground truth (Supp. Figs. 3-4), and the target experiments (visual, CO2, combined) are compared via KS distances. Those in-sample/self-consistency checks support the learned single-cue and 8-inch combined-cue models. The problem is specifically the transfer from 8-inch to 12-inch for combined cues. The suggested test is feasible: the authors already run 4- and 12-inch visual-sphere experiments (Supp. Fig. 11), so a 12-inch combined-cue experiment is a modest extension. If the rescaling passes, the central claim would be substantially strengthened; if it fails, the human-head prediction should be downgraded to a qualitative illustration rather than a quantitative validation. No further verdict change is needed: CONDITIONAL is the right call.","tokens_in":23491,"tokens_out":7151,"duration_ms":70017,"concrete_test":"Run a 12-inch black-sphere CO2 experiment (same 0.24 L/min release, same chamber protocol) with enough mosquitoes to match the 8-inch dataset, and compare the resulting (d, v, v·d) density heatmaps and, if desired, the independently inferred f∥ and f⊥ to the rescaled 8-inch combined-cue model (d0 multiplied by 1.5). Report the KS distance between experimental and predicted CDFs against the 8-inch benchmark in Fig. 5O; if it exceeds that benchmark by more than the bootstrap uncertainty, the linear d0 rescaling fails and the human-head prediction is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that the model 'quantitatively replicate[s] the experimental mosquito densities around the human head' (Fig. 4F,H)—requires transferring the force learned for an 8-inch black sphere with CO2 to a 12-inch 'black sphere emitting CO2' via the statement in Supplementary Materials III.B: 'we can rescale the d0 parameters to approximate the response to a larger or smaller target.' This rescaling is not validated for the combined cue. The training data for combined cues exist only for the 8-inch sphere; no 12-inch combined-cue experiment is reported, and Supplementary Fig. 11 tests size scaling only for the visual cue (4-inch and 12-inch spheres), not for CO2 or for visual+CO2. The physical basis of a single linear d0 rescaling is also questionable: d0 bundles the visual detection range, which plausibly scales with target radius through the mosquito's angular resolution, and the CO2 detection range, which for a fixed 0.24 L/min flow rate is set by plume advection/diffusion and should be nearly independent of the mounting sphere's radius. Rescaling both by the same factor 1.5 (8-inch to 12-inch) may thus inflate the CO2 contribution. Finally, the human-head comparison is presented only as side-by-side density heatmaps; unlike the target comparisons in Fig. 5O, no KS distance, error bars, or other quantitative metric is reported. The central transfer claim therefore rests on an untested scaling and visual inspection.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript combines 3D infrared tracking of freely flying Aedes aegypti mosquitoes with sparse Bayesian dynamical-systems inference to learn stochastic Langevin models of host-seeking flight. The authors record trajectories in an empty chamber and around visual (black sphere), CO2 (white sphere with 0.24 L/min CO2), and combined visual+CO2 targets, then infer spatially resolved longitudinal and transverse behavioral forces. The learned models reproduce trajectory-density statistics for the fitted sphere conditions and are invoked to predict mosquito densities around a human subject approximated as a 12-inch black sphere emitting CO2. The paper also reports that combined-cue responses are not a linear superposition of single-cue responses, and it provides a web application for interactive simulation.","tokens_in":23857,"tokens_out":4207,"duration_ms":40134,"significance":"If the central transfer claim holds, the framework would provide a quantitative, mechanistic description of mosquito host-seeking in windless conditions and demonstrate an advance over trajectory-statistics-only studies. The paper's strengths include the unprecedented size of the 3D tracking dataset, synthetic-data validation of the inference pipeline (Supp. Figs. 3-4), publicly archived data and code, a sparsity-selected model with explicit BIC comparison, and an independent out-of-sample test in the human-head experiment. The careful Bayesian formulation with posterior uncertainties and the non-additivity analysis are valuable contributions. However, the headline prediction of human-head densities rests on an untested target-size rescaling, so the significance of the transfer claim is currently uncertain.","major_comments":[{"comment":"The central predictive claim that the model \"quantitatively replicate[s]\" mosquito densities around the human head (Fig. 4F,H) depends on transferring the combined-cue force learned for an 8-inch sphere to a 12-inch human head by rescaling the distance scale d0. This rescaling is not validated for the combined cue: Supplementary Fig. 11 tests size scaling only for the visual cue (4-inch and 12-inch spheres), and no 12-inch combined-cue experiment is reported. Moreover, d0 bundles the visual detection range, which could plausibly scale with target radius, and the CO2 detection range, which for a fixed flow rate is set by plume advection and diffusion and should be nearly independent of the mounting sphere's radius; rescaling both by the same factor is therefore not physically justified. The supplementary statement \"we can rescale the d0 parameters\" does not specify the functional form of the rescaling or quantify its uncertainty.","section":"Supplementary Materials III.B; main Fig. 4E-H"},{"comment":"The human-head comparison, which is the only out-of-sample test of the model, is presented only as side-by-side density heatmaps. Unlike the sphere-condition comparisons, which report Kolmogorov-Smirnov distances (Fig. 5O), no KS distance, error bars, or other quantitative metric is given for Fig. 4F,H. The word \"quantitatively\" in the claim is therefore not supported by any reported statistic; the authors should provide a quantitative comparison such as a KS distance or a distributional discrepancy measure between the experimental and predicted human-head densities.","section":"Main Fig. 4F,H"},{"comment":"The sphere-condition validations are in-sample checks: the simulated densities and CDFs are generated from models fitted to the same trajectories with which they are compared. They establish self-consistency of the inference pipeline but not predictive power. The paper's wording that simulations \"quantitatively match\" these data (Figs. 3E,H and 4D) and that the KS distances \"highlight the accuracy of our model in predicting mosquito behavior\" (Fig. 5O) overstates the evidence; the predictive claim should be restricted to the human-head transfer, and the in-sample nature of the sphere-condition comparisons should be explicitly acknowledged.","section":"Main Figs. 3E, 3H, 4D and Fig. 5M-O"}],"minor_comments":[{"comment":"The phrase \"Kolmogorov–Smirno (KS) distance\" should read \"Kolmogorov–Smirnov (KS) distance.\"","section":"Main text, paragraph after Fig. 5"},{"comment":"The abstract states that the model was trained on \"more than 20,000,000 data points,\" while the Results and Supplementary Table 2 report 53,669,795 data points across all experiments; please clarify which subset is used for training the sensory-response models.","section":"Abstract and Results"},{"comment":"The rescaling sentence \"we can rescale the d0 parameters\" is the only methodological description of how the 8-inch-sphere model is adapted to the 12-inch human head; a precise equation (e.g., d0_new = d0_old * R_new/R_old) and a test of sensitivity to this choice should be added.","section":"Supplementary Materials III.B"},{"comment":"The photographic inset of the human subject is useful, but the matching between the photograph and the schematic \"black sphere emitting CO2\" is not self-evident; a dimensioned overlay or annotation would improve clarity.","section":"Main Fig. 4E"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is technically interesting and the dataset is substantial, but the headline transfer claim needs stronger support. I would ask the authors to either provide a dedicated combined-cue experiment with a different target size, or to demonstrate through sensitivity analysis and a quantitative human-head comparison that the rescaling uncertainty does not change the main conclusions. In its current form, the paper is best framed as a self-consistent modeling framework with promising but not yet fully validated predictive power."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing you should know: this is a substantial data-and-methods paper whose central transfer claim—quantitative prediction of mosquito densities around a human head from a model trained on an 8-inch sphere—is shakier than the abstract implies. The paper itself is worth a serious look; the human-head prediction is the part to press on.\n\nWhat's actually new: the scale of the 3D tracking dataset (53.7M data points, 477k trajectories) and the application of sparse Bayesian inference to extract interpretable Langevin force fields for mosquito host-seeking. The synthetic tests in Supp Figs. 3-4 show the pipeline recovers known forces with small error; that's a real check. And the behavioral distinction—visual cues produce taxis (directed turning), CO2 produces kinesis (undirected slowing/tumbling), and the combined cue produces a distinct orbiting state with non-additive integration—is a clean, physically interpretable result. The code and data are on Zenodo/GitHub, which I always weigh in the plus column.\n\nWhere the soft spots are: the human-head prediction depends on rescaling the learned force from an 8-inch sphere to a 12-inch head by a single linear factor on the length scale d0 (Supp. III.B). The paper never validates this rescaling for CO2 or for the combined cue. Supp. Fig. 11 tests target-size scaling only for visual 4-inch and 12-inch spheres, not for CO2 or visual+CO2. And physically, the CO2 detection range for a fixed 0.24 L/min flow rate is set by plume advection/diffusion and should be nearly independent of the mounting sphere's radius—so scaling both visual and CO2 length scales by the same factor is suspect. That is a load-bearing assumption for the headline claim. Also, the sphere-condition reproductions (Figs. 3E/H, 4D) are in-sample: same fits compared against the same trajectories. They establish self-consistency, not prediction. And the human-head comparison in Fig. 4F/H is side-by-side heatmaps with no KS distance or error bars, in contrast to the quantitative comparisons in Fig. 5O.\n\nThe paper does not pretend these are absent—the rescaling is stated in the Supplementary, and the non-additivity result is honestly reported. The concerns are addressable: a proper test would be a 12-inch combined-cue experiment, or at least a sensitivity analysis of the rescaling on the CO2 component.\n\nBottom line: if the transfer claim were removed or rephrased as a modeling assumption, the paper would be a strong empirical and methodological contribution. As is, the central claim rests on an untested scaling and visual inspection. This deserves peer review—it's not a desk reject—but a referee should require a quantitative, out-of-sample validation of the human-head prediction or a clear downgrading of that claim.\n\nRecommendation: send it, with the human-head transfer as the main question.","headline":"A strong data-and-inference paper whose headline transfer to human targets rests on an untested rescaling; the single-cue models and synthetic validation are the solid core.","tokens_in":24369,"tokens_out":2380,"would_cite":false,"duration_ms":22313,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a stochastic dynamical model of mosquito flight, inferred from more than 20 million tracking data points collected around simple visual and carbon dioxide stimuli, quantitatively predicts how Aedes aegypti mosquitoes…","keywords":["Aedes aegypti","host-seeking behavior","Bayesian inference","Langevin dynamics","3D flight tracking","carbon dioxide cues","visual cues","dynamical systems inference"],"falsifier":"Re-run the tracking and inference pipeline on a 12-inch black sphere releasing CO2 at 0.24 L/min and compare the empirically inferred force field and density with the prediction obtained by rescaling the 8-inch-sphere model; if the rescaled model does not reproduce the 12-inch-sphere data, the transfer to the human head is unsupported. Separately, repeat the human-head experiment with the body's heat and skin odors blocked, for instance with a heat-reflective suit under the white outfit: if the measured head density changes when those cues are removed, the sphere-only model is missing cues that matter.","tokens_in":23315,"feed_emoji":"🦟","tokens_out":13453,"duration_ms":112659,"temperature":0.7,"pith_summary":"The paper sets out to show that mosquito host-seeking flight, which has resisted quantitative prediction, can be captured by a stochastic dynamical model learned directly from tracking data. Using 3D infrared tracking of female Aedes aegypti flying around a black sphere, a carbon dioxide source, and their combination, the authors infer the forces steering mosquito flight through sparse Bayesian inference on more than 20 million trajectory data points. The learned model reproduces the measured flight densities around each stimulus and, rescaled to head size, quantitatively replicates the mosquito density measured around a real human head. If the claim holds, host-seeking behavior becomes a force field that can be simulated, giving trap design and bite-risk assessment a predictive tool where current capture devices are only 10 to 50 percent effective.","feed_headline":"Sphere-trained model predicts mosquito swarms around human heads","feed_subtitle":"Learned from 20 million mosquito flight measurements, the model transfers from spheres to real targets.","key_machinery":"The load-bearing object is the force-field representation of the Langevin dynamics $dv/dt = f(r,v) + \\xi$. The behavioral force is decomposed as $f = f_\\parallel \\hat{v} + f_\\perp (I - \\hat{v}\\hat{v})\\cdot\\hat{d}$, splitting it into a longitudinal 'throttle' $f_\\parallel$ (acceleration or braking along the flight direction) and a transverse 'turning' force $f_\\perp$ (steering toward or away from the target), each expanded in a tensor-product basis of Laguerre polynomials in speed $v$ and distance $d$ and Legendre polynomials in the alignment $\\hat{v}\\cdot\\hat{d}$. Sparse Bayesian regression — a sparsity-promoting Gaussian prior, expectation-maximization updates of coefficients and noise, sequential thresholding, and Bayesian-information-criterion model selection — condenses tens of millions of tracking data points into a few nonzero coefficients. The learned force fields are the mechanism: once inferred they can be simulated forward to generate whole trajectory ensembles whose density statistics are compared with experiment, and they are what gets rescaled through the distance scale $d_0$ to move from an 8-inch training sphere to a 12-inch human head.","core_discovery":"The central discovery, stated on the paper's own terms, is that the full repertoire of mosquito host-seeking flight in windless conditions — free flight, visual attraction, carbon-dioxide-induced tumbling, and orbiting around combined cues — is described by a single stochastic Langevin equation $dv/dt = f(r,v) + \\xi$, in which the behavioral force splits into a longitudinal throttle $f_\\parallel$ along the flight direction and a transverse turning force $f_\\perp$ perpendicular to it, both depending only on flight speed, distance from the target, and flight direction relative to the target. Sparse Bayesian inference extracts these forces from the tracking data and reveals cue-specific mechanisms: visual cues act through a speed potential that lowers the preferred speed near the target plus a taxis-then-repulsion turning force; carbon dioxide acts through a kinesis-like deceleration from about 0.7 m/s to 0.2 m/s within 0.4 m; combined cues produce an amplified visual turning force and sustained orbiting. The combined-cue response is not a linear superposition of the individual responses, and the combined-cue model, rescaled to head size, quantitatively reproduces the measured mosquito density around a real head wearing a black hood.","pith_inferences":["If the force-field picture holds, the inferred $f_\\perp$ field can be read as a sensory map: its turning wall near 0.4 m coincides with the visual range estimated from the eye's 12.3° minimum resolvable angle, so the model could measure perception ranges directly from flight behavior alone.","The non-additivity of combined cues suggests a testable gating hypothesis: CO$_2$ may amplify visual steering rather than act as an independent attractant, which would predict a visual turning force in the combined model that is strongest exactly where CO$_2$ is detectable; the paper documents that pattern but leaves the amplification unquantified.","Because the learned models are cheap to simulate, the $d_{50}$ metric could be inverted into a design loop — optimizing trap geometry, bait strength, and multi-target layouts in simulation before field trials — an application the paper lists as future work rather than demonstrating.","The windless chamber makes the CO$_2$ zone static; advecting the plume and re-inferring the forces would test whether the same kinesis-and-taxis fields simply acquire an upwind bias, or whether the inferred forces are wind-conditional rather than universal stimulus responses."],"forward_implications":["Simulated trajectories of the learned models match the experimental density distributions around targets for no cue, visual, CO$_2$, and combined cues, with small Kolmogorov–Smirnov distances, and the validation statistics (mean squared displacement, directional and speed correlations) were not used during inference.","Bite risk, quantified as the radius $d_{50}$ containing 50 percent of trajectories, shrinks from about 0.65 m with no cues, to 0.4 m with a visual target, to 0.25 m with CO$_2$, and to 0.2 m with combined cues, so combined stimuli concentrate mosquitoes in a tighter zone around the host.","The response to combined cues is not additive: the best non-negative linear superposition of the single-cue forces ($f_{\\parallel,\\mathrm{lin}} = 0.63 f_{\\parallel,\\mathrm{visual}} + 0.39 f_{\\parallel,\\mathrm{CO}_2}$; $f_{\\perp,\\mathrm{lin}} = 1.63 f_{\\perp,\\mathrm{visual}}$) leaves large residuals near the target and fails to reproduce the combined-cue density, implying nonlinear integration of s","Because the combined-cue sphere model, rescaled to head size, reproduces the density around a human head, the framework transfers from artificial stimuli to a realistic human target under windless conditions.","The inference pipeline requires no human-defined behavioral labels and is presented as directly applicable to other stimuli, such as odor and heat, and to other vector species."],"supporting_citations":[{"why":"Introduces the sparse Bayesian learning scheme whose sparsity-promoting Gaussian prior regularizes the inferred force coefficients.","marker":"[36]"},{"why":"Establishes the basis-function expansion of forces inferred from stochastic trajectory data, the template for representing and learning the behavioral force.","marker":"[44]"},{"why":"Describes the dual-infrared-camera tracking system used to record 3D mosquito positions.","marker":"[45]"},{"why":"Prior account of how CO2 plumes trigger upwind flight and modulate visual responsiveness in mosquitoes, the behavioral baseline this model builds on.","marker":"[6]"},{"why":"Sparse identification of dynamical systems; supplies the sequential thresholding used to eliminate small force coefficients.","marker":"[55]"},{"why":"Bayesian information criterion, used to select the sparse model that best trades fit against complexity.","marker":"[58]"},{"why":"Theory showing that speed reduction alone concentrates moving particles, the mechanism invoked for CO2-induced accumulation.","marker":"[64]"},{"why":"Evidence that CO2 excites mosquitoes and sensitizes them to other host cues, cited for the non-additive combined response.","marker":"[21]"},{"why":"Measurement of the mosquito eye's visual resolution, used to link the roughly 0.4 m attraction zone to visual range.","marker":"[27]"}],"fun_headline_variants":["Bayesian model predicts mosquito pursuit from 20M flight points","Single stochastic equation explains mosquito host-seeking flight","Learned from 20M tracks, model foretells mosquito human-target swarms","Visual and CO2 cues decoded in mosquito flight model","Sparse Bayesian inference maps mosquito flight to human targets"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The head prediction assumes a mosquito treats a human head as a black CO2-emitting sphere that differs from the 8-inch training sphere only in size, so that one distance rescaling transfers the learned forces, and that the rest of the human body, its heat, and skin odors add no cues that change the flight pattern.","fun_headline_variants_meta":{"raw":{"variants":["Bayesian model predicts mosquito pursuit from 20M flight points","Single stochastic equation explains mosquito host-seeking flight","Learned from 20M tracks, model foretells mosquito human-target swarms","Visual and CO2 cues decoded in mosquito flight model","Sparse Bayesian inference maps mosquito flight to human targets"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000185,"raw_usage":{"total_tokens":1307,"prompt_tokens":913,"completion_tokens":394,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":529,"completion_tokens_details":{"reasoning_tokens":311}},"tokens_in":529,"tokens_out":394,"duration_ms":4144,"temperature":1.0,"reasoning_tokens":311,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:13:11.602961+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the tracking and inference pipeline on a 12-inch black sphere releasing CO2 at 0.24 L/min and compare the empirically inferred force field and density with the prediction obtained by rescaling the 8-inch-sphere model; if the rescaled model does not reproduce the 12-inch-sphere data, the transfer to the human head is unsupported. Separately, repeat the human-head experiment with the body's heat and skin odors blocked, for instance with a heat-reflective suit under the white outfit: if the measured head density changes when those cues are removed, the sphere-only model is missing cues that matter.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces the sparse Bayesian learning scheme whose sparsity-promoting Gaussian prior regularizes the inferred force coefficients."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the basis-function expansion of forces inferred from stochastic trajectory data, the template for representing and learning the behavioral force."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Describes the dual-infrared-camera tracking system used to record 3D mosquito positions."},{"cited_title":"Van Breugel, J","cited_arxiv_id":null,"evidence_quote":"Prior account of how CO2 plumes trigger upwind flight and modulate visual responsiveness in mosquitoes, the behavioral baseline this model builds on."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Sparse identification of dynamical systems; supplies the sequential thresholding used to eliminate small force coefficients."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Bayesian information criterion, used to select the sparse model that best trades fit against complexity."},{"cited_title":"Farrell, J","cited_arxiv_id":null,"evidence_quote":"Theory showing that speed reduction alone concentrates moving particles, the mechanism invoked for CO2-induced accumulation."},{"cited_title":"Majeed, S","cited_arxiv_id":null,"evidence_quote":"Evidence that CO2 excites mosquitoes and sensitizes them to other host cues, cited for the non-additive combined response."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Measurement of the mosquito eye's visual resolution, used to link the roughly 0.4 m attraction zone to visual range."}],"review_version":1}