{"id":"5dab871d-0588-453b-9b11-7b454a1ff99d","arxiv_id":"2509.01102","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A new width-normalized discretization method quantifies fossil movement paths, and its application to Cruziana semiplicata suggests three recurring behavioural morphotypes.","lead":"Fossil trails are turned into quantifiable movement data by digitizing paths, dividing them into width-based steps, and comparing turning-angle statistics with t-tests. Applied to the trilobite trace Cruziana semiplicata, the method identifies three recurring path morphologies, interpreted as behavioural variants that persist across localities.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Morphotype persistence is not established: the three groups are defined post hoc from the same p-value matrix used to test locality effects, and p>0.1 is treated as evidence of similarity despite tiny samples.","rationale":"The reader's conditional verdict is appropriate. The methodology and toolkit contribution are genuinely useful, and the paper is transparent about its steps. However, the headline empirical claim about three persistent morphotypes rests on a post hoc analysis of the same p-value matrix used to define and then remove 'different' specimens, creating a circular validation loop. The additional reliance on p > 0.1 as evidence of behavioural similarity, especially with very small samples and high variance, means the cross-locality persistence claim is not established. The reader emphasized the p > 0.1 issue; I agree, but the more structurally load-bearing problem is that the three-group taxonomy itself is not independently validated. This does not change the verdict—conditional remains correct—but it sharpens the specific condition: the empirical claim requires re-analysis without threshold-based post hoc grouping. No independent code or data are currently provided, so the proposed test cannot be run by an external reviewer.","tokens_in":20790,"tokens_out":3444,"duration_ms":42775,"concrete_test":"Release the digitized path coordinates and analysis code, then re-analyze the specimen-level turning-angle summaries without the paper's thresholds: fit a Gaussian mixture model (or other model-based clustering) to each specimen's mean/variance (or full angle distribution), choose the number of clusters by BIC with bootstrap stability, and test whether cluster membership proportions differ among localities using a small-sample-appropriate test (e.g., Fisher-Freeman-Halton exact test). If the optimal cluster count is not 3, or if clusters 2 and 3 are absent in non-Spanish localities, the persistence claim collapses. Also compute the minimum detectable effect/power for the existing sample sizes to show whether p > 0.1 comparisons could ever detect a meaningful behavioural difference.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—three behaviourally distinct morphotypes persisting across localities—depends on two unsupported statistical moves. First, §3.7.4 defines 'different' specimens by p < 0.1 against ≥25% of all specimens in the same pairwise Welch's t-test matrix, then §3.8.4 removes them and re-runs the locality tests; any apparent disappearance of locality differences is therefore partly built into the selection rule. The three morphotypes themselves (§3.9) are read off the black (p > 0.1) blocks of that same matrix, so there is no independent evidence that there are exactly three groups. Second, the paper repeatedly treats non-rejection as positive evidence of similarity (e.g., §3.8.1, §3.8.3). With Wales N=5, Poland N=3, and Russia variance σ² ≈ 15.2 while other groups range 5.8–8.8, a p > 0.1 result has low power and cannot distinguish 'same behaviour' from 'inadequate sample.' No power analysis or multiple-comparison correction is reported, although the full specimen matrix involves roughly 4,000 comparisons. These problems bear directly on the abstract's claim that morphotypes 'persisted across multiple geographic localities.'","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a quantitative methodology for discretizing fossil movement paths, converting trace-fossil trajectories into turning-angle distributions, and then applying Welch's t-tests at specimen, subgroup, and locality levels. The method is demonstrated on Cruziana semiplicata from Spain, Oman, Poland, Russia, and Wales. Three previous assertions about C. semiplicata behaviour are tested: whether resting (Rusophycus) changes subsequent movement, whether Spanish subgroups show one stereotyped behaviour, and whether geographic groups differ. The authors report that resting does not alter behaviour, that locality-level differences largely disappear after removing specimens identified as 'different' by pairwise t-tests, and that three morphotypes (one dominant, two subordinate) persist across localities. The central biological conclusion is that C. semiplicata tracemakers show a dominant stereotyped behaviour with two recurring variants, and that Seilacher's two-ichnosubspecies split is not supported.","tokens_in":21033,"tokens_out":3142,"duration_ms":39497,"significance":"If the central morphotype claim were established, this would be a useful contribution: it imports movement-ecology metrics into ichnology and offers an open, repeatable pipeline for quantifying trace-fossil trajectories. The digitization, width-normalized segmentation, and turning-angle calculation are clearly described and represent a genuine methodological advance. The paper also makes concrete, testable claims about behaviour persistence and morphotype structure. However, the statistical evidence for the three morphotypes and their cross-locality persistence is not currently convincing because the groups are defined post hoc from the same pairwise t-test matrix used to test locality effects, and non-rejection of the null hypothesis is repeatedly treated as evidence of similarity without power analysis or multiple-comparison control. The methodological toolkit itself is promising, but the demonstration of its utility rests on inferences that need stronger statistical support.","major_comments":[{"comment":"The 'different' specimens are defined using the same specimen-vs-specimen t-test matrix that is then used to reinterpret group differences. Section 3.7.4 sets a p<0.1 threshold against ≥25% of specimens, §3.8.4 removes those 40 specimens and re-runs the group and subgroup t-tests, and §3.9 uses the black blocks of the same matrix to define morphotypes 2 and 3. This is circular: the selection rule and the final morphotype grouping are both derived from the same pairwise test results, so the disappearance of locality differences after outlier removal is partly built into the procedure. An independent classification (e.g., clustering on the turning-angle distributions themselves, or a pre-registered outlier rule not derived from the full pairwise matrix) is needed before the persistence claim can be evaluated.","section":"§3.7.4, §3.8.4, §3.9"},{"comment":"Non-rejection of H0 is repeatedly treated as positive evidence of similarity or persistence. For example, §3.8.1 says nine pairs with p>0.1 'suggest on average a similarity', and §3.8.3 reports 'no evidence to reject H0 between the Oman and Spain L groups' as support for cross-locality similarity. This reasoning is invalid without a power analysis or an equivalence test: with Wales N=5, Poland N=3, and Russia variance (σ, as reported in §3.8.3) of 15.20 against 5.84–8.75 for other groups, a p>0.1 result has very low power to distinguish 'same behaviour' from 'inadequate sample'. The load-bearing conclusion that morphotype 1 persists across localities is therefore not established even if the morphotype description is accurate.","section":"§3.5, §3.8.1, §3.8.3"},{"comment":"The existence of three morphotypes is inferred from visual inspection of the p-value matrix, not from an independent quantitative cluster analysis. Working from the black (p>0.1) blocks of the same matrix that was used for outlier removal, the authors group specimens into morphotypes and then use those groups to explain locality differences. This post hoc procedure does not validate the number of morphotypes or their distinctness. The paper should provide an independent clustering or model-selection step (for example, on the full direction-adjusted turning-angle data) and quantify support for three groups versus one or two groups, with appropriate uncertainty.","section":"§3.9, Figure 3.10"},{"comment":"No multiple-comparison correction is applied anywhere, despite thousands of pairwise Welch's t-tests. The specimen matrix alone involves 136 specimens (Table 3.3) and thus on the order of 9,000 pairwise comparisons, of which many are non-independent because the same specimens appear in multiple tests. Using an uncorrected p<0.1 threshold across this many comparisons guarantees a substantial number of false 'different' classifications, which directly affects the identification of the 40 'different' specimens and hence the morphotype grouping. Either a correction (e.g., FDR) or an explicit justification for not using one is needed.","section":"§3.7.4, Table 3.3"}],"minor_comments":[{"comment":"The text reports 'Russia had a much higher variance (σ = 15.20)' while other groups have variances between 5.84 and 8.75. σ is standard notation for standard deviation, not variance. Please clarify whether 15.20 is a variance or standard deviation, and correct the notation throughout.","section":"§3.8.3"},{"comment":"In §3.5 and Figure 3.12 the source for Wales is cited as 'Crimes 1968; 1970, Plate 5', but the reference list contains only Crimes (1970). Please add the 1968 reference or correct the citation.","section":"References"},{"comment":"The caption ends with 'while = strong evidence (p < 0.01)'; this should read 'white = strong evidence'.","section":"Figure 3.10 caption"},{"comment":"The paper states it 'provides an open-source groundwork', but no repository, code link, or data availability statement is included. Since the methodology is the main contribution, making the MATLAB/R scripts and the digitized coordinate data available would strengthen reproducibility.","section":"Throughout"},{"comment":"The choice of segment-distance multiplier (0.5) is stated but not motivated or tested for sensitivity. At minimum, report whether the main conclusions change for other multipliers (e.g., 0.25, 1.0), since autocorrelation and turning-angle distributions depend on this choice.","section":"§3.6.2, §3.6.3"}],"recommendation":"major_revision","confidential_remarks":"The methodological core is sound and the paper may make a useful contribution to ichnology, but the central empirical claim — three morphotypes persisting across localities — rests on circular and underpowered statistics. The authors should be asked to re-analyze the specimen-level data with an independent classification method and to replace 'p > 0.1 means similar' arguments with equivalence tests or power analyses. If the morphotype result cannot survive those changes, the paper may still be publishable as a methods contribution with a more modest demonstration, but the current abstract overstates what is established."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the width-normalized discretization is a keeper; the three-morphotype story is not nailed down by the statistics as presented.\n\nWhat the paper does well: it lays out a clear, repeatable pipeline for turning fossil trails into turning-angle data, and the choice to express segment length as a multiple of trail width is a sensible way to compare specimens of different sizes. The supplementary table with every specimen's width, length, and straightness index is useful, and the authors are candid about treating speed as a constant. The direction-adjustment step for preferred turning bias is a thoughtful touch.\n\nThe soft spot is the empirical claim. The paper treats p > 0.1 from Welch's t-tests as evidence that two sets of paths are behaviourally similar. That's only legitimate with adequate power, and Wales (N=5), Poland (N=3), and Russia (variance ~15.2) do not provide it. More importantly, the three morphotypes are read off the black blocks of the same specimen-vs-specimen p-value matrix that later gets used, minus 'different' specimens, to show locality differences vanish. The thresholds (p<0.1, 25%) are arbitrary; no sensitivity analysis is reported; and with roughly nine thousand pairwise tests, some chance hits are guaranteed. The abstract's persistence claim is therefore stronger than the analysis supports.\n\nNone of that kills the methodological contribution. The discretization pipeline stands on its own, and the case study is at least a roadmap for how to misuse, or after revision properly use, movement statistics on fossil traces. But the paper would not be publishable with the current inference chain. It needs equivalence tests or power analysis, an independent clustering step for morphotypes, and sensitivity to the segment multiplier and thresholds. Also, if it claims open-source groundwork, ship the code and the coordinate data.\n\nWho is this for? Ichnologists and paleontologists who want quantitative behaviour comparisons. I'd send it to review, but with the expectation of major revision. The method is worth referee time; the case study conclusion, as written, is not.","headline":"The discretization toolkit is a real contribution; the three-morphotype persistence claim is built on circular t-test logic and needs rework before it can be believed.","tokens_in":21615,"tokens_out":2699,"would_cite":true,"duration_ms":32888,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that fossil movement paths can be discretized into width-normalized steps and turning angles, and that this quantitative method reveals three recurring behavioural morphotypes in Cruziana semiplicata that persist across loc","keywords":["trace fossils","Cruziana semiplicata","movement ecology","turning angles","behavioural morphotypes","ichnology","trilobite","quantitative paleontology"],"falsifier":"Re-analyse the p-value matrix with a multiple-comparison correction (e.g., Benjamini–Hochberg) and compute the statistical power of the pairwise t-tests under the observed variances (Russia, for example, has variance 15.20 with 12 paths). If the checkerboard clustering into morphotypes 2 and 3 does not survive correction, or if power is below conventional levels for the group comparisons, the three-morphotype persistence claim is not established.","tokens_in":20645,"feed_emoji":"🐾","tokens_out":7492,"duration_ms":86424,"temperature":0.7,"pith_summary":"Movement behaviour rarely fossilizes except as trails and trackways, and ichnologists have had few ways to compare those paths quantitatively. This paper adapts movement-ecology measures to trace fossils by discretizing a path into segments whose length is a fixed multiple of trail width, then computing turning angles at every step; with an assumed average velocity, segment length stands in for time. Applied to the trilobite trace fossil Cruziana semiplicata from Spain, Oman, Poland, Russia, and Wales, the method is used to test three older claims about the tracemaker's behaviour. The authors report that temporary resting does not change the course of the path, that one dominant and two subordinate movement morphotypes recur across localities, and that the previously suggested split of the species into two ichnosubspecies is not supported.","feed_headline":"Trilobite trail analysis finds three behaviours, not two","feed_subtitle":"Width-relative turning angles turn fossil trackways into testable behavioural data.","key_machinery":"The central object is the turning-angle distribution of a discretized path. A trail is traced from photographs into coordinates, resampled at intervals equal to a chosen multiple of the trail width (here 0.5 times width), and each triple of consecutive points is converted into one turning angle. Each specimen, subgroup, or locality is then represented by the distribution of these angles, and pairs of distributions are compared with Welch's two-sample t-tests; the resulting p-value matrices are clustered by eye into morphotypes. The turning-angle distribution carries the argument because it summarizes path morphology in a form that can be tested statistically, while the width-relative segment","core_discovery":"The central claim is that fossil movement paths carry quantifiable behavioural signal once they are discretized into width-relative steps and summarized as turning-angle distributions. On Cruziana semiplicata, this signal takes the form of three morphotypes: one dominant stereotyped movement pattern and two subordinate behavioural variants. The dominant pattern is statistically similar before and after resting traces, and it persists across five geographically separate localities once the subordinate specimens are set aside. The authors therefore conclude that locality-level behavioural differences are a sampling artefact of rare morphotypes, and that the proposed two-ichnosubspecies split i","pith_inferences":["Because segment length is tied to trail width, the method implicitly assumes movement speed scales with body size; if that scaling is wrong, the inferred 'time' axis and hence the turning-angle comparison would be biased. Testing this on modern tracemakers of different sizes would settle it.","The three morphotypes are defined from clusters in a t-test matrix rather than from an explicit statistical clustering model; a mixture-model fit to the turning-angle distributions would show whether three is the natural number of components or an artefact of thresholds.","If the dominant morphotype is truly stable across five regions within a narrow time window, the same analysis applied to other Cruziana ichnospecies should reveal similarly conserved morphotype structure, making behaviour a phylogenetically informative trait.","The paper's inferred time axis could be calibrated, not just assumed, by using speed estimates from living analogues; that would turn qualitative statements about 'temporary resting' into quantitative durations."],"forward_implications":["Ichnologists can treat trail and trackway specimens as movement datasets and apply the statistical toolkit of movement ecology to fossil behaviour.","For C. semiplicata, the path after a resting trace is statistically indistinguishable from the path before it, supporting the view of a fixed behavioural programme.","The dominant morphotype is shared across Spain, Oman, Poland, Russia, and Wales once rare subordinate specimens are set aside, so locality-level differences are a sampling effect.","The proposed two-ichnosubspecies split of C. semiplicata is replaced by three morphotypes—one common, two subordinate—that occur in varying proportions.","Because the output measures are unitless and width-relative, the same workflow can be applied across different ichnospecies, environments, and geological time intervals."],"supporting_citations":[{"why":"Source of the Spanish slabs and of the three behavioural hypotheses the study tests: fixed programme after resting, identical behaviour across four subgroups, and a possible split into two ichnosubspecies.","marker":"Seilacher, 2007"},{"why":"Supplies the Oman specimens of C. semiplicata and the identification of the tracemaker used to frame the behavioural interpretation.","marker":"Fortey & Seilacher, 1997"},{"why":"Source of the Russian specimens and a review of the spatial and temporal distribution of C. semiplicata that motivates the cross-locality comparison.","marker":"Jensen et al., 2011"},{"why":"Source of the long Polish trackway used as one of the locality groups in the group-level t-tests.","marker":"Radwański & Roniewicz, 1972"},{"why":"Source of the Welsh specimens, the straighter comparison group in the locality-level analysis.","marker":"Crimes, 1970"},{"why":"Defines the two-sample t-test that underlies every specimen-, subgroup-, and group-level comparison in the paper.","marker":"Welch, 1947"},{"why":"Supplies the movement-ecology paradigm whose four factors—internal state, navigation capacity, motion capacity, and external factors—are used to interpret the three morphotypes as behavioural variants.","marker":"Nathan et al., 2008"},{"why":"Frames the autocorrelation issue that motivates the choice of sampling interval and the use of turning-angle distributions for path comparison.","marker":"Dray et al., 2010"}],"fun_headline_variants":["Trilobite tracks show three behavioral types, not two","Quantifying fossil movement reveals three trilobite behaviors","Width-relative turns decode trilobite trail behavior","Fossil trackways: three morphotypes, not a sampling error","New method turns fossil trails into behavioral data"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The argument relies on treating p > 0.1 from small-sample t-tests as evidence that two movement patterns are the same or that behaviour persists; with as few as three to five paths in some localities and no power analysis, a non-significant p-value is weak support for that conclusion.","fun_headline_variants_meta":{"raw":{"variants":["Trilobite tracks show three behavioral types, not two","Quantifying fossil movement reveals three trilobite behaviors","Width-relative turns decode trilobite trail behavior","Fossil trackways: three morphotypes, not a sampling error","New method turns fossil trails into behavioral data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000287,"raw_usage":{"total_tokens":1494,"prompt_tokens":689,"completion_tokens":805,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":433,"completion_tokens_details":{"reasoning_tokens":725}},"tokens_in":433,"tokens_out":805,"duration_ms":9631,"temperature":1.0,"reasoning_tokens":725,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T12:52:21.433172+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-analyse the p-value matrix with a multiple-comparison correction (e.g., Benjamini–Hochberg) and compute the statistical power of the pairwise t-tests under the observed variances (Russia, for example, has variance 15.20 with 12 paths). If the checkerboard clustering into morphotypes 2 and 3 does not survive correction, or if power is below conventional levels for the group comparisons, the three-morphotype persistence claim is not established.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the Russian specimens and a review of the spatial and temporal distribution of C. semiplicata that motivates the cross-locality comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Source of the long Polish trackway used as one of the locality groups in the group-level t-tests."}],"review_version":1}