{"id":"9d9a54d4-3947-4a69-b409-7e9c533fa601","arxiv_id":"1908.00431","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"A simulation combining conflict-intensity kriging and a Markov decision process maps likely inland origins of slaves departing West African ports between 1816 and 1836.","lead":"This paper models where enslaved people in the 19th century Oyo region were probably captured, based on the port they left from. It combines conflict maps with a route-choice simulation to produce origin probability maps, but the model is not yet validated.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The normalized conflict-kriging surface is treated as the capture probability density without validation; any bias in that surface propagates into every port-conditional origin map, so the central claim rests on an unestablished proxy.","rationale":"The reader identifies the same weakest assumption: the annual conflict-intensity surface is treated as a probability density for slave capture locations without independent evidence. My stress-test agrees and narrows the concern to its most load-bearing form. The entire simulation pipeline starts with draws from this surface, so the final conditional origin maps are only as valid as this proxy. The paper's own text acknowledges the lack of optimization, the future need for linguistic validation, and the disclaimer that the maps are an approximation rather than historical truth. These passages directly corroborate the concern. The concrete test proposed here would settle the question by comparing model output to independent linguistic origin data for known ships, which is the most direct available ground truth. Since the evidence is currently insufficient to establish the central claim, the reader's REJECT verdict remains appropriate, and my read does not change it.","tokens_in":70567,"tokens_out":2640,"duration_ms":32138,"concrete_test":"Use independently transcribed slave-name and ethnicity data from the 1832 ships departing Lagos and Ouidah, the same ships shown in Figure 7, to construct observed regional origin proportions. Run the model with its stated parameters to produce predicted origin proportions for those ports and years, then compare the observed and predicted distributions using a chi-square or permutation test. If the disagreement exceeds Monte Carlo and transcription uncertainty, the conflict-kriging density is not a valid source distribution and the conditional origin maps are unsupported. If such transliterated data are not yet available in sufficient quantity, the test can be repeated on any future transcribed ship logs with known ports and years.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that the annual kriging surface in Section 3.1.3 be a valid probability density for slave capture locations. The paper states: 'we can also view the resulting surface as an implied probability density function, where the higher points of the ridge near conflict locations represent regions of increased probability of slave capture.' This is asserted, not established. Recorded conflict events, such as battles and destroyed towns, may correlate with slave raiding, but no evidence shows they are an unbiased spatial sample of capture intensity; captures during raids and marches could occur away from documented battle sites, and historical records favor towns and well-known events. Because every simulated origin is drawn from this surface by direct inversion (Section 3.3.1), any bias in this density propagates through the Markov decision process into every port-conditional origin map. The paper itself flags this gap: Section 3.3.3 admits there is 'no mathematical optimization strategy' and treats linguistic validation as future work, while Section 4 disclaims historical truth and frames the output as an approximation. The port-total data used for tuning (Section 2) constrain overall volumes, not spatial origins, so they cannot validate the spatial source distribution. Thus the unvalidated conflict-as-capture proxy is the weakest load-bearing link in the argument.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a two-stage statistical pipeline for mapping inland slave-capture origins in the Oyo region, ca. 1816-1836. In the first stage, the authors krige a conflict-intensity surface from recorded events such as battles and destroyed towns, normalize the surface to an annual density, and simulate slave capture locations from it by direct inversion. In the second stage, they feed these simulated capture locations into a Markov decision process over a manually coded trade network, with randomized point-of-sale rewards, to produce maps of the conditional probability of origin given a port of departure. The paper also describes an interactive Shiny application and frames the output as a data-driven answer to the historical question of where slaves sold at particular ports originated.","tokens_in":70780,"tokens_out":5812,"duration_ms":63919,"significance":"If the central claim held, the paper would offer a useful template for quantitative spatial inference in digital history, and the interactive visualization is a genuine public-facing contribution. The authors are transparent about several limitations, use standard tools (kriging, MDP, KDE), and clearly separate the simulation stage from the historical validation stage. However, the central claim of valid port-conditional origin maps is not supported: the capture-density proxy is unvalidated, key parameters are fixed heuristically, the reward distribution is unspecified, and the only empirical comparison targets port totals rather than spatial origins. The paper is therefore better read as a proof-of-concept than as an established answer to the research question it poses.","major_comments":[{"comment":"Section 2 is explicitly incomplete: the text states 'Section in Progress pending collaboration: Describe the conflict data - what historical accounts were used?', 'The data were collected in ??', and the trade map relies on 'CITATION MISSING'. Because the entire pipeline depends on the provenance and coding of conflict events, city locations, and network edges, the missing data description prevents replication and evaluation of the central claim.","section":"Section 2"},{"comment":"The kriged conflict surface is normalized into an 'implied probability density function' for slave capture and then used as the source distribution for all simulated slaves via direct inversion (Section 3.3.1). This is the load-bearing assumption of the paper, but it is asserted rather than validated: recorded battles and destroyed towns need not be an unbiased spatial sample of capture intensity, and captures during raids or marches could occur away from documented sites. The port-total data used for tuning (Section 2) constrain only aggregate volumes, not spatial origins, so they cannot validate this surface. The manuscript itself flags the lack of validation in Sections 3.3.3 and 4.","section":"Section 3.1.3"},{"comment":"The Matern covariance parameters are fixed heuristically ('we found that a 10km range accomplished this'; smoothness fixed at '4.5 times differentiable'), and the sill and nugget are fit only to the 1828 data and reused for all years 1816-1836. No sensitivity analysis is provided, even though the normalized surface is directly proportional to the capture density and therefore to every conditional origin map. The authors should also state the smoothness parameter nu explicitly, since '4.5 times differentiable' is not a standard parameterization.","section":"Section 3.1.3"},{"comment":"The distribution of random reward vectors is never specified, and Section 3.3.3 admits 'our model includes a considerable amount of parameters with no mathematical optimization strategy.' Since port assignment is deterministic under equal rewards (Section 3.3.1), the stochastic port catchments shown in Figure 5 are entirely driven by the unspecified reward variance and the conflict cost C. Without specifying these distributions and performing optimization or systematic sensitivity analysis, the reported conditional probabilities are not reproducible and their uncertainty is not quantified. The proposed chi-square tuning relies on linguistic data that are not yet available, so it remains future work.","section":"Sections 3.3.1 and 3.3.3"},{"comment":"The conflict surface enters the model twice: as the capture density and as an additive edge cost, with C scaled to an annual maximum of 3. This is not circular by definition, but it means the same unvalidated proxy drives both the numerator and the routing denominator of the conditional maps, potentially amplifying artifacts. A concrete test would be to rerun the pipeline with C=0 or with C estimated from independent route-choice evidence and compare the port-conditional origin maps.","section":"Section 3.2.3"}],"minor_comments":[{"comment":"Placeholder text remains in the manuscript, including 'CITATION MISSING', 'website.com', 'maybe cite slavebiographies.org', and 'The data were collected in ??'; these must be completed before any resubmission.","section":"Throughout"},{"comment":"Figure 6 is referenced but no figure or quantitative summary is included in the text; the authors should include the figure with error bars, sample sizes, and a description of the fit.","section":"Figure 6"},{"comment":"The text refers to a '2-valued marker for intensity of battle' while Section 2 defines four intensity levels (0, 1, 5, 10); please clarify which variable is actually modeled.","section":"Section 3.1.1"},{"comment":"The discount factor is set to gamma = 1 even though gamma in [0,1) was defined earlier; if the horizon is finite, state this explicitly.","section":"Section 3.2.1"},{"comment":"The caption says 'increasing variance in rewards' but no numerical variances are given; please report the actual reward distributions used.","section":"Figure 5"},{"comment":"There are several typographical and notation inconsistencies, including 'M´tern' for Matern and 'Translatlantic' for Transatlantic; a careful proofread is needed.","section":"Throughout"}],"recommendation":"reject","confidential_remarks":"This manuscript is an incomplete draft: Section 2 is explicitly marked 'in progress', citations and figures are missing, and the core validation gap is not a local fix. Establishing that conflict events are a valid capture-intensity proxy, or replacing that assumption, would require new historical data and a substantial re-analysis. I would be open to reconsidering a revised version that completes the data description, specifies all parameter distributions, and provides sensitivity analyses against alternative source models."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a look if you work on spatial models in digital humanities, but this is clearly a draft. The new thing here is the pipeline: kriged conflict intensity, simulated capture locations, a Markov decision process over a trade network, and port-conditional origin maps. I don't know of another paper that connects slave origins to departure ports through an explicit decision model, and the interactive app idea is a sensible output for historians. The writing is honest; Section 3.3.3 flatly says there is no mathematical optimization strategy, and Section 4 disclaims historical truth. That candor counts. \n\nThe soft spot is load-bearing. Simulated capture locations are drawn from the normalized conflict kriging surface (Section 3.1.3), and that surface is asserted to be an implied probability density of slave capture. There is no evidence that recorded battles and destroyed towns are an unbiased spatial sample of capture intensity. Raids and captures along marches could easily be off the documented battle sites. Because every simulated origin is drawn from this surface, any bias propagates into every port-conditional map. The port totals used for tuning constrain volumes, not spatial origins, so they cannot validate the source distribution. The paper itself flags the missing validation via linguistic data as future work, which is honest but does not fill the gap.\n\nOther issues: parameters like the Matern range, smoothness, conflict cost multiplier C, and reward variance are hand-set with no sensitivity analysis; no code or data are provided; the Data section is literally marked “Section in Progress”; and some citations are missing. In its current form the central claim that these maps are meaningful conditional probabilities of origin is not supported.\n\nWho is this for? A historian or digital-humanities researcher interested in the modeling template might get something from the setup. A statistician will want to see the proxy validated and parameters estimated before taking the maps seriously. I would not cite it yet, and I would not send this version to referees. The right move is to tell the authors to finish the data section, release code and data, add sensitivity analysis, and ideally validate against the linguistic name data they mention. If they do that, the pipeline deserves a serious look.","headline":"A promising but unfinished preprint: the modeling pipeline is new, the validation is absent, and the central conflict-as-capture proxy is unproven.","tokens_in":71337,"tokens_out":1703,"would_cite":false,"duration_ms":20508,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-14T15:56:48.225557+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}