REVIEW 4 major objections 6 minor 2 references
Forecasting e-scooter substitution of direct and access trips by mode and distance
T0 review · 4 major / 6 minor · reviewed 2026-08-14 · deepseek-v4-flash
Pith's one-line read E-scooter substitution of Manhattan trips is forecast at 75,000 daily rides, mostly replacing short carpool, bike, and taxi trips.
desk verdict The paper's 75K daily Manhattan e-scooter forecast is internally inconsistent with the pilot data used to fit the model, but the multi-city demand framework and the distance-based mode-substitution decomposition are genuinely new and worth engaging with. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a two-stage statistical decomposition. Stage one is a log-log trip-generation regression (Eq. 1) that converts zip-code demographics and a citywide fleet-size variable into predicted e-scooter trips per zone. Stage two is the nonlinear multifactor model (Eq. 4): it treats the predicted trips as the dependent variable and, as independent variables, observed trips by mode and distance from Manhattan's travel survey, multiplied by two parameter sets — $F_m$, the fraction of each mode's trips that e-scooters could replace (fixed across distance), and $P_d = \beta_d / \delta_d$, a distance-decay competition probability that shrinks as the average trip distance $\delta_d$ grows. A separate access-trip factor $F_{transit,i} = \beta_5 t_i^{access} + \beta_6 t_i^{egress}$ captures e-scooter substitution for first- and last-mile transit access. Fitting both parameter sets by least squares, with bootstrap confidence intervals, reveals which modes are statistically similar to e-scooter trips and how that similarity decays with distance.
What would settle it
A direct test is to compare the forecast with realized operations: a Manhattan e-scooter program capped at 2,000 vehicles should show roughly 75,000 daily trips; if the observed count falls outside the model's out-of-sample error range (about ±20% on log trips, coefficient of variation 0.27), the transferability assumption fails. A second test is to survey riders' prior mode and check whether the shares replacing carpool, bike, and taxi are close to 32%, 13%, and 7.2%.
Extended reading notes
Core claim
The paper's central claim is that Manhattan has a large latent e-scooter market that can be quantified even before local ridership data exist. Using a log-log regression $$\ln R_i = \beta_0 + \beta_P \ln(\text{Population}_i \times \text{AgeRatio}_i) + \beta_L \ln(\text{LandArea}_i) + \beta_S \ln(\text{Scooters}) + \varepsilon_i$$ estimated on 60 zip-code observations from three pilot cities, the paper predicts 75,000 daily e-scooter trips for a 2,000-scooter Manhattan fleet. A second, nonlinear multifactor model $$R_{esco,i} = C + \sum_m F_m \sum_d P_d N_{m,i,d} + \sum_d (1-P_d) F_{transit,i} N_{transit,i,d} + \gamma_i$$ with $P_d = \beta_d / \delta_d$ attributes those trips to existing modes, finding statistically significant substitution from carpool ($F=0.636$), bike ($F=0.226$), and taxi ($F=0.184$), with the distance-competition term falling from 0.986 at 0.5 miles to 0.09 at 5.5 miles. The model also estimates an access/egress factor for public transit, $F_{transit,i} = \beta_5 t_i^{access}$, implying e-scooters can replace a share of first- and last-mile transit access trips. Taken together, the paper claims e-scooters would systematically replace short carpool, bike, taxi, and transit-access trips rather than merely add new travel.
Load-bearing premise
The load-bearing assumption is that Manhattan riders will respond to scooter supply the way riders in the three pilot cities did; the paper's closing remarks list omitted factors such as weather, trip chaining, and spatial correlation, so a different response in New York would change the 75,000-trip forecast and every substitution share.
Editorial extensions
If this is right
- A 2,000-scooter Manhattan fleet would generate about 75,000 daily trips and about $77 million in annual revenue at a $1 plus $0.15 per minute fare.
- E-scooter competition is a short-distance phenomenon: the competition probability drops from 0.986 at 0.5 miles to 0.09 at 5.5 miles, so the mode would serve a last-mile niche rather than long trips.
- Carpool is the most exposed mode, with up to 32% of carpool trips replaceable, followed by bike at 13% and taxi at 7.2%; walking and auto substitution are small.
- E-scooters would draw not only direct trips but also access/egress trips to public transit, and revenue from that access segment is nearly flat with distance, unlike direct-trip revenue.
- Because e-scooters are dockless, they would reach Manhattan areas that station-based Citi Bike does not cover, helping explain the forecast's 60% premium over Citi Bike ridership.
Reading between the lines
- A testable extension is to apply the same two-stage structure to another city with pilot data; the fleet-size coefficient of 0.781 implies regulators can use scooter caps to tune total trip volume, not just placement.
- The strong distance decay suggests operators could price by distance rather than by time, or concentrate parking and rebalancing on sub-2-mile trips, to flatten the revenue curve.
- Since transit-access substitution appears nearly distance-flat, transit agencies could treat e-scooters as a stable first- and last-mile feeder whose revenue does not erode on longer transit trips.
- Because the fleet-size effect is identified from only three city-level observations, the exact 75,000 figure is conditional on that comparison; the modal substitution shares are the more portable insight.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a two-stage model to forecast e-scooter demand in Manhattan and attribute that demand to substituted travel modes. In the first stage, a log-log regression (Eq. 1) is estimated on 48 zip-code observations from Portland, Austin, and Chicago, using population×age ratio, land area, and fleet size as covariates. Applying this model to Manhattan with an assumed 2,000-scooter fleet yields a forecast of 75,000 daily e-scooter trips (Section 5.1). The second stage fits a nonlinear multifactor model (Eq. 4) that decomposes these 75,000 trips into direct substitutions from bike, walk, carpool, taxi, auto, and public transit, plus access/egress substitutions to transit, as a function of distance. The model reports, for example, that e-scooters could replace 32% of carpool trips, 13% of bike trips, and 7.2% of taxi trips, and it projects $77 million in annual revenue (Section 5.4). The paper positions these as the first multi-city e-scooter demand forecast with a fleet-size elasticity and a distance-based mode-substitution decomposition.
Significance. If the forecast were reliable, the paper would provide a useful planning tool for cities considering e-scooter programs and would extend the micromobility literature by explicitly modeling substitution from multiple modes and access/egress trips. The authors are to be credited for assembling multi-city pilot data, including a fleet-size variable, and for providing a bootstrap procedure for the second-stage parameters. However, the central quantitative claims are not currently supported: the first-stage model explains only 31% of variance, the forecast implies per-scooter utilization an order of magnitude above the pilot cities used for estimation, and the second-stage model is fitted to the same forecast it is supposed to explain. As a result, the headline numbers (75K daily trips, $77M revenue, and the substitution shares) are not credible in their present form.
major comments (4)
- [§4.1, Eq. (1), Table 5] The first-stage trip generation model is too weak to support the 75K daily forecast. The regression has R²=0.314 (adjusted 0.267) from only 48 training records, and the 'Scooters' variable is constant within each city, so its coefficient (0.7812) cannot be separated from city-specific unobservables such as pilot program design, regulatory environment, or marketing effort. The negative land-area elasticity (−1.108) then extrapolates to Manhattan's much smaller zip codes (average 0.5 sq mi vs 9.71 in Portland, 18.51 in Austin, 4.12 in Chicago; Table 1), far outside the estimation range. The out-of-sample validation in Section 4.2 is limited to 12 records and a single Hoboken observation with an error of 0.42, which is not sufficient to establish transferability. The 75K point forecast therefore rests on an extrapolation from a poorly identified, weakly fitting model.
- [§5.1 and §5.4] The headline forecast of 75K daily e-scooter trips with 2,000 scooters implies 37.5 trips per scooter per day. The paper's own pilot data give 2.9 trips per scooter per day in Portland (700K trips over about four months with 2K scooters, Section 3.1) and roughly 2.1–2.7 in Chicago (821K trips with 2.5K scooters), so the Manhattan forecast is 13–18 times the utilization observed in the estimation cities. With the paper's assumption of 12-minute average trips (Section 5.4), this requires 7.5 hours of riding per scooter per day before rebalancing, charging, and deadheading, which is not plausible for a shared dockless fleet. This is an internal inconsistency, not merely a transferability caveat: the model applied to Manhattan produces a per-scooter productivity that contradicts the very data used to estimate it. The $77M annual revenue estimate inherits this implausibility.
- [§4.4, Eq. (4), Table 6] The multifactor model uses the forecasted 75K trips as its dependent variable, so the reported substitution shares (32% carpool, 13% bike, 7.2% taxi) are in-sample fitted quantities that mechanically reproduce the first-stage forecast; they are not independent predictions. The bootstrap procedure resamples the 318 TAZ observations conditional on the fixed forecast, so it does not propagate the substantial first-stage estimation error. Moreover, the constant C=73.365 alone contributes about 23K trips (Section 5.2), meaning a large share of the decomposition is attributed to an unmodeled constant rather than to mode substitution. Several mode factors are not statistically significant (F_walk t=0.350, F_citi_bike t=1.093, F_auto t=1.677), which weakens the conclusion that carpool, bike, and taxi are the dominant substituted modes. These issues undermine the second-stage claims as presented.
- [§4.4, Eq. (2), Table 7] The distance-decay relationship P_d = β_δ / δ_d is imposed by assumption rather than estimated from the data; only the single parameter β_δ is calibrated. Consequently, the claim that the substitution probability 'drops by an order of magnitude' from 0.5 to 5.5 miles (Table 7) follows mechanically from the chosen functional form and is not a data-driven finding. The paper should at least test alternative functional forms or acknowledge that this structure is a modeling assumption, and it should be cautious in presenting the distance-based substitution shares as empirical results.
minor comments (6)
- [§4.2, Figure 5] The text refers to 'Mean Absolute Error' while the figure caption says 'Mean Absolute Deviation'; please make the terminology consistent.
- [§4.2 and Table 1] The city of Austin, Texas is referred to as 'Texas' in several places (e.g., 'Texas data points have relatively higher error bound' and the Table 1 header). Please use 'Austin' consistently.
- [Figure 4] Figure 4(a) is described as a q-q plot of e-scooter ridership in Portland, but the model is estimated on pooled data from three cities; please clarify which residuals or data are shown in the figure.
- [Eq. (6) and Table 6] The notation for the access-time coefficient is inconsistent: Eq. (6) uses β_5 while Table 6 reports β_6=0.493 and also lists β_7, β_8, β_9 with N/A t-stats. Please align the subscripts and explain the zero-valued coefficients.
- [§4, §5.2, and Abstract] The paper states in Section 4 that the model 'is not meant to make behavioral predictions of which modes would be substituted by e-scooter,' yet the abstract and Section 5.2 present 'could replace 32% of carpool, 13% of bike, and 7.2% of taxi trips.' This tension between the stated scope and the reported conclusions should be resolved explicitly.
- [§4.1, Table 5] The text says 'All variables are substantially significant,' but the constant has p=0.0949 (significant only at the 10% level); please phrase this more precisely.
Circularity Check
Substitution shares are in-sample fits to the first-stage 75K forecast, while the 75K demand forecast itself is a genuine out-of-sample extrapolation.
-
fitted input called prediction
[Section 4.4 (Eqs. 4-5) and Section 5.2 (Table 6)]
"4.4: 'The dependent variables are the forecasted e-scooter trips per zone while the independent variables are the existing trips by mode.' Eq. (5): 'min_z=sum (R_escooter,i − R_fitted,i)^2'. 5.2: 'e-scooters can replace, in total, up to 32% of carpool, 13% of bike, 7.2% of taxi, 1.9% of walking, and 1.8% of auto trips.'"
The multifactor model's dependent variable is the same R_escooter,i forecast from Eq. (1) (Section 4.4), and Eq. (5) fits F_m and P_d to minimize the gap to that forecast. The 'replace' percentages in Section 5.2 are products of the fitted F_m and P_d, so they are in-sample calibrated quantities, not independent predictions. The fitted R_fitted is forced by the objective to approximate the forecasted demand it is meant to explain; the substitution breakdown therefore inherits the 75K forecast and cannot confirm it. The external modal trip counts provide descriptive content, but the forecasting language overstates the evidential status.
full rationale
The first-stage 75K daily trip forecast is a genuine out-of-sample extrapolation of Eq. (1) from Portland, Austin, and Chicago to Manhattan, with a weak but real out-of-sample check against 12 held-out zip codes and a Hoboken observation; that part is not circular. The circularity concern attaches to the substitution analysis: Section 4.4 defines the multifactor model's dependent variable as the forecasted e-scooter trips from Eq. (1), and the objective function fits the mode factors and distance-decay parameters to reproduce that same forecast. The reported substitution percentages (32% carpool, 13% bike, 7.2% taxi) are therefore in-sample fitted quantities, not independently predicted outcomes. This is partial circularity rather than full circularity because the modal trip counts themselves are external data from the NYMTC RHTS and Citi Bike, and the paper sometimes carefully uses 'estimates' rather than 'predicts' for the substitution shares. Self-citations such as He et al. (2020) are used to source or validate the Manhattan modal trip inputs, but they are not invoked as a uniqueness theorem or as a substitute for estimation, so they are not load-bearing circularity. The implausibly high per-scooter utilization implied by 75K daily trips with 2,000 scooters is a serious correctness risk, but it is an external-validity/sanity-check concern, not a circularity argument.
Assumptions & free parameters
free parameters (9)
- beta_const =
-7.929
- beta_P (population*age ratio elasticity) =
0.573
- beta_L (land area elasticity) =
-1.108
- beta_S (fleet size elasticity) =
0.7812
- Fleet size assumption for NYC =
2000 scooters
- P_d distance-decay coefficient beta_delta =
0.493
- Mode substitution factors F_m =
bike 0.226, walk 0.021, citibike 0.094, carpool 0.636, taxi 0.184, auto 0.109, transit direct 0.000
- Constant C in multifactor model =
73.365
- Access/egress coefficients =
reported as beta 0.069 or 0.493 (table inconsistent)
assumptions (7)
- ad hoc to paper Log-log linear functional form for e-scooter trip generation (Eq. 1)
- ad hoc to paper Distance-decay form P_d = beta_delta / delta_d (Eq. 2)
- ad hoc to paper Additive decomposition of e-scooter demand into direct and access substitutions (Eq. 4)
- domain assumption 2010/2011 NYMTC RHTS modal trip data remain representative of Manhattan's trip landscape
- domain assumption Zip-code to TAZ allocation by population shares
- domain assumption Average trip duration of 12 minutes (1.6 miles) for revenue calculation
- domain assumption E-scooters only substitute access/egress for public transit, not access trips to other modes
Cite this review
Pith. "Pith review of Forecasting e-scooter substitution of direct and access trips by mode and distance." pith.science (2026). https://pith.science/paper/N7MGJCAP
@misc{pith2026190808127,
author = {Pith},
title = {Pith review of: Forecasting e-scooter substitution of direct and access trips by mode and distance},
year = {2026},
howpublished = {\url{https://pith.science/paper/N7MGJCAP}},
note = {Machine review of arXiv:1908.08127}
}
read the original abstract
An e-scooter trip model is estimated from four U.S. cities: Portland, Austin, Chicago and New York City. A log-log regression model is estimated for e-scooter trips based on user age, population, land area, and the number of scooters. The model predicts 75K daily e-scooter trips in Manhattan for a deployment of 2000 scooters, which translates to 77 million USD in annual revenue. We propose a novel nonlinear, multifactor model to break down the number of daily trips by the alternative modes of transportation that they would likely substitute based on statistical similarity. The model parameters reveal a relationship with direct trips of bike, walk, carpool, automobile and taxi as well as access/egress trips with public transit in Manhattan. Our model estimates that e-scooters could replace 32% of carpool; 13% of bike; and 7.2% of taxi trips. The distance structure of revenue from access/egress trips is found to differ from that of other substituted trips.
Figures
Reference graph
Works this paper leans on
-
[1]
INTRODUCTION “Micromobility services” is a relatively new term (in the context of urban mobility) defined to encapsulate the set of small vehicle shared mobility modes including electric scooters (e-scooters), docked and dockless shared bikes, electric skateboards, electric mopeds, and electric pedal-assisted (pedelec) bikes (see Zarif et al., 2019). The ...
work page 2018
-
[4]
METHODOLOGY Because e-scooter ridership data in New York are not readily available to the public, the research design involves a two-step approach: (1) estimate and apply a demand model to forecast the potential demand in different zones within Manhattan; (2) estimate a nonlinear multifactor model to break down that demand into different modal trips such ...
work page 2018
Reviewed August 14, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.