{"id":"8f3de682-e8e7-4374-a2f3-a3038ea3e459","arxiv_id":"2508.06513","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Perceived street safety is positively tied to restaurant ratings in Washington, DC, and this tie weakens as neighborhood car dependency rises.","lead":"In Washington, DC, restaurants with better-reviewed indoor photos and safer-looking streets tend to get higher Yelp ratings, but the street safety link weakens in neighborhoods with higher car dependency. The study is a reminder that improving sidewalks and streetscapes may pay off more in walkable areas than in car-oriented ones.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that car dependency moderates the safety–rating effect rests on a neighborhood proxy whose construct validity is explicitly conceded; the reported odds-ratio for safety is also internally inconsistent (87.18 for a 1-unit change vs. 8.7× for 0.1).","rationale":"The reader's conditional verdict is appropriate. The statistical finding is plausible and the model is fit with standard diagnostics, but the central claim depends on the construct validity of the Car Dependency Index, which the authors explicitly concede is limited. The reported odds-ratio interpretation (8.7× for a 0.1-point increase) is inconsistent with Table 3 (OR=87.18 for a 1-unit change), indicating a scaling ambiguity that affects how readers interpret the main street-level effect. The sign reversal at high CDI is an important, non-obvious implication that the Discussion does not explicitly address, and the Discussion's statement that streetscape quality remains important in car-dependent areas appears to contradict the negative total effect at high CDI. These issues are fixable: the model can be re-estimated with a trip-level car-access variable for a validation subset, and the safety scale can be clarified or re-percentaged. Given that the core analysis is potentially sound but the key moderator and effect-size reporting are not yet secure, CONDITIONAL is the right verdict. My concern agrees with the reader's weakest-assumption identification: the Car Dependency Index may not measure actual car use.","tokens_in":16280,"tokens_out":2243,"duration_ms":24900,"concrete_test":"Re-estimate the ordinal logistic model after replacing CDI with a trip-level measure of car access for a subset of EDEs, e.g., using the Advan Research foot-traffic data to compute the share of visitors whose home census tract has high car modal share or whose trip distance exceeds a threshold (e.g., 5 km), then test whether the interaction term (Perceived Safety × actual car-access share) remains negative and significant. If the interaction disappears or reverses, the conclusion that streetscape effects diminish for car-dependent customers is not supported. Also re-run the model with CDI dichotomized at the median to check robustness, and recalculate the safety odds ratio using the correct unit (1.0 or 0.1, depending on the actual coding of the safety score) to resolve the factor-of-ten discrepancy.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central empirical claim is that higher car dependency weakens (and eventually reverses) the positive effect of perceived street safety on Yelp ratings. This claim depends on the Car Dependency Index (CDI) actually measuring whether customers access EDEs by car. CDI is built from car modal share, population density, and employment density of census block groups within 1 km of the EDE: CD = 0.5×MS + 0.5×(1−PED). The authors themselves state in the Discussion that this index 'may not fully capture whether visitors actually use cars to access EDEs' and that it omits trip origin and trip distance. If CDI is effectively a suburban/urban location proxy, the interaction term could reflect unmeasured neighborhood attributes (parking supply, transit quality, land-use mix) rather than the travel behavior of actual customers. A second, independent problem is the safety effect size: Table 3 gives coefficient 4.468, odds ratio 87.18, but the Results text says a 0.1-point increase raises the likelihood of a higher rating category by 8.7 times. With the safety variable described in Table 2 as mean 6.0, SD 0.3, range 4.7–6.6, a 0.1-unit increase cannot plausibly produce OR 87.18; the authors appear to conflate a 0.1-point change on a 0–1 scale with 0.1 unit on their actual scale. This inconsistency undermines the headline interpretation and should be corrected before the finding is accepted. The sign reversal at CDI ≈ 66 is a striking implication that the Discussion does not address; instead, it claims streetscape quality remains important in car-dependent areas, which is difficult to reconcile with a negative total effect at high CDI.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks whether the association between perceived streetscape safety (measured by a computer vision model applied to Google Street View imagery) and Yelp ratings of eating and drinking establishments (EDEs) in Washington, DC, is moderated by neighborhood car dependency. The authors build a Car Dependency Index (CDI) from adjacent census block groups' car modal share and population/employment density, then estimate an ordinal logistic regression with a perceived-safety × CDI interaction. They report that higher car dependency weakens the positive safety–rating association, with the interaction coefficient -0.068 (p < 0.001). The paper also includes indoor aesthetics from Yelp photos, WalkScore, foot traffic, and several controls.","tokens_in":16635,"tokens_out":5044,"duration_ms":56127,"significance":"If the moderation result holds, it would be a useful contribution to servicescape and walkability research by showing that the value of streetscape improvements depends on the transport context, with implications for context-sensitive planning. Strengths include the integration of multiple novel data sources (Yelp review photos, Street View imagery, mobile-device foot traffic), the explicit testing of the proportional odds assumption and multicollinearity, and a candid limitations section. However, the central empirical claim is currently undermined by an arithmetic/scale error in the reported safety effect, an unresolved construct-validity issue with the CDI, and the paper's silence on the sign reversal implied by its own interaction model. These issues are load-bearing and require revision.","major_comments":[{"comment":"The reported effect of perceived safety is internally inconsistent. The text states that a 0.1-point increase in the perceived safety score (described as ranging from 0 to 1) raises the likelihood of a higher rating category by 8.7 times, but exponentiating the coefficient 4.468 times 0.1 gives about 1.56, not 8.7. In addition, Table 2 reports Perceived Safety with mean 6.0, SD 0.3, and range 4.7–6.6, which contradicts the claimed 0–1 scale and makes a 0.1-point change far smaller than the reported SD. The odds ratio of 87.18 in Table 4 corresponds to a one-unit change on the actual scale (if the coefficient is 4.468 per unit), not a 0.1-point change. Please correct the scale description, the odds-ratio interpretation, and the corresponding statement in the Highlights, and report an easily interpretable effect size (e.g., per SD or per realistic increment).","section":"Section 4 (Street-level effects)"},{"comment":"The Car Dependency Index is built from neighborhood-level averages (car modal share, population density, employment density of adjacent block groups), and the authors concede that it may not capture whether visitors actually use cars to access the EDE, nor trip origin or trip distance. Since the central claim is that car dependency moderates the safety–rating relationship, this construct-validity limitation is load-bearing. The interaction could be driven by unmeasured neighborhood attributes correlated with CDI, such as parking supply, transit accessibility, or land-use mix, rather than by the actual travel behavior of customers. To support the interpretation, please add a robustness check controlling for plausible neighborhood confounders, or validate CDI against observed visitor travel behavior using the Advan foot-traffic data that the paper already uses for visitor income.","section":"Section 3.1.4 and Section 5"},{"comment":"The model's interaction implies a total effect of perceived safety of 4.468 − 0.068 × CDI, which becomes negative when CDI exceeds approximately 65.7. Given that CDI ranges from 1.1 to 99.1 with a mean of 54.6, a substantial share of sampled EDEs fall in the region where the model predicts safety has a negative association with ratings. The paper does not acknowledge or discuss this sign reversal, despite its direct relevance to the policy recommendations about improving streetscape quality in car-dependent areas. Please address this explicitly, for example by plotting marginal effects of safety across the CDI range and discussing whether a negative safety effect is plausible or likely an artifact of extrapolation or confounding.","section":"Section 4 and Section 5"},{"comment":"The definition of the Car Dependency Index is ambiguous. The text defines PEDi as the sum of the normalized population density and employment density, but the formula CD_i = 0.5 × MS_i + 0.5 × (1 − PED_i), with PED as a sum of two 0–1 variables, can produce negative values and would not map to a 0–100 range after multiplying by 100. The reported sample range of 1.1–99.1 suggests that PED was actually an average of the two densities or that the normalization was different. Because CDI is the key moderating variable, please clarify the exact construction, including how each density was normalized and how the final index was scaled to lie in the reported range.","section":"Section 3.1.4"}],"minor_comments":[{"comment":"The variable names are inconsistent between Table 2 and Table 3: Table 2 lists 'Neighborhood Income level (10k)' while Table 3 uses 'Income Level of Neighborhood'; please harmonize the terminology.","section":"Section 3.1.2"},{"comment":"The perceived safety scores from Hwang et al. (2023) are described as TrueSkill scores, but the paper does not explain how these scores were transformed into the reported range of 4.7–6.6. Please provide the transformation or rescaling details.","section":"Section 3.1.3"},{"comment":"There is a grammatical error in the final paragraph of the Discussion: 'this study provides suggests targeted interventions' should be 'this study suggests targeted interventions.'","section":"Section 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is more applied urban analytics than core physics/society, but it may still fit the journal's interdisciplinary scope. The self-citation pattern is not excessive. The main concern is that the headline interaction result is currently obscured by the safety-effect reporting error and the unresolved CDI construct validity; the authors should be encouraged to address these before publication. The sign-reversal issue is particularly important for the paper's own policy conclusions."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The useful thing here is the interaction between car dependency and perceived street safety, which I don't think is in the cited literature. The authors test it with an ordinal logistic on 744 Yelp-scraped EDEs in DC, using computer vision for street safety and indoor aesthetics. That's a legitimate extension, and the data work looks honest: they report a Brant test for proportional odds, VIFs, and the standard errors don't look fishy. If the interaction holds up, it gives planners context-sensitive guidance about when streetscape investment matters for local business ratings.\n\nThe problems are mainly in the reporting and interpretation. The Highlights say a 0.1-point increase in safety raises the likelihood of a higher rating by 8.7 times, but the Table 3 coefficient is 4.468, and exp(4.468*0.1) is about 1.56, not 8.7. They also describe safety as ranging from 0 to 1 while Table 2 shows mean 6.0, SD 0.3, range 4.7–6.6. That's not a harmless typo; it's the headline effect.\n\nMore substantively, the model itself implies that above a Car Dependency Index of roughly 66 (from 4.468/0.068), the total effect of safety on ratings is negative. The paper never acknowledges this; instead the Discussion says streetscape quality remains important in car-dependent areas, which is hard to reconcile with a negative effect at high CDI. This needs a direct address, because it could mean the interaction is too strong or the model is picking up something else.\n\nThird, the CDI is a neighborhood proxy built from car modal share and density, and the authors concede in the limitations that it may not capture whether visitors actually drive. That's load-bearing for the moderation claim. I agree with the reader that this is the weakest point, and it deserves more than a sentence in limitations—ideally a robustness check with trip-level or origin-destination data, or at least an acknowledgment that the result is about neighborhood context, not individual travel behavior.\n\nCredit where earned: they test model assumptions, use a validated aesthetic model, control for visitor income and footfall, and are candid about computer vision bias. Self-citation is not a real problem here; the safety model is their own validated tool, and using it is appropriate.\n\nThis paper will interest urban planners and design researchers working on streetscapes and business vibrancy. It deserves a serious referee, but I'd send it back for major revision before publication. I would not cite it in its current form, mainly because of the odds-ratio error and the undiscussed sign reversal.","headline":"A useful moderation hypothesis in an otherwise solid empirical paper, but the safety odds ratio is misreported and the car-dependent reversal is left undiscussed.","tokens_in":17193,"tokens_out":2417,"would_cite":false,"duration_ms":29590,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that the positive relationship between perceived streetscape safety and Yelp ratings for eating and drinking establishments weakens as the neighborhood around the establishment becomes more car-dependent, and at high car…","keywords":["Walkability","Perceived Safety","Car Dependency","Servicescape","Point-of-Interests","Computer Vision","Yelp Ratings","Ordinal Logistic Regression"],"falsifier":"Observe actual arrival modes at the sampled establishments, for example from mobile-location traces or parking utilization, and re-run the ordinal logistic model with that real car-access measure replacing the neighborhood Car Dependency Index; the central moderation claim would be falsified if the safety interaction is not negative when actual car access is used.","tokens_in":16083,"feed_emoji":"🚗","tokens_out":8573,"duration_ms":95447,"temperature":0.7,"pith_summary":"The paper tests whether the visual quality of the street around a restaurant, bar, or café still shapes customer ratings when people mostly drive there. Using 744 Yelp-listed eating and drinking establishments in Washington, DC, it pairs user review photos with Google Street View imagery, scores both with computer vision models, and models Yelp rating categories with ordinal logistic regression. It finds that indoor aesthetics and perceived street safety are both positively linked to ratings, but the safety effect shrinks as a neighborhood-level Car Dependency Index rises; at high car dependency the effect turns negative. This matters because it suggests that the value of streetscape investment depends on the dominant travel mode of the area, not on the streetscape alone.","feed_headline":"Safe streets lift Yelp ratings—until car dependency takes over","feed_subtitle":"Indoor looks and street safety both predict ratings, but the safety effect fades and flips as car dependency rises.","key_machinery":"The load-bearing object is the Car Dependency Index (CDI), a neighborhood score for each establishment computed as 0.5 times the normalized commuting car modal share plus 0.5 times the inverse of the normalized sum of population and employment density across block groups within a 1-km buffer. The decisive term is the interaction Perceived Safety × CDI in an ordinal logistic regression of the five-category Yelp rating; this term carries the paper's claim that the safety effect depends on how car-oriented the surrounding area is. Two computer vision pipelines supply the perceptual inputs: one scores interior visual appeal from Yelp review photos using a model trained on aesthetic ratings, and one scores perceived safety from Google Street View images using a model trained on crowdsourced safety judgments.","core_discovery":"The central discovery is a statistical moderation: the positive association between the perceived safety of the surrounding streetscape and an eating establishment's rating score weakens as the car dependency of its neighborhood increases. In the main ordinal logistic model, the Perceived Safety coefficient is 4.468 (p<0.001) and the interaction with the Car Dependency Index is -0.068 (p<0.001), so the safety advantage declines by about 0.068 log-odds per unit of car dependency. The paper also reports positive associations for interior perceived aesthetics and for WalkScore, showing that both indoor and outdoor visual qualities predict higher rating categories, with the outdoor effect conditional on car dependency.","pith_inferences":["A consequence the authors do not state is the crossover point: solving $4.468 - 0.068 \\times \\text{CDI} = 0$ gives CDI $\\approx 65.7$, so perceived safety's total effect on ratings turns negative above roughly that index value.","The same logic would predict that other destination types, such as shops, services, and offices, also show a diminished streetscape premium in car-dependent neighborhoods; this is testable with the same Yelp and street-view pipeline applied to other point-of-interest categories.","The moderation could reflect selection rather than perception: drivers sort into car-oriented destinations, so their satisfaction is less tied to the walking environment; using the foot-traffic origin data to construct actual arrival-mode shares would separate these mechanisms.","In practice, this suggests that street investments in car-oriented districts may need to be paired with parking or transit-access changes before they move customer ratings, while in walkable districts the streetscape itself is the lever."],"forward_implications":["In walkable parts of a city, a one-point higher WalkScore is associated with a 3% higher odds of being in a better rating category, so pedestrian-friendly location itself appears to carry a satisfaction premium.","Perceived indoor aesthetics consistently predict higher ratings, so interior design quality matters across all car-dependency contexts.","Perceived street safety raises the odds of a higher rating category by a large factor at low car dependency, but this advantage shrinks as CDI rises, meaning the same streetscape improvement will not produce the same rating gain in a walkable district and a car-oriented one.","In highly car-dependent areas, the total association between perceived safety and ratings becomes negative, implying that safe streets alone will not raise customer satisfaction there.","For planning practice, the paper implies that street-level interventions should be calibrated to the local transportation context rather than applied uniformly across a city."],"supporting_citations":[{"why":"Establishes the servicescape concept that frames the paper's indoor-plus-outdoor environment model.","marker":"Bitner, 1992"},{"why":"Provides the DINESCAPE scale for dining environments that motivates the indoor aesthetic measurement.","marker":"Ryu & Jang, 2008"},{"why":"Extends servicescapes to streetscapes and links streetscape quality to local business attractiveness, the relationship this paper conditions on car dependency.","marker":"Koo et al., 2023"},{"why":"Documents that dominant travel mode shapes how people perceive the built environment, the premise for the safety-by-car-dependency interaction.","marker":"Handy, 2005"},{"why":"Supplies the urban design quality framework used to justify measuring walkability and streetscape perception.","marker":"Ewing & Handy, 2009"},{"why":"Supplies the Place Pulse 2.0 crowdsourced safety-judgment dataset used to train the perceived-safety model.","marker":"Dubey et al., 2016"},{"why":"Builds the street-view safety scoring model that produces the Perceived Safety variable.","marker":"Hwang et al., 2023"},{"why":"Validates the aesthetic model on restaurant review images, supporting the interior aesthetics score's use.","marker":"Pan et al., 2024"}],"fun_headline_variants":["Streetscape safety boosts Yelp ratings—unless streets are car-first","In car-dependent DC, streetscape safety matters less for Yelp scores","Yelp ratings: streetscape safety fades where cars dominate","Car dependency tempers the streetscape boost to restaurant ratings","Streetscape effect on Yelp ratings shrinks in car-centric neighborhoods"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The Car Dependency Index, a neighborhood score built from commuting mode share, population density, and employment density, is assumed to reflect whether customers actually reach the establishment by car, even though it ignores the visitor's origin, trip distance, and actual mode of arrival.","fun_headline_variants_meta":{"raw":{"variants":["Streetscape safety boosts Yelp ratings—unless streets are car-first","In car-dependent DC, streetscape safety matters less for Yelp scores","Yelp ratings: streetscape safety fades where cars dominate","Car dependency tempers the streetscape boost to restaurant ratings","Streetscape effect on Yelp ratings shrinks in car-centric neighborhoods"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00018,"raw_usage":{"total_tokens":1223,"prompt_tokens":781,"completion_tokens":442,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":397,"completion_tokens_details":{"reasoning_tokens":351}},"tokens_in":397,"tokens_out":442,"duration_ms":5241,"temperature":1.0,"reasoning_tokens":351,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T12:00:43.234162+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Observe actual arrival modes at the sampled establishments, for example from mobile-location traces or parking utilization, and re-run the ordinal logistic model with that real car-access measure replacing the neighborhood Car Dependency Index; the central moderation claim would be falsified if the safety interaction is not negative when actual car access is used.","supporting_citations":[],"review_version":1}