Pith. sign in

REVIEW 2 cited by

Urban Safety Perception Through the Lens of Large Multimodal Models: A Persona-based Approach

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.00610 v1 pith:XY2475OM submitted 2025-03-01 cs.CY cs.AI

classification cs.CYcs.AI
keywords urbansafetymodelmodelsperceptionspersona-basedsocio-demographicapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Understanding how urban environments are perceived in terms of safety is crucial for urban planning and policymaking. Traditional methods like surveys are limited by high cost, required time, and scalability issues. To overcome these challenges, this study introduces Large Multimodal Models (LMMs), specifically Llava 1.6 7B, as a novel approach to assess safety perceptions of urban spaces using street-view images. In addition, the research investigated how this task is affected by different socio-demographic perspectives, simulated by the model through Persona-based prompts. Without additional fine-tuning, the model achieved an average F1-score of 59.21% in classifying urban scenarios as safe or unsafe, identifying three key drivers of perceived unsafety: isolation, physical decay, and urban infrastructural challenges. Moreover, incorporating Persona-based prompts revealed significant variations in safety perceptions across the socio-demographic groups of age, gender, and nationality. Elder and female Personas consistently perceive higher levels of unsafety than younger or male Personas. Similarly, nationality-specific differences were evident in the proportion of unsafe classifications ranging from 19.71% in Singapore to 40.15% in Botswana. Notably, the model's default configuration aligned most closely with a middle-aged, male Persona. These findings highlight the potential of LMMs as a scalable and cost-effective alternative to traditional methods for urban safety perceptions. While the sensitivity of these models to socio-demographic factors underscores the need for thoughtful deployment, their ability to provide nuanced perspectives makes them a promising tool for AI-driven urban planning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Time Series Foundation Models are Flow Predictors

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Zero-shot time series foundation models outperform trained statistical and deep learning baselines for origin-destination crowd flow prediction on three real-world mobility datasets.

  2. Interpretable Multimodal Framework for Human-Centered Street Assessment: Integrating Visual-Language Models for Perceptual Urban Diagnostics

    cs.CV 2025-06 reject novelty 4.0 of 10

    MSEF fine-tunes VisualGLM-6B with GPT-4-generated soft labels to assess streetscape walkability, safety, and vibrancy, reporting F1 0.84 and 89.3% perception agreement.

Pith tools