{"id":"aa967b00-e3ae-4bf7-9a32-e95820316dac","arxiv_id":"2507.09367","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A step-by-step cookbook and open-source codebase for building a multi-agent VR transportation simulator where pedestrians, cyclists, drivers, and transit users interact in real time while their physiological and neural responses are recorded.","lead":"This paper describes an open-source, virtual-reality simulation platform that lets real people interact as pedestrians, cyclists, drivers, transit riders, and automated-vehicle passengers in one shared city environment, while sensors record their brain, eye, heart, and stress signals. The authors publish hardware specifications and scripts so that research labs can build or adapt the system, aiming to make high-fidelity transportation experiments more accessible.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The real-time, synchronized multimodal capability claim rests on unvalidated UDP latency and fNIRS-in-VR data quality; no measurements are reported.","rationale":"The paper is a systems and cookbook contribution, not a quantitative empirical claim; its value lies in the modular architecture, hardware specifications, and open-source scripts. The authors are honest about limitations in Section 5, which supports conditional acceptance. The critical gap is that the defining capability, namely real-time, synchronized, multimodal, multi-agent interaction, has no measured evidence. No latency numbers, no synchronization drift data, no fNIRS signal quality validation under the intended motion conditions, and no demonstrated multi-agent run involving all claimed roles are provided. A careful reader cannot tell whether a replication attempt would yield aligned data streams or unusable signals. Because the concern is addressable through published validation, the verdict remains conditional rather than reject. The reader's weakest_assumption correctly identifies this area; my stress test agrees and proposes a concrete measurement protocol that would settle whether the concern lands.","tokens_in":24330,"tokens_out":5114,"duration_ms":58188,"concrete_test":"Run the Figure 7 three-agent scenario (driver, cyclist, pedestrian) with two or more human participants for 10 minutes. Log end-to-end UDP latency and clock drift between agent computers using a common hardware time base (e.g., PTP or an oscilloscope on a synchronization pulse), and record fNIRS signals during walking and cycling versus a sitting baseline. If peak end-to-end latency exceeds 50 ms, drift accumulates beyond the sensor sampling period, or fNIRS motion artifact amplitude (e.g., spline or wavelet artifact detection) exceeds the expected hemodynamic response, then the real-time aligned multimodal claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that a lab can replicate the platform and run real-time, synchronized multi-agent experiments collecting aligned behavioral, physiological, and neural data. That requires (a) low-latency, drift-free synchronization across agent computers and (b) usable fNIRS signals during motion. The paper states in Section 3.1 that agents are 'synchronized using a low-latency communication protocol based on User Datagram Protocol (UDP), ensuring real-time data exchange and temporal alignment,' and in Section 3.2.1 that 'all sensing streams are time-synchronized with simulator events via software-level integration,' but reports no latency, jitter, or drift measurements, and no ground-truth validation of the synchronization. Section 5 concedes that combined fNIRS and VR 'can pose integration difficulties' due to motion artifacts, sensor displacement, and physical interference, yet the use cases (Figures 9 and 11) show sample traces from a single participant and no multi-agent scenario involving a transit user interacting with others. Therefore the platform's defining capability is asserted rather than demonstrated; the load-bearing assumption is that off-the-shelf UDP networking and headband fNIRS under a Varjo XR-4 deliver research-grade alignment and signal quality during walking, cycling, and riding.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes a modular, open-source multi-agent VR simulation platform intended to study pedestrians, cyclists, drivers, automated vehicles, and public transit users in a shared virtual environment. It provides hardware specifications, system architecture, integration of fNIRS, eye tracking, and wrist-based biosensors, and presents three use cases: multimodal travel transitions, human-automated vehicle interaction, and road user well-being. The authors claim real-time synchronized multi-agent interaction as well as replicability through a 'cookbook' approach and open-source scripts.","tokens_in":24500,"tokens_out":5074,"duration_ms":57638,"significance":"If validated, the platform would fill a genuine gap in the literature: existing multi-agent human-in-the-loop simulators typically support only two agent types and largely exclude public transit users. The paper's strengths are its detailed hardware component list, explicit modular architecture, open-source repository, and the breadth of sensing modalities it attempts to integrate. These assets make it a potentially useful reference for other laboratories building similar systems. However, the manuscript does not provide measured evidence for the two load-bearing capabilities it advertises: real-time multi-agent synchronization and usable fNIRS signals during whole-body movement. The contribution is therefore best judged as a promising systems description rather than a demonstrated, ready-to-replicate platform.","major_comments":[{"comment":"The abstract and Section 3.1 state that agents are 'synchronized using a low-latency communication protocol based on User Datagram Protocol (UDP), ensuring real-time data exchange and temporal alignment,' and Section 3.2.1 asserts that 'all sensing streams are time-synchronized with simulator events via software-level integration.' No latency, jitter, drift, or ground-truth synchronization measurements are reported anywhere in the manuscript. The only multi-agent interaction shown is the schematic in Figure 8; Figures 9-11 present single-participant traces. Because real-time synchronized multi-agent interaction is the platform's defining claimed capability, the paper should include at least one measurement of end-to-end synchronization error from a multi-agent run to substantiate it.","section":"Section 3.1 (System Architecture) and Section 3.2.1"},{"comment":"The Discussion concedes that combined fNIRS and VR 'can pose integration difficulties' due to motion artifacts, sensor displacement, and physical interference, and states that 'achieving long-term or continuous monitoring with high spatial and temporal precision remains technically demanding.' Nevertheless, Figure 9(C) presents fNIRS traces from cycling, walking, and transit as evidence of synchronized neural monitoring. No signal-quality metrics, artifact-rejection counts, or comparisons against a reference are reported. To support the multimodal sensing capability claim, the authors should report at least one validation check (e.g., channel reliability, signal-to-noise ratio, or motion-artifact rejection rate) for the fNIRS headband during representative activities.","section":"Section 5 (Discussion) and Figure 9C"},{"comment":"The abstract claims the platform 'enables interaction across public transit users, pedestrians, cyclists, automated vehicles, and drivers.' Section 3.1.4 describes only a single user walking on an omnidirectional treadmill and sitting in a seated mode; no scenario involves a transit user simultaneously interacting with another human-controlled agent. Use Case 1 (Section 4.1) is a sequential, single-participant transition through cycling, walking, and transit, not a multi-agent transit interaction. A concrete demonstration of a transit user interacting with at least one other human-controlled agent is needed to support the 'all road users' claim.","section":"Abstract; Section 3.1.4; Section 4.1"},{"comment":"The paper bills itself as a 'step-by-step cookbook' that is 'accessible to users from all technical backgrounds with minimal coding background,' but the manuscript text contains no step-by-step assembly, configuration, calibration, or scenario-authoring instructions. Appendix A provides only a hardware component table, and the actual scripts are relegated to an external OSF link [8]. Without at least one complete replication example or a summary of the key setup steps in the paper, the replicability claim cannot be evaluated from the manuscript itself.","section":"Section 1 (Introduction) and Appendix A"}],"minor_comments":[{"comment":"Section 3.1.4 contains a typo: 'V ARJO XR-4' should read 'Varjo XR-4'.","section":"Section 3.1.4"},{"comment":"Table 2 lists one fNIR2000C headband but three Empatica EmbracePlus wristbands; the text should clarify which sensors are per-agent and how simultaneous physiological recording across multiple agents is configured.","section":"Appendix A, Table 2"},{"comment":"The Discussion refers to 'real users across four distinct modes,' while the Introduction and Abstract enumerate five agent types (public transit users, automated vehicles, pedestrians, cyclists, drivers); please reconcile the count.","section":"Section 5 (Discussion)"},{"comment":"Reference [49] is incomplete (no year or URL), and the OSF repository link in Reference [8] would benefit from a version identifier or DOI for reproducibility.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper's advertised capabilities outrun its demonstrated evidence: there is no measured synchronization validation, no fNIRS signal-quality validation during movement, and no scenario with a transit user interacting with another human-controlled agent. I would encourage the editor to request such validation data or a clearly scoped revision that frames the paper as a system description with explicit open validation needs. The self-citations are appropriate given the authors' prior work in this line of research and are not a concern in themselves."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What you should know: this is a systems paper, not a quantitative study. The contribution is a detailed, open-sourced blueprint for a multi-agent VR simulator that puts five road-user modes—pedestrian, cyclist, driver, automated vehicle, and public transit user—in one synchronized Unity environment, with integrated fNIRS, eye tracking, EDA, and in-VR questionnaires. The transit-user inclusion and e-bike support are genuinely new for this line of work, and the authors are transparent about what is and isn't validated.\n\nThe paper does several things well. The literature review is thorough, and Table 1 gives a useful comparison of prior multi-agent simulators. The hardware specifications and modular architecture are concrete enough for another lab to attempt a build. The TTA-based scenario setup (equal time-to-arrival for agents) is a nice, practical trick. The limitations section is honest, explicitly flagging fNIRS+VR integration difficulties, motion sickness, and the cost/space barriers.\n\nThe soft spot, as you might expect, is that the headline capability—real-time, synchronized, multimodal interaction across all five modes—is asserted rather than demonstrated. There are zero latency, jitter, or drift measurements for the UDP synchronization. No scenario shows a public transit user simultaneously interacting with other agents. The use cases are single-participant, illustrative traces. And the fNIRS-in-VR data quality question is real: the authors themselves cite motion artifact and sensor displacement as known problems, and they offer no ground-truth validation during walking, cycling, or riding. These are addressable gaps, and for a cookbook paper they are not fatal—but they are load-bearing if someone reads the abstract as 'this is ready for research-grade experiments.'\n\nMy take: this deserves serious peer review, but a conditional one. The authors should be asked to measure and report synchronization latency/drift, provide at least one demonstration of a transit user interacting with other agents, and add a validation note on fNIRS signal quality during motion. They should also commit the repository with a hash so the 'open-sourced scripts' claim is auditable.\n\nWho's this for? Researchers building or extending human-in-the-loop transportation simulators, and people planning multimodal VR experiments. It's a useful reference, and the codebase could become a community resource. I'd bring it to reading group, and I'd cite it if I worked in this area—though I'd cite it as a platform description, not as evidence of validated performance.\n\nRecommendation: send to peer review, with requests for validation data and a transit-user demo.","headline":"A useful, honest cookbook for a five-mode multi-agent VR transportation simulator that is one validation study short of its headline claim.","tokens_in":25075,"tokens_out":2688,"would_cite":true,"duration_ms":30583,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that one open-source, modular simulation platform can host pedestrians, cyclists, drivers, automated vehicles, and public transit riders in a single shared virtual environment in real time, while collecting aligned…","keywords":["transportation simulation","virtual reality","multi-agent systems","human-in-the-loop","road user behavior","functional near-infrared spectroscopy (fNIRS)","eye tracking","open-source platform"],"falsifier":"Run a scripted motion protocol in the pedestrian module: have the participant walk a fixed route on the omnidirectional treadmill and then stand still while the vehicle and cyclist agents follow scripted paths, all with the fNIRS headband under the VR headset. If inter-agent clock drift exceeds roughly a frame or a reaction step (tens of milliseconds) by the end of a session, or if the fNIRS channels during walking cannot be distinguished from a resting baseline after standard motion-artifact correction, the claimed real-time multimodal alignment and neural sensing capability are not supported.","tokens_in":24069,"feed_emoji":"🚦","tokens_out":8418,"duration_ms":83896,"temperature":0.7,"pith_summary":"The paper's central claim is that a single shared virtual-reality simulation can host all major road-user roles at once — pedestrians, cyclists, drivers, automated vehicles, and public transit riders — in real time, with each participant moving through a common urban scene and influencing the others. It offers a step-by-step cookbook of hardware choices, wiring, and open-source scripts so that a research group with comparable equipment can replicate the platform rather than buy a proprietary simulator. A sympathetic reader should care because existing simulators mostly isolate one mode at a time or rely on scripted 'other' agents, which limits how well findings about trust, yielding, and negotiation transfer to real streets. The paper also integrates brain-activity (fNIRS), eye-tracking, and wrist-sensor data into that shared scene, aiming to turn each traffic encounter into a synchronized multimodal dataset.","feed_headline":"One recipe builds a real-time simulator for every road-user type","feed_subtitle":"Open-source scripts and hardware specs let labs run synchronized VR tests across all road-user roles in one shared scene.","key_machinery":"The central mechanism is the modular agent-computer topology: one computing engine per road-user role, each attached to a distinct physical interface (an omnidirectional treadmill for the pedestrian, a smart trainer with elevation and wind simulation for the cyclist, an actuated cockpit with force-feedback wheel and pedals for the driver, and a walk-then-seat configuration for the transit user), all exchanging state over a low-latency UDP protocol into a single Unity scene. This per-agent computer plus UDP plus shared-scene arrangement is what turns independent hardware into one time-synchronized multiplayer environment; a second load-bearing mechanism is the human sensing stack — a headband fNIRS sensor, the VR headset's embedded eye tracker, and wrist biosensors — whose streams are software-aligned to scenario events, alongside in-VR questionnaires and an N-back task for cognitive load.","core_discovery":"The core discovery claimed by the paper is that it is feasible to assemble, from commodity VR, motion, and sensing hardware, a synchronized human-in-the-loop simulation in which a pedestrian on an omnidirectional treadmill, a cyclist on a smart trainer, a driver in an actuated cockpit, and a seated public transit user all occupy the same virtual city block and respond to each other's live behavior. Each agent runs on a dedicated computer coordinated over UDP, so physical inputs such as steering angle, pedal cadence, and walking direction drive a shared Unity scene, and each participant's state is captured by fNIRS, embedded eye tracking, and wrist-based biosensors that are software-synchronized to simulation events. The paper presents three use cases — a continuous cycling-walking-transit journey, human encounters with an automated vehicle using multimodal external human-machine interfaces, and mode-specific physiological responses to traffic and infrastructure variations — as demonstrations that the platform can produce layered, time-aligned behavioral, neural, and physiological data across roles.","pith_inferences":["If the platform's synchronization holds, a natural next step is cloud-based distributed operation so that participants at different sites share one scene; the paper itself notes that network latency and drift remain unresolved, so this is conditional on future validation.","The same agent-computer-plus-sensing architecture could be reused beyond transportation, for example to study crowd navigation in buildings or mixed pedestrian-robot spaces, though the paper does not make that claim.","A testable extension would be publishing end-to-end latency, clock drift, and sensor-accuracy numbers alongside the scripts; their absence is the main reason the real-time multimodal claim currently rests on an assumption rather than a measurement.","The equal-time-to-arrival staging technique suggests a general method for creating controlled social encounters in VR; varying the TTA distribution could test how timing uncertainty changes yielding and conflict decisions."],"forward_implications":["A lab that follows the cookbook and uses the open scripts can run synchronized real-time experiments with multiple human participants in different road-user roles within one shared scene, without building proprietary infrastructure.","Researchers can study a single participant across mode transitions (cycling to walking to transit) and obtain continuous fNIRS, eye-tracking, and wrist-sensor streams aligned to each phase, which single-mode simulators cannot offer.","The automated vehicle module supports supervised autonomous operation with takeover controls, enabling studies of trust, supervisory attention, and takeover timing with multiple interacting human road users.","The equal-time-to-arrival scenario controller lets experimenters stage naturalistic negotiation moments, such as unsignalized crossings, where yielding and conflict-resolution behavior can be observed across modes.","Multimodal sensing combined with in-VR questionnaires makes it possible to relate physiological and neural state to behavior without breaking immersion, supporting mechanism-level accounts of road-user decisions."],"supporting_citations":[{"why":"The open-source codebase that the paper says makes the platform replicable and adaptable.","marker":"[8]"},{"why":"Prior immersive bicycle-simulator work with psycho-physiological measures that the cyclist module builds on.","marker":"[38]"},{"why":"A coupled bicycle–automated-vehicle driving simulator that this platform extends to additional agent types.","marker":"[63]"},{"why":"Prior multi-agent distributed immersive VR research whose limitations and requirements the platform addresses.","marker":"[85]"},{"why":"The commercial city-generator asset used to create the customizable virtual urban environment.","marker":"[73]"},{"why":"Evidence and method for collecting questionnaires inside VR without breaking immersion.","marker":"[87]"},{"why":"Field measurements of cyclist speeds used to set the cyclist agent's speed in the synchronized-arrival scenario.","marker":"[28]"},{"why":"Field study of average adult walking speed used to set the pedestrian agent's speed in the same scenario.","marker":"[57]"}],"fun_headline_variants":["Open-source VR puts every road user in one shared scene","Simulator unites pedestrians, cyclists, drivers, transit in real time","Cookbook-style guide for building a multi-agent transport simulator","One platform, all road users: synchronized VR transportation sim","Human-centered simulator with open-source scripts for all modes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the off-the-shelf sensors — especially the brain-activity headband worn together with the VR headset — and the UDP network alignment stay accurate and time-synchronized while participants actually walk, pedal, and drive, since the paper reports no latency, drift, or sensor validation measurements.","fun_headline_variants_meta":{"raw":{"variants":["Open-source VR puts every road user in one shared scene","Simulator unites pedestrians, cyclists, drivers, transit in real time","Cookbook-style guide for building a multi-agent transport simulator","One platform, all road users: synchronized VR transportation sim","Human-centered simulator with open-source scripts for all modes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000404,"raw_usage":{"total_tokens":2126,"prompt_tokens":991,"completion_tokens":1135,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":1052}},"tokens_in":607,"tokens_out":1135,"duration_ms":13040,"temperature":1.0,"reasoning_tokens":1052,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T17:57:36.365571+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a scripted motion protocol in the pedestrian module: have the participant walk a fixed route on the omnidirectional treadmill and then stand still while the vehicle and cyclist agents follow scripted paths, all with the fNIRS headband under the VR headset. If inter-agent clock drift exceeds roughly a frame or a reaction step (tens of milliseconds) by the end of a session, or if the fNIRS channels during walking cannot be distinguished from a resting baseline after standard motion-artifact correction, the claimed real-time multimodal alignment and neural sensing capability are not supported.","supporting_citations":[{"cited_title":"Psycho-physiological measures on a bicycle simulator in immersive virtual environments: How pro- tected/curbside bike lanes may improve perceived safety","cited_arxiv_id":null,"evidence_quote":"Prior immersive bicycle-simulator work with psycho-physiological measures that the cyclist module builds on."},{"cited_title":"A coupled driving simulator to investigate the interaction between bicy- cles and automated vehicles","cited_arxiv_id":null,"evidence_quote":"A coupled bicycle–automated-vehicle driving simulator that this platform extends to additional agent types."},{"cited_title":"Mad-ive: Multi-agent distributed immersive virtual environ- ments for vulnerable road user research—potential, challenges, and re- quirements","cited_arxiv_id":null,"evidence_quote":"Prior multi-agent distributed immersive VR research whose limitations and requirements the platform addresses."},{"cited_title":"Fantastic city generator","cited_arxiv_id":null,"evidence_quote":"The commercial city-generator asset used to create the customizable virtual urban environment."},{"cited_title":"The influence of in-vr questionnaire design on the user experience","cited_arxiv_id":null,"evidence_quote":"Evidence and method for collecting questionnaires inside VR without breaking immersion."},{"cited_title":"Field studies of pedestrian walking speed and start-up time","cited_arxiv_id":null,"evidence_quote":"Field study of average adult walking speed used to set the pedestrian agent's speed in the same scenario."}],"review_version":1}