{"id":"833bcad1-cad1-4ddb-a8d5-feebc9a7a316","arxiv_id":"2504.14065","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A procedural, open-source pipeline generates object-based 3D worlds from Dutch public data, evaluated for visual fidelity and shown with live transit data.","lead":"AnywhereXR is an open-source pipeline that creates 3D environments from public Dutch geospatial data for virtual and augmented reality applications. The paper reports a user evaluation and a live public transport case study, positioning the tool as a lightweight basis for immersive digital twins.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The proof-of-concept is plausible, but the paper's own survey undercuts the central 'high-fidelity street-level' claim: mean fidelity scores are 3.1–3.8 out of 7 and street-level perspectives receive significantly lower ratings.","rationale":"The paper is an honest proof-of-concept with a real open-source artifact, a reproducible workflow, and an explicitly acknowledged limitation about data availability. The reader's conditional acceptance is reasonable. My stress-test focuses on a different, more immediate load-bearing point: the central claim of 'high-fidelity' immersive environments is not well supported by the paper's own evaluation, especially for street-level perspectives, which are the main differentiator claimed against Google Maps and Cesium. The survey results show moderate ratings overall and a statistically significant penalty for street-level views. Since the paper explicitly invokes the visual Turing test as its guiding principle and claims street-level fidelity as a key advantage, the evidence gap is material. However, this is a weakness in the strength of the claim rather than a fatal flaw in the system's feasibility; a controlled matched-perspective study could resolve it. The reader flagged the evaluation as a red flag but chose data availability as the weakest assumption; I partially agree with the reader, but I would place the primary burden on the fidelity evidence because it bears directly on the central claim as stated, not only on future scalability. The verdict remains conditional: the paper should be accepted with required revisions, specifically a more rigorous fidelity evaluation and softened or re-evidenced street-level claims.","tokens_in":28248,"tokens_out":5055,"duration_ms":50117,"concrete_test":"Re-run the fidelity assessment with matched image pairs: for each of the 10 Dutch sites, capture AnywhereXR and reference images from identical camera positions, viewing angles, time-of-day proxies, and field of view, using at least 60 naive raters and a preregistered threshold (e.g., mean rating at or above 5 on the 7-point scale for each of buildings, trees, terrain, and roads). If street-level ratings remain below threshold under matching, the paper should qualify or drop the 'high-fidelity street-level' claim and reframe the contribution as a procedural open-data pipeline with moderate current fidelity.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The load-bearing weakness is the empirical support for 'high-fidelity' in exactly the regime where AnywhereXR claims its main advantage over existing platforms. Section 4.2.5 reports mean visual-fidelity ratings of only 3.1 (trees) to 3.8 (terrain) on a 7-point scale, and the conjoint models (Model 3 and Model 4 in Table A.2) find a significant negative effect of street-level perspective (Frequentist coefficient -0.31, p < .05; Bayesian estimate -0.13 with 95% CI [-0.24, -0.03]). Yet Section 2.4 and Section 6 motivate AnywhereXR specifically by the failure of mesh-based platforms such as Google Maps and Cesium at street level. These results do not establish that the generated environments achieve 'high fidelity' in the ground-level regime central to the paper's contribution. The survey itself is a convenience sample with only 29 complete responses, image pairs are not controlled for season, lighting, camera pose, or field of view, and raters compared AnywhereXR screenshots against Google Earth images or photographs of the physical sites. The observed street-level deficit could therefore partly reflect stimulus mismatch rather than the actual quality of the generated environment. This does not invalidate the feasibility proof or the open-source artifact, but it makes the central 'high-fidelity, street-level' claim conditional on evidence the current paper does not provide.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript introduces AnywhereXR, a Unity-based procedural generator that assembles object-based 3D environments from open Dutch geospatial data sources (PDOK, AHN, 3D BAG, and the Bomenregister). Five generation stages are described: land cover/use, elevation, water, buildings, and trees. The authors report a survey evaluation (29 complete responses) in which participants rated visual fidelity of AnywhereXR screenshots against Google Earth images or photographs for 10 Dutch locations, alongside mixed-model and conjoint analyses. A case study demonstrates integration of live Dutch public transport data via an OV-API. The paper claims that AnywhereXR provides lightweight, high-fidelity, street-level-capable environments that can serve as a basis for open immersive digital twin applications.","tokens_in":28577,"tokens_out":4919,"duration_ms":44668,"significance":"If the central claims hold, AnywhereXR would be a notable open-source contribution to the digital twin community: it demonstrates that publicly available, object-level geospatial data can be assembled into interactive 3D environments without proprietary meshes, enabling applications such as participatory planning, co-design, and real-time data integration. The paper ships source code and evaluation data (Zenodo/GitLab), which supports reproducibility, and the modular architecture could be reused for other regions. The main significance is the proof of concept and the articulation of a workflow, rather than the specific user-study results. However, the paper's own evaluation provides only modest support for the 'high fidelity' and 'street-level advantage' claims, and several reporting inconsistencies in the statistical tables need to be resolved.","major_comments":[{"comment":"The central claims of 'high-fidelity' environments and superior street-level quality are not supported by the reported evaluation. The aggregated fidelity scores are only 3.1 (trees) to 3.8 (terrain) out of 7, and the conjoint models find a significant negative effect of street-level perspective (Model 3: -0.31, p < .05; Model 4: -0.13, 95% CI [-0.24, -0.03]). Yet the abstract, highlights, Section 2.4, and Discussion (Section 6) motivate AnywhereXR specifically by its ground-level advantage over mesh-based platforms. The manuscript should either substantially temper these claims or provide additional task-based or comparative evidence that directly addresses street-level fidelity.","section":"Abstract, Section 4.2.5, Section 6, Table 1"},{"comment":"The interaction terms in Table A.2 are reported inconsistently. For example, 'Terrain x Ede Grotestraat Downtown' appears twice with different coefficients and signs (0.84 (0.37) and -1.68 (0.43); similarly for 'Terrain x Neijmegen Lent' and 'Terrain x Wageningen Skyline Rijn'). The second block appears to be intended as Trees x location interactions based on the text's interpretation that 'mostly the trees negatively impact ratings', but the table labels them as Terrain interactions. This makes the statistical results as printed impossible to interpret and the interaction effects must be re-reported correctly.","section":"Table A.2 and Section 4.2.5"},{"comment":"The term 'lightweight' is a load-bearing descriptor in the title and Highlights, but the paper provides no quantitative performance evidence. No measurements are given for generation time, memory use, polygon counts, or frame rates in the generated environments; the only number is a reference to '60 frames per second' in the transport case study. For a systems paper, some basic performance characterization is needed to substantiate the lightweight claim.","section":"Section 3, Section 5.2, and title"},{"comment":"The objective image-difference measures (number of missing trees, missing objects, noticeable buildings, roads, terrain; perspective) appear to be coded by the authors without a documented protocol, and the coding includes fractional values (e.g., location 'rcb' has #missing trees = 1.10). Please describe how these values were derived, who performed the coding, whether any inter-rater reliability was assessed, and how non-integer counts arise. This matters because the conjoint models (Model 3 and 4) directly rely on these measures.","section":"Table A.3 and Section 4.2.4"}],"minor_comments":[{"comment":"The text reports 29 participants who filled out the questionnaire completely, but Table A.2 lists 'Num. groups: ResponseId 28' for all models. Please clarify whether one participant was excluded (e.g., due to missing covariates) or whether the table is incorrect.","section":"Section 4.2.2 and Table A.2"},{"comment":"There are several typos and spelling inconsistencies: 'reposiories' in Section 3, 'adaption' in the Highlights, 'forumla' in Section 4.2.4, and 'Schevingen' used repeatedly instead of 'Scheveningen' (e.g., Table A.2, Section 4.2.5, and Figure captions).","section":"Throughout"},{"comment":"The sentence describing the best and worst locations reads ambiguously: 'the lowest scores from a very iconic location in Rotterdam, the Cube Houses, shown at street-level (see image pair second from top in Figure 17; 4.2 versus 2.4)'. Please clarify which number corresponds to which location and whether '4.2 versus 2.4' refers to overall ratings or a category.","section":"Section 4.2.5"},{"comment":"The precision of AHN 3 is described as 'approximately 0.5 square meters'; this should be stated as a resolution (e.g., 0.5 m cell size) rather than an area, or the area unit should be clarified.","section":"Section 3.2.2"},{"comment":"The same Zenodo DOI is given for both the source code and the evaluation data, but the text says the code is at a GitLab URL while data are at Zenodo; please confirm whether both are at the same Zenodo record or provide a separate DOI for the data.","section":"Data availability statement"}],"recommendation":"major_revision","confidential_remarks":"The paper is fundamentally a systems/proof-of-concept contribution, and its value lies in the open-source implementation and the demonstrated workflow rather than in the user study. I would suggest the editors encourage the authors to align the abstract and discussion with the actual evidence: the evaluation supports feasibility and encourages further development, but not the strong 'high-fidelity' and 'street-level superiority' claims as currently stated. The statistical table errors are fixable but must be corrected. There is also a notable reliance on self-citations among the authors; this is not problematic per se, but the related-work framing should be checked for balance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Jascha, quick read of AnywhereXR. The genuinely new thing is the assembly: a modular Unity pipeline that turns Dutch open data (PDOK, AHN, 3D BAG, Bomenregister) into object-based 3D scenes, with a live public transport feed as a working digital-twin case. Individually the pieces are known—Earcut, ray tracing, flood-fill, standard game-engine rendering—but the integration into dynamic-resolution, object-level environments with real-time data is a real proof of concept. The authors deserve credit for shipping code and data on Zenodo and for being upfront about the mismatch between open-data quality and visual fidelity.\n\nThe soft spot is the evaluation, and it is load-bearing for the 'high-fidelity, street-level' claim. The survey's own numbers: mean fidelity scores 3.1–3.8 out of 7, and street-level perspectives rated significantly worse in both frequentist and Bayesian models. The paper motivates AnywhereXR by arguing that mesh-based platforms like Google Maps fail at street level; the survey doesn't show AnywhereXR succeeds there yet. The authors acknowledge the limitations—convenience sample, imperfect image matching, data gaps—and frame it as a first evaluation, which is honest. But the strength of the conclusion should be scaled down to match: the feasibility proof and the open pipeline stand; the perceptual claim is conditional on a more controlled study.\n\nMinor points: the 'anywhere' name overpromises given that high-quality object-level open data exists in few countries; the paper says this too, but the title sells it harder. Table 1's comparison is qualitative and a bit self-serving; treat it as illustrative. The statistical models are overkill for 29 raters but not wrong.\n\nBottom line: this deserves a serious referee. The engineering contribution is real and reproducible, the paper is readable, and the evaluation problems are fixable with a larger, better-matched study. I'd ask for revisions rather than desk reject. I'd cite it for the pipeline work in my own XR/digital twin writing.","headline":"A solid open-source pipeline paper whose own survey undercuts the street-level fidelity claim; the engineering contribution and reproducible artifacts carry it through peer review.","tokens_in":29080,"tokens_out":1284,"would_cite":true,"duration_ms":13088,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AnywhereXR claims that open geospatial data alone can feed a procedural pipeline that builds object-based, street-level immersive digital twins, demonstrated for the Netherlands.","keywords":["immersive digital twin","procedural generation","open geospatial data","extended reality","digital twin","participatory co-design","data visceralization","3D environment generation"],"falsifier":"Run the same AnywhereXR pipeline on a region outside the Netherlands that lacks an open 3D building registry and tree registry but has only 2D vector tiles and coarse elevation. If the resulting environment cannot reach the fidelity levels reported for Dutch scenes, roughly 3 to 4 on a 7-point scale, or requires manual post-processing to fix water and buildings, then the claim that open data is sufficient for on-the-fly immersive digital twins is falsified for that data regime.","tokens_in":28096,"feed_emoji":"🗺️","tokens_out":5578,"duration_ms":44999,"temperature":0.7,"pith_summary":"AnywhereXR is a procedural pipeline that turns publicly available geospatial data into object-based, street-level 3D environments inside a game engine, without manual modeling. The paper claims that for data-rich countries such as the Netherlands, open data alone can support high-fidelity immersive digital twins: environments where each tree, building, road, and water body is a separate object that can be selected, modified, or animated. If this holds, participatory planning, co-design, and visceral data experiences could run on open infrastructure rather than on proprietary globe services. The authors support the claim with a five-stage implementation, an expert survey of visual fidelity, and a case study that injects live public-transport positions into the generated scene. The broader ambition is that the workflow generalizes beyond the Netherlands, though the authors acknowledge that most countries do not yet publish the required object-level data.","feed_headline":"Open data can build street-level 3D worlds on the fly","feed_subtitle":"A Dutch pipeline turns public maps, elevation, buildings, and tree records into object-based immersive digital twins.","key_machinery":"The load-bearing mechanism is the five-stage AnywhereXR pipeline, running inside the Unity game engine: land cover and use polygons from PDOK vector tiles are triangulated with the Earcut algorithm, AHN elevation data is cleaned with a circular post-filter, water polygons are intersected with terrain boundaries to fix elevation mismatches, 3D BAG buildings are imported from b3dm tiles and textured with aerial imagery, and tree positions are extracted from Bomenregister image tiles by flood-filling crown regions. The pipeline's signature technique is dynamic terrain detail: soil information is baked into a constant-size texture and the mesh resolution is sampled from that texture based on the viewer's position, so close-up fidelity stays high without a fixed high-polygon mesh. These stages work because every object remains a separate Unity entity, which is what allows later stages, simulations, or user edits to touch individual trees, buildings, or roads.","core_discovery":"The central claim is that a lightweight, modular procedural generator can create object-based, high-fidelity immersive environments from open geospatial data. AnywhereXR queries public Dutch databases — PDOK for land cover, AHN for elevation, 3D BAG for building geometry, and the partially open Bomenregister for trees — and extrudes them into a Unity scene in five stages: land cover and use, elevation, water, buildings, and trees. Unlike 3D tile services built from aerial reconstructions, each environmental entity is an individual mesh, which makes the model adaptable to alternative future states and real-time data. The paper does not claim photographic parity: an online survey of 29 complete responses rated visual fidelity between 3.1 and 3.8 on a 7-point scale across trees, buildings, terrain, and roads, and street-level scenes scored lower than aerial views. The claim is that this fidelity is already sufficient for applications such as spatial planning and stakeholder engagement, and that the modular design is the path to \"anywhere\" coverage.","pith_inferences":["If the \"anywhere\" ambition is taken literally, the paper's own data-availability caveat means the approach will first succeed where governments publish open object-level cadastral, elevation, and tree data; elsewhere it will need generative AI or citizen-science data to fill gaps.","The finding that one missing tree hurts ratings more than many missing trees suggests that perceptual fidelity is driven by salient anchor objects rather than aggregate completeness; a follow-up experiment could test whether placing characteristic objects in known locations improves presence more than adding generic detail.","Extending the NDOV case, the same socket-based API pattern could ingest other real-time feeds, such as weather, energy use, or pedestrian counts, to make the twin a live decision-making surface rather than a static scene.","The object-based representation makes AnywhereXR a natural testbed for comparing the usability of low- versus high-fidelity environments in participatory planning, complementing prior co-design work."],"forward_implications":["The same five-stage pipeline can be pointed at another country's open datasets once they exist, making \"anywhere\" a data-availability problem rather than a modeling problem.","Object-based environments support scenario editing, such as changing tree density, road networks, or building states to simulate future or alternative designs.","Live data feeds, like the NDOV public-transport case, can be merged into the scene to turn a static 3D model into a digital twin that reflects current conditions.","Because terrain complexity is sampled dynamically from a texture, the approach keeps rendering cost low while preserving street-level detail.","The survey's marginal-effect results imply that completeness, such as a single missing characteristic tree, matters more to perceived fidelity than many other visual defects."],"supporting_citations":[{"why":"Supplies the prior demonstration that open geospatial data and game engines can create immersive virtual environments, which AnywhereXR builds on and extends.","marker":"[63]"},{"why":"Supplies the idea of generating digital twins on demand through reusable infrastructure rather than hand-building each twin.","marker":"[56]"},{"why":"Provides the FAIR data principles that motivate relying on open and reproducible data and software.","marker":"[70]"},{"why":"Provides the glTF-based streaming format underlying the b3dm building tiles that AnywhereXR parses into Unity meshes.","marker":"[73]"},{"why":"Supplies the UV-mapping explanation used for placing facade textures onto building walls and roofs.","marker":"[75]"},{"why":"Defines the visual Turing test that guides the design principle of veridical, high-fidelity replication.","marker":"[77]"},{"why":"Supports the argument that plausibility can matter more than raw fidelity, legitimizing environments that fall short of photorealism.","marker":"[85]"},{"why":"Shows that objectively lower-fidelity representations can be well received in participatory co-design, supporting the claimed application value.","marker":"[36]"},{"why":"Identifies Gaussian splatting as a future route for filling in scene detail when object-level data is scarce.","marker":"[111]"}],"fun_headline_variants":["Dutch open data spins up on-the-fly 3D digital twins","Five-stage pipeline turns public geodata into immersive twins","AnywhereXR: modular generator builds open-source 3D worlds","Open data pipeline yields object-based 3D twins on the fly","On-the-fly 3D environments for participatory decision-making"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that governments will keep publishing high-quality, open, object-level geospatial data outside the Netherlands; if that data does not appear, the \"anywhere\" part of the approach cannot be realized.","fun_headline_variants_meta":{"raw":{"variants":["Dutch open data spins up on-the-fly 3D digital twins","Five-stage pipeline turns public geodata into immersive twins","AnywhereXR: modular generator builds open-source 3D worlds","Open data pipeline yields object-based 3D twins on the fly","On-the-fly 3D environments for participatory decision-making"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001569,"raw_usage":{"total_tokens":6233,"prompt_tokens":885,"completion_tokens":5348,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":501,"completion_tokens_details":{"reasoning_tokens":5259}},"tokens_in":501,"tokens_out":5348,"duration_ms":31806,"temperature":1.0,"reasoning_tokens":5259,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T11:56:45.233698+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same AnywhereXR pipeline on a region outside the Netherlands that lacks an open 3D building registry and tree registry but has only 2D vector tiles and coarse elevation. If the resulting environment cannot reach the fidelity levels reported for Dutch scenes, roughly 3 to 4 on a 7-point scale, or requires manual post-processing to fix water and buildings, then the claim that open data is sufficient for on-the-fly immersive digital twins is falsified for that data regime.","supporting_citations":[{"cited_title":"Shadbolt, K","cited_arxiv_id":null,"evidence_quote":"Provides the glTF-based streaming format underlying the b3dm building tiles that AnywhereXR parses into Unity meshes."},{"cited_title":"Schilling, J","cited_arxiv_id":null,"evidence_quote":"Supplies the UV-mapping explanation used for placing facade textures onto building walls and roofs."},{"cited_title":"Flavell, UV Mapping, Apress, Berkeley, CA, 2010, Ch","cited_arxiv_id":null,"evidence_quote":"Defines the visual Turing test that guides the design principle of veridical, high-fidelity replication."},{"cited_title":"J., Making climate change visible: A critical role for landscape professionals, Landscape and Urban Planning 142 (2015) 95–105","cited_arxiv_id":null,"evidence_quote":"Identifies Gaussian splatting as a future route for filling in scene detail when object-level data is scarce."}],"review_version":1}