{"id":"1e2edd64-1757-42e7-821e-8bccc5d56e8e","arxiv_id":"2607.22613","paper_version":1,"verdict":"REJECT","confidence":"HIGH","novelty_score":3.0,"correctness_risk":"high","formal_verification":"none","parameter_count":1,"one_line_summary":"Proposes an LLM + AR pipeline for visualizing 'place identity,' demonstrated only as a ChatGPT/DALL-E exercise on Monastiraki Square, with the AR overlay unimplemented and the benefit claims unevaluated.","lead":"This paper sketches a plan to enrich real city squares with AI-generated historical images and stories in an augmented-reality headset, and demonstrates the first steps on one square in Athens. It is a concept sketch: ChatGPT answers and DALL-E images exist, but the AR overlay is never built and no one tests whether it actually deepens sense of place.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The pipeline's content layer is admitted to be factually unreliable, and no verification stage exists; therefore the central claim that the tool preserves/enhances cultural memory is unsupported.","rationale":"The reader's weakest_assumption correctly identifies the fidelity of LLM/generative outputs as the load-bearing point. The paper's own methodology (§4.1) and demonstrable failures (§4.2) make this concern concrete and internally sourced. The central claim—that the pipeline enhances identification and appreciation of place identity—depends on the content being trustworthy enough to represent cultural and political memory. Without a verification stage, the tool risks fabricating the memory it claims to preserve. This is not a peripheral issue; it directly invalidates the asserted 'robust methodology.' My independent reading found no reason to move away from REJECT. The paper could become viable with an implemented AR layer and a sourcing/verification step, but as submitted, the central argument is unsupported. The reader's verdict stands unchanged.","tokens_in":9015,"tokens_out":2308,"duration_ms":29471,"concrete_test":"Run a factual audit of the pipeline's content for a sample of documented public squares (including Monastiraki). Using the exact prompt strategy from §4.1, query ChatGPT/Gemini and generate DALL-E images for ancient/Byzantine/future periods for at least 10 squares. Have two historians independently rate the factual accuracy of the text and image content against primary/secondary archival sources (e.g., ancientathens3d.com, bibliographic records), using a pre-registered scale (e.g., 0–3: 0=no factual basis, 1=major errors, 2=partially correct, 3=fully corroborated). Also record whether the LLM cited any verifiable sources. If the mean accuracy score is below a threshold (e.g., <2) or a substantial fraction of claims are uncorroborated, the content layer cannot reliably carry cultural memory, and the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim (Conclusion, §4) requires that LLM-generated text and images faithfully convey cultural and political memory of a place. This is explicitly undermined in the methodology: §4.1 concedes 'the provided images are not derived from accurate historical images and sources but rather serve as a means of generating imaginative representations' and 'The main issue with these technologies lies in their limited accuracy in generating photographs.' §4.2 demonstrates the failure: ChatGPT cannot produce historical illustrations, and Gemini does not recognize Monastiraki Square. Since the methodology includes no verification or sourcing step, the AR overlay would present fabricated content as historical memory. Even if the user experience is engaging, the core function—anchoring memory to physical attributes—fails if the content is hallucinated. The paper's own limitation statements (LLM bias, prompt injection, image inaccuracy) confirm this fragility. This is not a disagreement with current consensus; it is an internal inconsistency between the assertion (§3.3: 'profound understanding of historical facts') and the admitted output characteristics (§4.1). Thus the load-bearing assumption—that LLMs/generative models can reliably carry cultural memory—is insecure.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a technological pipeline for enhancing 'place identity' in urban environments by combining Google Street View imagery, ChatGPT as a tour-guide-like information source, DALL-E for generating era-specific images, and the Apple Vision Pro headset for AR overlay. The methodology is demonstrated on Monastiraki Square in Athens, where ChatGPT answered factual questions and DALL-E produced three images representing ancient, Byzantine, and speculative future scenes. The conclusion claims that this constitutes a 'robust methodology' and that 'findings demonstrate' AR can significantly enrich physical experience and preserve cultural memory. The executed portion is limited to a single anecdotal demonstration without user evaluation or factual validation.","tokens_in":9069,"tokens_out":3444,"duration_ms":38814,"significance":"The idea of using generative AI and AR to mediate cultural and political memory in public spaces is timely and of potential interest to the digital heritage and human-computer interaction communities. The authors are honest in acknowledging that DALL-E images are imaginative rather than historically accurate, and they mention known LLM limitations. However, the significance of the paper as a research contribution is currently low: the central claim—that the methodology enhances place identity or preserves cultural memory—is not supported by the reported evidence. No controlled evaluation, measurable outcome, or comparison with existing AR heritage applications is provided. If the paper were reframed as a speculative position piece, it might be of interest; in its present form it does not meet the standard for a research paper.","major_comments":[{"comment":"The conclusion states, 'we have outlined a robust methodology to enhance the identification and appreciation of place identity' and 'Our findings demonstrate that AR can significantly enrich the physical experience of a location.' The only evidence is a single informal trial at Monastiraki Square: ChatGPT answered descriptive questions and DALL-E generated three images. There is no user study, no baseline, no quantitative or qualitative measure of 'enrichment,' and no analysis of whether place identity was actually affected. The words 'robust,' 'demonstrate,' and 'findings' overstate the evidentiary value of an anecdotal demonstration.","section":"Conclusion (§4) and §4.2"},{"comment":"There is an internal contradiction between the assertion in §3.3 that ChatGPT has 'profound understanding of historical facts' and 'comprehensive knowledge of the past' and the acknowledgment in §4.1 that 'the provided images are not derived from accurate historical images and sources' and that 'the main issue with these technologies lies in their limited accuracy in generating photographs.' The methodology has no verification or sourcing step. §4.2 demonstrates that ChatGPT cannot produce historical illustrations and that Gemini does not even recognize Monastiraki Square. Without verification, the AR overlay would present unverified generative content as cultural memory, undermining the paper's core promise. The paper's own limitation statements confirm that the content layer is unreliable.","section":"§3.3 compared to §4.1"},{"comment":"The Monastiraki case study shows only that a popular LLM recognizes a famous square and can generate evocative images from prompts. This is unsurprising given training data and does not validate the proposed methodology. Moreover, the theoretical framework in Section 2 does not connect to the implemented pipeline: 'place identity' and 'memory' are never operationalized or measured, so there is no way to evaluate whether the pipeline achieves its goal or to compare it with existing AR heritage applications such as ARCHEOGUIDE.","section":"§4.2 and §2"}],"minor_comments":[{"comment":"The text refers to 'the representation of these eras as shown in Figure 2,' but Figure 2 is the methodology diagram; the era images appear in Figure 3. This cross-reference error should be corrected.","section":"§4.2"},{"comment":"A full paragraph describing Apple Vision Pro features is repeated verbatim twice within the same section, and the second occurrence begins with 'All these features...' followed by another repetition. This is likely a copy-paste error.","section":"§3.2"},{"comment":"The heading '3.3' is used twice: 'LLMs and Generative AI in the infrastructure world' and immediately after 'LLM Tools to describe place identity.' The section numbering and titles need revision.","section":"§3.3"},{"comment":"The abstract mentions 'Visual Reality' instead of 'Virtual Reality.' Also, the keyword 'Visual Reality' appears to be a typo for 'Virtual Reality.'","section":"Abstract"},{"comment":"References 18 and 26 appear to describe the same ARCHEOGUIDE project with different titles and venues. The citation style is inconsistent (e.g., [Google Scholar] label in reference 13).","section":"References"}],"recommendation":"reject","confidential_remarks":"This manuscript is far below the standard for a research publication. The central claim is unsupported by the evidence, and the manuscript shows signs of unfinished editing, including duplicated paragraphs, repeated headings, and unclear figure references. A serious revision would require new empirical work (e.g., a user study with measurable outcomes and a historical-verification protocol), not merely textual changes. The idea is not without merit, but the current paper is better suited as a short position statement than as a refereed research article."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a concept note, not a completed research result. The authors walk through a simple pipeline: feed a Google Street View image to ChatGPT, ask it to describe a square's history, have DALL-E generate images for ancient, Byzantine, and speculative future eras, and (in principle) overlay these on an Apple Vision Pro headset. That last step is never implemented — the paper only quotes Apple's spec sheet. There is no user study, no baseline, no metric, and no measurement of any effect on place identity. The conclusion's claim of a \"robust methodology\" and findings that \"demonstrate\" deeper engagement is not supported by the evidence.\n\nThat said, the paper earns some credit. The authors are unusually candid about the technology's limits. They state explicitly that the DALL-E images are \"imaginative representations\" not derived from historical sources, and they report a concrete failure case: Gemini fails to recognize Monastiraki Square at all. The conceptual discussion of place identity draws on a reasonable slice of the heritage-AR literature, including Archeoguide and later AR heritage projects. The pipeline is transparent and easy to reproduce, though the exact prompts and model versions are not reported, which limits reproducibility.\n\nThe soft spots are the load-bearing ones. The central claim — that LLM-generated content can anchor or convey cultural and political memory — depends on the content being faithful. The paper itself admits that this is not the case: the generated images are not accurate, and there is no verification or sourcing stage in the methodology. So the tool would present fabricated content as historical memory. The authors mention that ChatGPT can point to websites for further reading, but that is not part of the pipeline. The paper also contains duplicated paragraphs (Section 3.2 is pasted twice), and the abstract overpromises by mentioning digital twins and sustainability, which never appear in the methodology.\n\nWho is this for? Possibly readers interested in early-stage demos of LLM+AR for heritage interpretation, or as a workshop submission. It is not a serious research paper, and it does not deserve referee time at a journal. As a demo report, it would be fine if the authors reduced the claims to what was actually done: a short experiment showing that ChatGPT can answer questions about a famous square and DALL-E can produce artistically plausible but historically unreliable era images. I would not cite this in my own work, and I would not bring it to a reading group as a substantive contribution — though the candid limitation statements might be useful as a teaching example of overclaiming in AI-for-heritage papers.","headline":"A candid, well-scoped demo note about a ChatGPT/DALL-E/AR pipeline for visualizing place identity, but the central claim is asserted rather than demonstrated — not a research paper.","tokens_in":9786,"tokens_out":1584,"would_cite":false,"duration_ms":19815,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper proposes that an AR headset paired with ChatGPT and DALL-E can overlay a place's ancient, Byzantine, and future selves onto the live urban view, strengthening collective place identity through a demonstration on Monastiraki Squar","keywords":["place identity","augmented reality","large language models","cultural memory","digital twins","urban heritage","generative AI","Monastiraki Square"],"falsifier":"Take a public square with well-documented archival images from several eras, run the pipeline's prompt set through the same LLM and image generator, and compare the generated images and narratives against the archival record; if the generated content consistently deviates in identifiable ways (e.g., anachronistic architecture or wrong rulers) for well-documented sites, the central claim that this enhances authentic memory collapses. A simpler controlled experiment: have participants use the AR overlay and measure whether their stated place identity and knowledge of the square's history increas","tokens_in":8717,"feed_emoji":"🏛️","tokens_out":3876,"duration_ms":41601,"temperature":0.7,"pith_summary":"Taking Monastiraki Square in Athens as its test case, this paper argues that augmented reality headsets paired with large language models can make a place's layered history — and one speculative future — visible on top of the present scene, and that doing so strengthens the public's sense of place identity. The proposed pipeline passes a current street-view image to ChatGPT for recognition and historical narration, uses DALL-E to generate images of the square in ancient, Byzantine, and 2074 states, and overlays these on the live AR view through the Apple Vision Pro. The authors claim this enriches the physical experience and acts as a repository of cultural memory. The evidence offered is a single anecdotal demonstration, with the authors acknowledging that the generated images are imaginative rather than historically accurate.","feed_headline":"AI overlays a square's past and its 2074 future onto live view","feed_subtitle":"The proposal couples ChatGPT and DALL-E with the Apple Vision Pro to make cultural memory visible; historical accuracy remains the open ques","key_machinery":"The pipeline is the central machinery: (1) capture a current view via Google Street View, (2) prompt ChatGPT to recognize the place and answer historical/cultural questions, (3) feed those answers into DALL-E to generate images for selected time periods, (4) overlay the results onto the live view using Apple Vision Pro's spatial mapping, eye tracking, and hand-gesture interaction, and (5) loop user feedback for adjustment. Place identity is the conceptual object: a mix of physical features, cultural associations, and collective memory that the authors argue can be appreciated and evaluated through this loop.","core_discovery":"The central claim is that the identity of a place can be surfaced and revitalized by joining generative AI to augmented reality: an LLM recognizes the space, supplies historical and cultural narration, and drives an image generator to create era-specific visuals, which an AR device then pins to the live environment. This, the authors maintain, lets users experience the evolution of a place and deepens identification with it. The demonstration on Monastiraki Square shows ChatGPT answering questions about the square's mixed architectural heritage and producing DALL-E images labeled ancient, Byzantine, and future; the authors stress that these are not archival reconstructions but imaginative re","pith_inferences":["A key unstated consequence is the risk of fabricated memory: since the generated images are not sourced from archives, the pipeline may replace documented history with plausible-looking invention, especially for less-documented places; a verification step is needed and is not in the methodology.","The demo's own failure (Gemini not recognizing the square, ChatGPT unable to produce historical imagery) hints that model accuracy, not AR hardware, is the bottleneck; this could be tested by swapping different LLMs into the pipeline and scoring their historical outputs against archival records.","The approach could be extended to contested or traumatic sites, where political memory is itself the subject; whether AI-generated futures help or distort reconciliation is a question worth investigating but the paper does not address.","The authors' suggestion to fine-tune an open LLM for space identity implies a path toward domain-specific models, which could be evaluated with a dataset of urban squares and their documented histories."],"forward_implications":["If the pipeline works as intended, residents and tourists can engage with a square's layered history in situ, making cultural heritage part of the everyday visual field rather than something confined to plaques or museums.","Planners and heritage professionals could use the same loop to test how proposed urban changes affect the perceived identity of a place by overlaying future scenarios before construction.","The method positions LLMs as repositories of cultural and political memory, suggesting a new role for generative models in heritage conservation and public history.","Successful integration would blur the line between digital twin and live environment, making the 'digital twin of a space' a real-time, narratively driven overlay rather than a static model."],"fun_headline_variants":["AR plus LLMs lets you see a place's past and future","ChatGPT and DALL-E bring a square's past and future to AR","LLM-driven AR overlays a place's history and imagined future","Generative AI and AR make cultural memory visible in place","Augmented reality plus AI reveals a square's evolving identity"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The pipeline's value rests on the assumption that ChatGPT and DALL-E produce historical and cultural content that is faithful enough to the place's real past and future plausibility — the paper provides no verification step and even concedes the images are imaginative, so if the models fabricate, the tool would preserve invented memory rather than cultural memory.","fun_headline_variants_meta":{"raw":{"variants":["AR plus LLMs lets you see a place's past and future","ChatGPT and DALL-E bring a square's past and future to AR","LLM-driven AR overlays a place's history and imagined future","Generative AI and AR make cultural memory visible in place","Augmented reality plus AI reveals a square's evolving identity"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000924,"raw_usage":{"total_tokens":3829,"prompt_tokens":809,"completion_tokens":3020,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":2931}},"tokens_in":553,"tokens_out":3020,"duration_ms":22319,"temperature":1.0,"reasoning_tokens":2931,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T11:11:47.191190+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a public square with well-documented archival images from several eras, run the pipeline's prompt set through the same LLM and image generator, and compare the generated images and narratives against the archival record; if the generated content consistently deviates in identifiable ways (e.g., anachronistic architecture or wrong rulers) for well-documented sites, the central claim that this enhances authentic memory collapses. A simpler controlled experiment: have participants use the AR overlay and measure whether their stated place identity and knowledge of the square's history increas","supporting_citations":[],"review_version":1}