{"id":"07d71ad8-67f3-4340-b156-d95a33c575f4","arxiv_id":"2412.14418","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"A new large-scale multi-season, multi-elevation, real-world image dataset of a university campus, with a described calibration pipeline.","lead":"This paper presents a real-world dataset of the Johns Hopkins Homewood Campus with 12,300+ images taken across seasons, times of day, and elevations, plus a calibration pipeline that aligns phone and drone imagery into one coordinate system. It is intended as a benchmark for testing 3D reconstruction algorithms under appearance and viewpoint variation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The unified campus claim rests on an unvalidated 60m summer aerial anchor (Sec. 4.3, Eq. 1); no quantitative localization or drift metric is reported, so all building alignments inherit an unknown error.","rationale":"The reader and I converge on the same load-bearing step: the summer 60m aerial anchor. I did not find an internal contradiction in the pipeline description, and temporal adjacency (Sec. 4.1) is a reasonable safeguard even if its effect is unmeasured; if the anchor fails, no amount of correct per-building registration can produce the claimed unified coordinate system. The paper gives a plausible multi-stage procedure but no quantitative validation anywhere: Table 2 only counts registered images for feature matchers and says nothing about pose accuracy, and the global alignment section contains no error metrics. Since the strongest claim is about enabling rigorous reconstruction research and about the reconstructed campus, the absence of anchor verification is the single most consequential gap. A conditional verdict is therefore appropriate, with release and independent localization checks as the condition; I would not change the reader's verdict.","tokens_in":7832,"tokens_out":5226,"duration_ms":44878,"concrete_test":"On release of the data (or from the authors' intermediate outputs), run COLMAP on the summer 60m aerial subset alone and report the anchor quality: fraction of images registered, median reprojection error, and relative poses between overlapping building pairs. Compare those relative poses against an external reference, e.g., campus GIS building footprints, known inter-building distances, or drone GPS logs, and compute the drift per building across the 80,000 m^2 area. If the anchor has drift above roughly 1 m in translation or 1 degree in rotation between adjacent buildings, the Procrustes alignment in Eq. 1 will not yield a coherent campus; additionally, re-estimate Eq. 1 leaving out one building at a time to check whether the other buildings' poses move significantly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of a 'coherent, large-scale sparse reconstruction' depends on Sec. 4.3: a subset of summer 60m aerial images is asserted to be 'reliably calibrated' into a campus-wide coordinate system, and every building is then transformed into that frame via Procrustes alignment of camera centers (Eq. 1). The paper reports no verification of this anchor frame: no number of registered aerial images, no reprojection error, no loop-closure statistics, no comparison against GPS/RTK or known campus distances, and no sensitivity analysis of the Procrustes fit. Because Eq. 1 uses only camera-center positions, any rotation, translation, or especially scale drift in the anchor is propagated identically into all ten buildings; a small per-building orientation error would break the coherent campus model while remaining hard to spot in a sparse point-cloud figure. This is a verification gap rather than an internal inconsistency, but it is load-bearing: the dataset's benchmark value is predicated on poses being correct, and the paper's own evidence does not yet distinguish a correctly aligned campus from one with substantial inter-building misalignment.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces a real-world dataset of over 12,300 images of ten buildings on the Johns Hopkins University Homewood Campus, acquired over one year across four seasons, daytime and nighttime, and elevations from ground level to 120m. The authors propose a three-stage calibration pipeline: temporal-adjacency-constrained matching for ground-level video frames to mitigate doppelganger matches, integration of drone ascending sequences to bridge ground and aerial perspectives, and Procrustes alignment of per-building reconstructions into a campus-wide anchor frame defined by summer 60m aerial images. The central claim is that this pipeline produces a coherent large-scale sparse reconstruction, and that the dataset enables benchmarking of reconstruction methods under appearance, scale, and viewpoint variation. Registration counts are compared against SIFT, SuperGlue, LoFTR, and RoMA, and qualitative sparse point-cloud visualizations are provided.","tokens_in":8071,"tokens_out":6992,"duration_ms":54928,"significance":"If the poses and alignments are accurate, the dataset fills a genuine gap: existing reconstruction benchmarks are either small-scale or controlled, single-elevation, or lack per-acquisition appearance consistency. The multi-season, multi-elevation coverage with organized appearance groups is valuable for evaluating NeRF/3DGS methods and structure-from-motion under appearance change. The pipeline ideas—temporal adjacency to avoid doppelganger matches and ascending sequences to connect ground and aerial views—are sensible and potentially useful. However, the paper currently substantiates the central claim only with qualitative figures and registration counts; no quantitative pose accuracy, reprojection error, external georeferencing check, or reconstruction benchmark is reported. The contribution is therefore conditional on additional verification.","major_comments":[{"comment":"The campus-wide coordinate system rests entirely on the assertion that the summer 60m aerial subset is 'reliably calibrated,' yet no quantitative evidence is given: the paper reports no number of aerial images used in the anchor, no reprojection error, no loop-closure statistics, no comparison with GPS/RTK or known campus distances, and no sensitivity analysis of the Procrustes fit. Because Eq. (1) aligns only camera-center positions, any scale, rotation, or translation error in the anchor propagates identically into all ten building reconstructions. Please provide quantitative validation of the anchor and of the final inter-building alignment, such as residuals of the Procrustes fit, known-distance checks, or pose error against an independent survey.","section":"4.3, Eq. (1)"},{"comment":"The doppelganger-mitigation claim is supported only by the number of images that register, not by whether the registrations are geometrically correct. Table 2 shows that several methods register all or nearly all images (e.g., LoFTR for Ames, Clark, Garland, Hackerman, and Mason), so the count alone cannot distinguish correct alignment from visually plausible but wrong matches. Please report pose accuracy against known building geometry, loop-closure consistency, or a quantitative comparison of reconstructions with and without the k=10 temporal adjacency constraint. In addition, Table 2's 'G' and 'D' columns are not consistently populated, and some rows show registered counts exceeding the stated number of images, making the comparison hard to interpret.","section":"4.1, Table 2"},{"comment":"The paper's central claim of 'a coherent, large-scale sparse reconstruction' is demonstrated only through sparse point-cloud figures. No reconstruction benchmark is run on the dataset, and no metric such as mean reprojection error, track length, pose uncertainty, or novel-view synthesis quality is reported. For a dataset intended to be a benchmark, calibration quality must be quantified; otherwise readers cannot tell whether downstream failures are due to the data or to the algorithm. Please add calibration statistics and, ideally, baseline reconstruction results (SfM/NeRF/3DGS) that use the provided poses.","section":"5 and Figures 1-2"},{"comment":"The manuscript does not state where the dataset will be hosted, under what license, or what metadata are included (timestamps, elevation tags, building labels, GPS where available). For a dataset paper, these release details are essential for the claimed community impact. Please include a release plan and a brief data-card-style description of the files, formats, and intended usage.","section":"Dataset release (all sections)"}],"minor_comments":[{"comment":"The header 'mA mV Elevation' does not define the abbreviations, and the 'mE' property from the text is not made explicit in the table. Also, UrbanScene3D and Quad 6K are both cited as [4], which appears to be a citation error since [4] is the Crandall et al. SfM paper.","section":"Table 1"},{"comment":"The statement 'All methods fail to register cross-view images correctly' is not supported by Table 2, where several methods register a large fraction of images; please clarify what 'correctly' means and whether the comparison withholds ascending sequences for all methods.","section":"4.2"},{"comment":"Eq. (1) uses the notation C^i_hall while the text defines the building-wise coordinate system as C^i_building; please make the notation consistent.","section":"4.3"},{"comment":"There are minor typographical issues, e.g., 'welldesigned' in the abstract and 'time highlight' in the Figure 2 caption; a careful proofread would improve readability.","section":"Abstract and Introduction"},{"comment":"A limitations section is missing; in particular, the manuscript should acknowledge that the anchor accuracy is not quantitatively verified and that PII blurring may affect reconstruction quality in some regions.","section":"Conclusion"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a reasonable dataset contribution in scope for this venue, but the reviewers should insist on quantitative calibration validation, a clarified Table 2, and dataset release details before acceptance. The absence of a limitations section and the under-specified global-alignment verification are the main concerns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Plainly: this dataset fills a real gap, but the paper's headline claim—a coherent campus-wide reconstruction—is not supported by the numbers it reports. The stress-test is right: the 60m summer aerial anchor is asserted as 'reliably calibrated' with no verification, and the Procrustes step (Eq. 1) propagates any scale or rotation error in that anchor to all ten buildings.\n\nWhat's genuinely new: the combination of real-world, large-scale, ground-plus-aerial imagery with consistent multi-view per appearance across seasons and times of day. OMMO is aerial-only, MatrixCity is synthetic, and MegaScenes lacks consistent multi-view. The ascending drone sequences are a clever bridge between ground and aerial perspectives, and the temporal adjacency constraint for doppelganger mitigation is a sensible practical choice.\n\nThe soft spots are proportionate but real. No data or code is released, so the 12.3K images cannot be used today. The anchor frame gets no quantitative check: no reprojection error, loop-closure statistics, GPS/RTK comparison, or even the number of aerial images that went into it. Table 2 reports how many images each matcher registered, but not whether the registrations are correct—a method could register everything by joining the wrong places. The paper also skips any downstream reconstruction benchmark, so we do not see whether the claimed poses actually support NeRF or 3DGS. These are verification gaps, not internal contradictions; the pipeline is plausible and the thinking is clear. But for a dataset paper, the data and the validation are the product, and both are missing here.\n\nWho gets value: researchers who want a benchmark for multi-appearance large-scale reconstruction with mixed ground and aerial views. If the authors release the data with pose-error numbers and a few reconstruction baselines, this becomes a real resource. As it stands, I would treat the claims as conditional.\n\nMy recommendation: send it to peer review—it deserves a serious look—but the revision must include data release and quantitative anchor validation. I would not cite it yet.","headline":"A genuinely useful dataset idea that is currently under-validated: the campus-wide alignment rests on an unchecked 60m anchor, and the paper ships no data or quantitative pose error.","tokens_in":8568,"tokens_out":2442,"would_cite":false,"duration_ms":17669,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces a 12,300-image, multi-season, multi-elevation campus dataset and a calibration pipeline that registers all of it into one campus-wide reconstruction.","keywords":["3D reconstruction dataset","multi-season imagery","multi-elevation capture","neural rendering benchmark","structure-from-motion calibration","doppelganger mitigation","unconstrained scene reconstruction"],"falsifier":"Survey a set of building corners with RTK GPS and compare them with the reconstructed camera and point positions after Procrustes alignment; if the median alignment error exceeds the pixel-projection tolerance that downstream rendering requires, the claim that the 60-meter aerial anchor is reliable collapses.","tokens_in":7683,"feed_emoji":"🏛️","tokens_out":7735,"duration_ms":62606,"temperature":0.7,"pith_summary":"This paper introduces a large-scale, real-world imagery collection of ten adjacent buildings on a university campus, with over 12,300 images taken across four seasons, day and night, different weather, and elevations from ground level to 120 meters. The authors argue that existing reconstruction benchmarks split up the properties a full evaluation needs: some offer appearance variation, others scale, others aerial or ground views, but none combine multi-appearance, multi-view, and multi-elevation in one real-world setting. The contribution is the dataset itself plus a multi-stage calibration pipeline that registers all images into a single campus-wide coordinate system, producing a coherent sparse reconstruction. If the claim holds, researchers can evaluate neural rendering and structure-from-motion methods under inconsistent illumination, repeated architecture, and extreme viewpoint changes without needing pixel-level test-view supervision.","feed_headline":"A 12,300-image campus dataset spans seasons, times, and altitudes","feed_subtitle":"Ten adjacent buildings, ground to 120 meters up, give 3D reconstruction a real-world multi-appearance benchmark.","key_machinery":"The carrying mechanism is the dataset's acquisition design combined with a three-stage calibration pipeline. The temporal adjacency constraint matches each ground image only to its 10 nearest video frames, cutting off the long-range matches that let visually similar front and back doors of a building collapse into one location; this is the doppelgänger mitigation. Ascending drone sequences, shot from ground level up to about 60 meters, give feature matchers an incremental perspective bridge between ground and aerial imagery. Finally, Procrustes alignment solves for a similarity transform that registers each building's local reconstruction onto an anchor coordinate system built from summer 60-meter aerial images, merging the campus into one frame. The dataset's defining structure is that every appearance condition is captured across many views, so a method can be handed a timestamp and asked to render held-out views gathered at that same time.","core_discovery":"The central claim is that a carefully planned acquisition — one multi-view appearance set per season, time of day, and weather condition per building, with handheld ground videos, circular drone flights at 60, 100, and 120 meters, and ascending drone sequences — supplies the missing benchmark for holistic scene reconstruction. The paper further claims that its calibration approach, which restricts ground-image matches to the 10 nearest video frames to suppress doppelgänger matches, uses ascending sequences to connect ground and aerial perspectives, and aligns each building's reconstruction to a summer 60-meter aerial anchor through Procrustes alignment, yields a coherent large-scale campus reconstruction at reasonable processing cost. The dataset is released as a testbed where repeated architectural motifs and appearance shifts make calibration genuinely hard rather than controlled away.","pith_inferences":["An independent geolocation check, such as surveyed ground-control points, would turn the asserted reliability of the summer 60-meter aerial anchor into a measured error; the paper does not report one.","Since the temporal constraint exploits known video order, a natural extension is to test whether global-context learned matchers can drop that requirement, or to use this dataset to train doppelgänger-robust matching.","The repeated appearance sets make the dataset a plausible training ground for time-conditioned appearance models, an evaluation the paper does not itself run.","The campus's uniform architectural style is a deliberate difficulty but also a domain restriction; conclusions drawn here may not transfer to heterogeneous city scenes without separate validation."],"forward_implications":["Reconstruction methods can now be tested on a real large-scale scene where illumination and appearance change while multi-view consistency is preserved, so held-out test views can be rendered from time metadata alone.","Structure-from-motion and feature-matching pipelines face a public stress test with repetitive architecture, where global matching fails and the paper's temporal constraint is shown to restore stable registration.","Multi-elevation coverage makes rooftop and upper-facade reconstruction a measurable benchmark instead of an unobserved region.","The 12,300-image collection supports fair comparison of appearance-conditioned neural radiance fields and Gaussian splatting variants on identical geometry.","Per-building registration followed by global alignment offers a template for scaling calibration to city-sized image collections."],"supporting_citations":[{"why":"supplies the existing appearance-diverse photo-tourism benchmark that lacks controlled multi-view appearance sets","marker":"[22]"},{"why":"provides a large ground-level scene benchmark with appearance diversity but no per-appearance multi-view structure","marker":"[25]"},{"why":"offers a large-scale aerial dataset with multi-appearance and multi-view, limited to a single elevation","marker":"[15]"},{"why":"gives the synthetic ground-plus-aerial comparison that motivates a real-world counterpart","marker":"[10]"},{"why":"represents large-scale street-level city reconstruction, lacking aerial coverage of roofs","marker":"[24]"},{"why":"supplies the Procrustes alignment method used to merge building reconstructions into the campus anchor frame","marker":"[6]"},{"why":"serves as a learned feature-matching baseline that fails on the dataset's doppelgänger imagery, motivating the temporal constraint","marker":"[21]"}],"fun_headline_variants":["Campus 3D dataset: 12k images, all seasons, 120m up","Multi-season, multi-elevation benchmark for 3D reconstruction","Ground-to-120m campus dataset: 12k images, seasons, times","Real-world 3D reconstruction testbed: multi-appearance campus","12,300-image campus benchmark spans seasons and altitude"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The campus-wide alignment rests on the assumption that the summer 60-meter aerial images are reliably calibrated and therefore form a trustworthy anchor; the paper asserts this reliability but gives no quantitative check of the anchor's accuracy.","fun_headline_variants_meta":{"raw":{"variants":["Campus 3D dataset: 12k images, all seasons, 120m up","Multi-season, multi-elevation benchmark for 3D reconstruction","Ground-to-120m campus dataset: 12k images, seasons, times","Real-world 3D reconstruction testbed: multi-appearance campus","12,300-image campus benchmark spans seasons and altitude"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000997,"raw_usage":{"total_tokens":4159,"prompt_tokens":817,"completion_tokens":3342,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":433,"completion_tokens_details":{"reasoning_tokens":3245}},"tokens_in":433,"tokens_out":3342,"duration_ms":22491,"temperature":1.0,"reasoning_tokens":3245,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:14:59.345431+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Survey a set of building corners with RTK GPS and compare them with the reconstructed camera and point positions after Procrustes alignment; if the median alignment error exceeds the pixel-projection tolerance that downstream rendering requires, the claim that the 60-meter aerial anchor is reliable collapses.","supporting_citations":[{"cited_title":"Seitz, and Richard Szeliski","cited_arxiv_id":null,"evidence_quote":"supplies the existing appearance-diverse photo-tourism benchmark that lacks controlled multi-view appearance sets"},{"cited_title":"Megascenes: Scene-level view synthesis at scale","cited_arxiv_id":null,"evidence_quote":"provides a large ground-level scene benchmark with appearance diversity but no per-appearance multi-view structure"},{"cited_title":"A large-scale outdoor multi- modal dataset and benchmark for novel view synthesis and implicit scene reconstruction","cited_arxiv_id":null,"evidence_quote":"offers a large-scale aerial dataset with multi-appearance and multi-view, limited to a single elevation"},{"cited_title":"Matrixcity: A large- scale city dataset for city-scale neural rendering and beyond","cited_arxiv_id":null,"evidence_quote":"gives the synthetic ground-plus-aerial comparison that motivates a real-world counterpart"},{"cited_title":"Mildenhall, Pratul P","cited_arxiv_id":null,"evidence_quote":"represents large-scale street-level city reconstruction, lacking aerial coverage of roofs"},{"cited_title":"Generalized procrustes analysis","cited_arxiv_id":null,"evidence_quote":"supplies the Procrustes alignment method used to merge building reconstructions into the campus anchor frame"},{"cited_title":"Superglue: Learning feature matching with graph neural networks","cited_arxiv_id":null,"evidence_quote":"serves as a learned feature-matching baseline that fails on the dataset's doppelgänger imagery, motivating the temporal constraint"}],"review_version":1}