{"id":"cf5931d7-a6dd-4e5d-a677-81ee5da3e2eb","arxiv_id":"2508.14507","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"DeepTelecom provides a multimodal LoD3 digital-twin channel dataset generated via LLM-assisted scene modeling and Sionna ray tracing.","lead":"DeepTelecom is a new 3D digital-twin channel dataset for wireless AI, built with an LLM-assisted scene modeling pipeline and GPU ray tracing. It packages synchronized images, videos, channel tensors, and path data as a proposed benchmark for 6G foundation models.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fidelity claim depends on unvalidated geometry/material annotations feeding Sionna; no measured-channel or independent-tracer comparison, and only partial release, so 'high fidelity' is currently unsupported.","rationale":"The reader's weakest assumption—unvalidated Sionna ray tracing on LLM-annotated LoD3 twins—is exactly the load-bearing point. The paper's novelty is a pipeline, and the pipeline is plausible, but the central value proposition is that the resulting dataset is high-fidelity and physically grounded. That requires validation against either measured channels or an independent, trusted simulator on the same scenes, plus enough release/statistics for the community to verify. The paper provides none of these. I do not see a stronger internal flaw: the equations for Fresnel reflection, propagation factor, and CIR/CFR are standard, and the use of Sionna is a reasonable engineering choice. The 'partial availability' statement is another independent obstacle to the 'complete open dataset' claim, but it reinforces rather than replaces the validation gap. A conditional acceptance with release-and-validation requirements is the right call, matching the reader's verdict, so no change is needed.","tokens_in":8451,"tokens_out":3256,"duration_ms":42303,"concrete_test":"Take one released outdoor and one indoor scene, and run the identical DeepTelecom configuration through Sionna and through an independent reference ray tracer (or compare with published channel-sounder measurements at matched RX locations). Report path-loss RMSE, delay-spread CDF, and dominant-path delay error. If the median path-loss error exceeds roughly 6 dB, or if the two tracers disagree on dominant path structure, the 'high-fidelity' claim is not supported; if they agree within a few dB, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that DeepTelecom is a high-fidelity, physically accurate LoD3 digital-twin channel dataset. For that to hold, two links must both be secure: (1) the reconstructed geometry and LLM-assigned EM materials faithfully represent the real scene, and (2) Sionna's ray tracer, fed that scene, outputs physically correct CIR/CFR. Neither link is quantitatively demonstrated. Section III-A.3 states that an LLM assigns materials in bulk from a reference table, but the table, mapping rules, and consistency checks are not reported; wrong or overly homogeneous material assignments can shift path amplitudes by several dB, which is exactly the kind of error a training dataset would silently encode. Section III-C invokes standard Fresnel/UTD physics, but no comparison is made against measured channel data or against an independent reference tracer (e.g., Wireless InSite/Volcano) on the same scene, so there is no evidence that the reconstructed digital twin's outputs match reality. Section IV gives only qualitative scale claims—'thousands of transmitter-receiver pairs and millions of raw ray paths'—with no histograms, no scenario statistics, no validation metrics, and no downstream benchmark, so even internal consistency and usability are unverified. Finally, the paper states the dataset is 'partially available online'; the 'complete open multimodal dataset' at the core of the abstract and Section V cannot yet be independently inspected. None of these points is an internal contradiction; they are missing evidence for the paper's strongest claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces DeepTelecom, a claimed LoD3 digital-twin channel dataset generation pipeline. It uses an LLM-assisted workflow to annotate reconstructed indoor (LiDAR/robot) and outdoor (OpenStreetMap/Google 3D Tiles) scenes with material properties, and then employs NVIDIA Sionna's GPU ray tracer to compute propagation paths, CIR/CFR, coverage heatmaps, and synchronized images/videos. The authors assert that this provides a large-scale, high-fidelity, multimodal benchmark for wireless AI research, and state that the dataset is 'partially available online.' The paper reports no quantitative validation of the generated channels and only gives qualitative scale claims.","tokens_in":1461,"tokens_out":1507,"duration_ms":45083,"significance":"If substantiated, DeepTelecom would be a useful contribution to the wireless-AI dataset landscape, combining material-aware LoD3 scene reconstruction with GPU-accelerated ray tracing and multimodal outputs. The use of open-source Sionna and a modular pipeline are commendable and could lower the barrier for generating diverse channel data. However, the paper's central claim of 'high fidelity' and 'physically accurate' channel generation is currently unsupported: there are no comparisons against measured data, no independent ray-tracer cross-checks, no quantitative dataset statistics, and only partial data release. These gaps must be addressed before the resource can serve as a trustworthy benchmark or training substrate.","major_comments":[{"comment":"The experimental analysis is purely qualitative. The only scale statement is that 'each scene yields thousands of transmitter-receiver pairs and millions of raw ray paths,' but no concrete statistics are given: number of scenes, frequency bands, bandwidths, number of TX-RX pairs per scenario, ray counts, material distributions, or delay/angular spreads. Without histograms or summary tables, the claimed 'diversity' and 'large scale' cannot be assessed. Please include a quantitative dataset card with scenario counts, parameter ranges, and per-scenario output sizes.","section":"Section IV"},{"comment":"The LLM-assisted material assignment is load-bearing for channel fidelity, yet the reference table of electromagnetic material properties and the prompting/validation rules are not reported. The text states that the LLM assigns materials 'in bulk' from a reference table, but without that table and a consistency-check procedure, a reader cannot reproduce the material annotations or judge whether they are physically plausible. Since erroneous or overly homogeneous permittivity/permeability values directly alter reflection/transmission coefficients in Eqs. (5)-(6), this missing information undermines the physical-accuracy claim.","section":"Section III-A.3"},{"comment":"No validation of the ray-tracing output is provided. The paper invokes standard Fresnel and UTD physics, but there is no comparison against measured channel data or an independent ray tracer (e.g., Wireless InSite, Volcano, or even Sionna's own reference implementation on a canonical scene). The fidelity claim requires that both the reconstructed geometry and the LLM-assigned materials produce physically correct paths. Please add at least a small validation set: e.g., path-loss versus distance for a simple corridor/street scene, delay spread and angular distributions for a known environment, or a comparison against published measurements. Without this, 'high fidelity' is an assertion rather than a demonstrated property.","section":"Section III-C"},{"comment":"The paper claims 'the complete open multimodal dataset' and 'complete open multimodal dataset that fuses visual, tensor, and tabular views,' but the footnote says the dataset is 'partially available online.' For a dataset paper, full availability of the core HDF5/CSV outputs and scene files is essential for reproducibility and independent verification. Please clarify exactly which components are publicly downloadable, which are withheld, and what the release timeline is. The current contradictory wording weakens the dataset contribution.","section":"Abstract and Section V"}],"minor_comments":[{"comment":"Typo: 'which is is given by' should be 'which is given by.'","section":"Eq. (8)"},{"comment":"Minor language issues: 'a large language model' should be 'an LLM' or 'a large language model' read correctly; 'editting' should be 'editing.'","section":"Section III-A.3 and Section III-A.1"},{"comment":"The array response vectors a_r(Omega_r,l) and a_t(Omega_t,l) are used but never defined. Please define them, including the array geometry and element pattern assumptions, since MIMO is a central application.","section":"Eq. (7)"},{"comment":"The symbol N_int is not explicitly defined in the text near Eq. (4). Define it as 'number of interactions so far' for clarity.","section":"Section III-C, Eq. (4)"},{"comment":"The caption says 'The overall DeepTelecom framework,' but Figure 1 actually illustrates the four-module workflow. Consider rewording to 'Framework diagram of the DeepTelecom data-generation pipeline.'","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"This is a dataset-announcement paper whose central claim ('high fidelity') is currently supported only by qualitative statements. The missing validations are additive in nature—comparison against measurements/reference ray tracing, quantitative statistics, and full data release—so the paper can plausibly be revised to meet the standard. However, if the authors cannot provide such evidence or the data remains only partially available, the contribution should be reframed as a pipeline description rather than a validated dataset."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: DeepTelecom is a genuinely useful pipeline paper — it combines LLM-assisted LoD3 scene reconstruction with Sionna's GPU ray tracer and packages synchronized images, videos, HDF5 tensors, and CSV manifests into one resource. That integration is new relative to DeepMIMO, ViWi, WiThRay, and WAIR-D, and the engineering is clearly described. The paper earns credit for shipping a working pipeline with a plausible plan for open data.\n\nWhat it doesn't do is back up the word 'high-fidelity.' There is no comparison against measured channel data or an independent reference tracer on the same scene. The material parameters are assigned by an LLM from a reference table, but the table, mapping rules, and consistency checks are not reported — an error of a few dB in per-surface reflection could be silently baked into every training example. Section IV gives only qualitative scale ('thousands of TX-RX pairs, millions of ray paths') with no histograms, no scenario statistics, and no downstream benchmark. The paper itself admits the dataset is only 'partially available online,' so the central artifact can't be independently inspected. None of this is an internal contradiction; it's missing evidence for the strongest claim.\n\nThe equations (Fresnel, UTD, CIR/CFR) are standard and correctly stated. The ray termination criteria and the Fibonacci sphere sampling are standard. The reliance on Sionna is reasonable — Sionna is a well-tested tool — but that only validates the simulator, not the scene representation fed into it.\n\nThe citation pattern is fine. The only self-citation is [14] for the robot localization platform, and it's not load-bearing. No invented entities, no circularity.\n\nWho is this for? Researchers who want a ready-made multimodal indoor/outdoor channel dataset for wireless AI training or benchmarking. They'll get value once the dataset is actually released and the fidelity question is addressed. Right now the paper is a well-engineered proposal with an unvalidated product.\n\nMy verdict: it deserves a serious referee, not a desk reject. But acceptance should be conditioned on full release plus at least one validation study — measured channels or a credible reference tracer on the same scene — and real dataset statistics. The authors have done the hard engineering; the missing step is proving the output means what they say it means.\n\nRecommendation: send to peer review with a request for major revision.","headline":"A well-engineered dataset pipeline whose 'high-fidelity' claim is currently unsupported by validation or full release; worth refereeing, but acceptance should hinge on evidence.","tokens_in":9353,"tokens_out":1764,"would_cite":false,"duration_ms":18746,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"DeepTelecom couples 3D digital twins with GPU ray tracing to build open wireless AI channel datasets.","keywords":["3D digital twin","wireless AI dataset","ray tracing","large language model","channel modeling","MIMO","multimodal dataset","6G"],"falsifier":"Choose one DeepTelecom scene, place a transmitter and receivers at the published positions, run a channel sounder at the same carrier frequency and bandwidth, and compare the measured power-delay profile, delay spread, and angular spectrum with the dataset's CIR/CFR. If the simulated statistics differ by more than the claimed LoD3 accuracy, the dataset's physical fidelity claim fails.","tokens_in":8388,"feed_emoji":"📡","tokens_out":6988,"duration_ms":67339,"temperature":0.7,"pith_summary":"The paper introduces DeepTelecom, a pipeline that turns detailed 3D digital-twin scenes—indoor LiDAR scans and outdoor urban models from open geospatial data—into large, multimodal wireless channel datasets. Its central claim is that an LLM-assisted workflow can annotate every surface with electromagnetic material properties and then run full ray-tracing propagation on GPU, producing synchronized outputs: rendered images and videos, coverage heatmaps, ray-path trajectories, and standard MIMO channel responses (CIR/CFR) with angles, delays, phases, and Doppler. The authors argue this closes a gap in existing wireless AI corpora, which are slow to produce, low in geometric and material fidelity, and limited in scenario variety. If the claim holds, DeepTelecom supplies a common training and benchmarking substrate for 6G research tasks such as localization, beamforming, and channel prediction.","feed_headline":"Digital twins turn city scans into AI-ready wireless channel datasets","feed_subtitle":"Each scene ships with the 3D model, ray paths, heatmaps, and channel tensors aligned in time and space.","key_machinery":"The central object is the LoD3 (third level of detail) digital twin with segmentable, material-parameterized surfaces, exported as a structured XML scene description. Its load-bearing role is to bridge geometry and electromagnetics: object names like 'window' or 'roadSurface' are bulk-annotated by an LLM into frequency-dependent permittivity and permeability values, which the ray tracer consumes to compute Fresnel coefficients, transmission, and diffraction. Around this object, the pipeline builds MIMO channel tensors via the superposition of path gains times transmit and receive array response vectors, with the ray sampler using golden-ratio Fibonacci sphere sampling to cover directions uni","core_discovery":"DeepTelecom's core claim is that a complete channel dataset can be generated from an LoD3 digital twin that keeps per-surface material semantics from geometry to simulation. The pipeline reconstructs indoor scenes from LiDAR point clouds and outdoor scenes from open 3D tiles, enforces a strict object-naming convention, and uses a large language model to convert the resulting XML scene description into consistent electromagnetic material parameters. A GPU-accelerated ray tracer then launches a near-uniform Fibonacci-sphere distribution of rays, applies Fresnel reflection/transmission and uniform theory of diffraction, prunes rays by interaction depth and power threshold, and records per-path","pith_inferences":["If the LLM assigns a wrong material profile to a surface class, every ray that interacts with that surface inherits the error; a sensitivity analysis that perturbs material parameters and measures CIR/CFR drift would show which annotated classes matter most.","Because the dataset is generated rather than measured, it is suited to controlled studies of scene diversity—for example, training on one city's twin and testing on another—which could quantify how much geometric and material variation a wireless foundation model needs.","The rendered images and videos are synthetic views of the same scene that generated the channels, so the dataset offers a controlled testbed for sim-to-real transfer in vision-aided communication, provided the renderer's camera distribution is aligned with real deployment views."],"forward_implications":["New scenes can be generated on demand from GPS coordinates or point-cloud scans, removing the manual calibration bottleneck that limits existing datasets.","Each scenario package includes the 3D model, simulation configuration, ray paths, heatmaps, and CIR/CFR tensors, so one download supplies both the physical scene and the channel labels for supervised learning.","The synchronized visual and channel data enable vision-aided wireless AI tasks—such as localization from camera images fused with channel state—without separate data collection campaigns.","The same pipeline can produce large numbers of transmitter-receiver pairs and millions of ray paths per scene, enough volume for foundation-model pretraining in 6G physical-layer research."],"supporting_citations":[{"why":"Supplies the GPU-accelerated differentiable ray-tracing engine that produces the propagation paths and channel tensors.","marker":"[13]"},{"why":"Defines the leading open deep-learning channel dataset whose precomputed, single-scenario format DeepTelecom is designed to extend.","marker":"[10]"},{"why":"Provides a ray-tracing simulator for smart wireless environments used as a point of comparison for versatility and scenario coverage.","marker":"[12]"},{"why":"A wireless AI research dataset that represents the prior state of wireless-AI corpora with limited fidelity and scenario diversity.","marker":"[4]"},{"why":"A vision-aided wireless dataset framework whose image-channel pairing DeepTelecom extends with synchronized multimodal output.","marker":"[11]"}],"fun_headline_variants":["Digital twins generate AI-ready wireless channel datasets","LLM-assisted scene builder streams ray paths for MIMO research","City-scale digital twins yield high-fidelity channel data","DeepTelecom: from 3D scenes to synchronized channel tensors","GPU ray tracing powers multimodal wireless AI training data"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The dataset's usefulness rests on the premise that the reconstructed 3D geometry and the LLM-assigned electromagnetic material parameters produce ray paths close to real radio propagation; the paper reports no comparison against measured channel data or an independent reference ray tracer.","fun_headline_variants_meta":{"raw":{"variants":["Digital twins generate AI-ready wireless channel datasets","LLM-assisted scene builder streams ray paths for MIMO research","City-scale digital twins yield high-fidelity channel data","DeepTelecom: from 3D scenes to synchronized channel tensors","GPU ray tracing powers multimodal wireless AI training data"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000169,"raw_usage":{"total_tokens":1088,"prompt_tokens":721,"completion_tokens":367,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":465,"completion_tokens_details":{"reasoning_tokens":288}},"tokens_in":465,"tokens_out":367,"duration_ms":4945,"temperature":1.0,"reasoning_tokens":288,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T18:28:36.006676+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Choose one DeepTelecom scene, place a transmitter and receivers at the published positions, run a channel sounder at the same carrier frequency and bandwidth, and compare the measured power-delay profile, delay spread, and angular spectrum with the dataset's CIR/CFR. If the simulated statistics differ by more than the claimed LoD3 accuracy, the dataset's physical fidelity claim fails.","supporting_citations":[{"cited_title":"Sionna: An open-source library for next-generation physical layer research,","cited_arxiv_id":null,"evidence_quote":"Supplies the GPU-accelerated differentiable ray-tracing engine that produces the propagation paths and channel tensors."},{"cited_title":"WiTh- Ray: A versatile ray-tracing simulator for smart wireless environments,","cited_arxiv_id":null,"evidence_quote":"Provides a ray-tracing simulator for smart wireless environments used as a point of comparison for versatility and scenario coverage."},{"cited_title":"ViWi: A deep learning dataset framework for vision-aided wireless communications,","cited_arxiv_id":null,"evidence_quote":"A vision-aided wireless dataset framework whose image-channel pairing DeepTelecom extends with synchronized multimodal output."}],"review_version":1}