{"id":"aaf4a1cd-3837-4451-9222-3c8f715b176e","arxiv_id":"2505.23730","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"DTBIA is a VR visual analytics system for exploring functional and structural Digital Twin Brain data, validated by qualitative expert case studies.","lead":"This paper presents DTBIA, a virtual reality system for exploring Digital Twin Brain simulation data, letting researchers move from whole brain regions down to voxels and slices. The authors report two expert case studies suggesting the immersive view helps neuroscientists compare simulated and real brain activity more intuitively.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'strong evidence' claim depends on two qualitative case studies in which most evaluators co-designed the system; only E11 and E12 are external, there is no baseline, and outcomes are verbal feedback, not measured performance.","rationale":"I agree with the reader's identification: the weakest load-bearing assumption is that the case-study design yields valid evidence of effectiveness. This is not an attack on the authors' integrity; co-design is a legitimate method for deriving requirements, and the paper is transparent about participation. The problem is the strength of the conclusion drawn from limited, non-independent, qualitative data. The two external experts are a valuable gesture, but two participants, in assisted sessions, cannot support 'strong evidence,' especially when the primary evaluators are co-designers. The system itself may well be useful: the design rationale is coherent, the use of FDEB and hierarchical exploration is reasonable, and the authors describe concrete domain requirements. However, the paper would need either a quantitative evaluation or a more modest conclusion ('feasibility demonstrated in case studies') to match the evidence. I did not identify an internal technical contradiction in the construction that would invalidate the system; the reported 38,036 top-10% links from a 22,703 by 22,703 matrix is inconsistent with 10% of that matrix, but that affects reporting precision rather than the core effectiveness claim. Since the reader's conditional verdict already reflects that the system may be useful but the validation is insufficient, my finding does not change the verdict.","tokens_in":19198,"tokens_out":3909,"duration_ms":41113,"concrete_test":"Run a preregistered within-subjects study with at least ten external neuroscience or brain-computing experts who had no prior contact with DTBIA. Give each expert the same T1-T3 tasks in DTBIA and in a 2D desktop baseline (e.g., DTBVis or an equivalent desktop tool), counterbalancing order; record completion time, ROI and voxel identification accuracy against known ground truth, number of correct insights, and NASA-TLX workload; have independent coders score think-aloud sessions blind to condition. If DTBIA does not show a significant quantitative advantage over the baseline on these measures, the Section 8 'strong evidence' claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim, in the abstract and Section 8, is that utility and effectiveness are 'validated' by two case studies, with 'strong evidence supporting the practicality and effectiveness of our system.' For that claim to hold, the case-study observations must be a valid measure of how the system performs for its intended users. The current design does not deliver that. Section 3.2 describes a year-and-a-half co-design process with ten ISTB experts (E1-E10) who defined tasks T1-T3 and requirements R1-R3. Section 6 then uses E1 and E5, two of those co-designers, as primary evaluators; the other eight members of the original group 'provide additional feedback' by watching the VR sessions on screen; only E11 and E12 are external. The sessions are also described as assisted ('we assisted the experts'), so even the external participants did not work independently. No task-completion times, accuracy, error rates, or insight counts are reported; there is no comparison baseline such as DTBVis or a 2D desktop view, and no systematic coding of think-aloud transcripts. The positive statements attributed to E1 and E5 after 18 months of shaping the design are therefore not an impartial test. Section 7.2 itself concedes a 'potential for subjective biases' and says 'further research is needed to assess the reliability and objectivity of the findings.' In short, the evidence supports a feasibility demonstration, not the 'strong evidence' of effectiveness claimed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents DTBIA, an immersive visual analytics system for exploring Digital Twin Brain (DTB) data. The system supports hierarchical exploration at region, voxel, and slice levels, with two VR brain models (real-scale and large-scale) and a 3D force-directed edge bundling (FDEB) algorithm for structural DTI connectivity. The authors derive tasks T1-T3 and requirements R1-R3 from a year-and-a-half co-design process with ten domain experts, then report two case studies (Sections 6.1 and 6.2) in which four experts actively used the system and eight others observed. The abstract and Section 8 state that the system's utility and effectiveness are 'validated' and that the studies provide 'strong evidence' for these claims.","tokens_in":19461,"tokens_out":3432,"duration_ms":33224,"significance":"If the validation claim were sound, DTBIA would be a noteworthy contribution to immersive visual analytics for neuroscience, and the co-design with domain experts over an extended period is a genuine strength. The system offers concrete, potentially useful features: hierarchical spatiotemporal navigation, side-by-side simulated-versus-biological comparison, and edge bundling for dense connectivity. However, the current evidence is a feasibility demonstration, not a validation: the case studies involve only four active participants, two of whom co-designed the system, there are no quantitative outcome measures, no baseline comparison, and the authors themselves concede a 'potential for subjective biases' in Section 7.2. The central claim is therefore not supported as written.","major_comments":[{"comment":"The central claim that the system's utility and effectiveness are 'validated' by the two case studies is not supported by the reported evidence. Only four experts actively used the system (E1, E5, E11, E12), and two of them (E1, E5) were part of the co-design team described in Section 3.2; the other eight original-group experts only observed screen streaming. The study reports no task-completion times, accuracy, error rates, insight counts, or other quantitative metrics, and there is no baseline comparison, for example against DTBVis or a 2D desktop view. The sessions were also assisted by the authors ('we assisted the experts', Section 6.1), so the external participants did not work independently. Section 7.2 itself states that 'further research is needed to assess the reliability and objectivity of the findings.' The abstract and Section 8 overstate the strength of this evidence.","section":"Sections 6 and 8; Abstract"},{"comment":"The evaluation is circular in a load-bearing way. Tasks T1-T3 and requirements R1-R3 were derived from the same ten experts (E1-E10) who later evaluated the system, with E1 and E5 serving as primary case-study participants and the remaining eight providing additional feedback. The two external experts (E11, E12) were not involved in task definition, but they constitute only a small fraction of the feedback. The positive statements from E1 and E5 after an 18-month co-design process cannot serve as an impartial test of the system's effectiveness for its intended user population. This circularity is a core weakness, not a presentation issue, and it directly affects the validity of the 'strong evidence' conclusion.","section":"Sections 3.2 and 6"},{"comment":"The 3D-FDEB algorithm, which is a claimed contribution for reducing visual clutter, is not specified precisely enough to be reproduced. In Eq. (1), the summation over Q is written as ∑_{Q∈E} ‖p_i − q_i‖ / Ce(P,Q), but the points q_i and the exact form of the compatibility function Ce(P,Q) are not defined. The free parameters kP and Ce are not given any default values or tuning procedure, and the 'top 10%' threshold for selecting connections is used without justification or sensitivity analysis (it appears again in Sections 5.1.1 and 5.2.2). Without this information, the reader cannot assess whether the bundling actually improves clarity, and the contribution is not reproducible.","section":"Section 4.3, Eq. (1)"}],"minor_comments":[{"comment":"The heading 'Reion-Level Exploration' contains a typo; it should read 'Region-Level Exploration'.","section":"Section 5"},{"comment":"The text says '12 experts participated in the evaluation,' but only four experts actively interacted with the system; the other eight observed a screen stream. This phrasing overstates the amount of hands-on evaluation.","section":"Section 6, opening paragraph"},{"comment":"The phrase 'involving with brain research experts' is awkward; consider 'involving brain research experts' or 'with brain research experts.'","section":"Abstract"},{"comment":"The thresholds are stated without rationale: the 'top 10%' in Section 4.3 and the '0.8' connection-weight threshold in Section 6.1. A brief explanation of how these values were chosen would help readers understand their role in the findings.","section":"Sections 4.3 and 6.1"},{"comment":"The sentence 'Based on the feedback received from the experts, the experts initially expressed the positive evaluation of the system's interface representation' is redundant and could be streamlined.","section":"Section 6.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within the scope of TVCG, but the evaluation section does not meet the standard for a validation claim. The circularity concern is substantial: the same experts who defined the tasks evaluated the system, and the two external experts were in the minority. If the authors cannot add a more rigorous, independent evaluation (e.g., quantitative task metrics, baseline comparison, or at least a clearly separated formative and summative assessment with more external participants), they should revise the claims to describe a feasibility study with those limitations stated prominently. The system itself appears interesting, and the long-term co-design is a strength that could be leveraged in a revised framing."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things you should know. First, this is a real systems paper: DTBIA is a VR tool for exploring Digital Twin Brain data, with a hierarchical region-to-voxel-to-slice workflow, 3D edge bundling for DTI, and a side-by-side compare mode for simulated vs. biological BOLD. The engineering story is coherent and the design is grounded in an 18-month co-design process with ten domain experts. Second, the paper's central claim—'strong evidence supporting the practicality and effectiveness'—is not supported by the evidence it presents. The two case studies are qualitative walkthroughs with a handful of experts, most of whom helped design the system. Only E11 and E12 are external, and even they were assisted during the sessions. There is no baseline, no task completion times, no error rates, no independent coding of think-aloud data. The paper's own limitations section acknowledges 'potential for subjective biases' and says more work is needed, so the abstract overstates what the evaluation can show.\n\nThe genuinely new piece is the integration: nobody else has put VR navigation plus hierarchical AAL parcellation plus 3D-FDEB edge bundling in front of DTB BOLD/DTI data. That is worth a look for people working in immersive analytics for neuroscience. The citation pattern is fine—FDEB, NeuroCave, and DTBVis are all properly credited.\n\nThe soft spots are mostly in the evaluation, but there are also some internal numeric inconsistencies: the system says it shows the top 10% of connections, and in one spot calls that '38,036 links.' With 22,703 voxels, 10% of the full connectivity matrix would be millions of edges, not tens of thousands. Threshold mentions also drift between 'top 10%' and 'top 20%' in the case study. Those are easy to fix but signal that the quantitative details were not triple-checked.\n\nBottom line: the system is plausible and well described, but the effectiveness claim is not proven. A good referee should send it back for a real comparative study—DTBIA against DTBVis or a desktop view, with external, independent participants and at least some quantitative measures. That is a fixable problem, not a fatal one. I'd send it to peer review and let the reviewers demand that experiment. It deserves a serious referee even though the current conclusions should be toned down.","headline":"A well-built immersive analytics system for Digital Twin Brain data, but the 'strong evidence' claim outruns the qualitative, co-designer-heavy evaluation.","tokens_in":20065,"tokens_out":3661,"would_cite":false,"duration_ms":33401,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An immersive VR analytics system, DTBIA, lets brain researchers explore Digital Twin Brain simulations from whole regions down to individual voxels and slices.","keywords":["Digital Twin Brain","Immersive analytics","Virtual reality","BOLD signal visualization","Diffusion tensor imaging","Edge bundling","Visual analytics","Brain-inspired computing"],"falsifier":"A controlled study in which external, non-design users attempt the three defined analytic tasks with DTBIA and with the 2D DTBVis baseline, measuring time, error, and insight quality, would settle the claim: if DTBIA shows no measurable advantage, or if external users cannot reproduce the case-study findings such as the three-time-point BOLD delay, the effectiveness claim would be refuted.","tokens_in":18988,"feed_emoji":"🧠","tokens_out":6534,"duration_ms":57807,"temperature":0.7,"pith_summary":"DTBIA is an immersive visual analytics system whose central claim is that exploring Digital Twin Brain (DTB) simulation results in virtual reality, rather than on a flat screen, makes it practical for researchers to compare simulated and real brain signals and to trace structural connections. The DTB data it targets are large: BOLD time series from 22,703 voxels sampled over 166 time points, plus a 22,703 by 22,703 DTI connectivity matrix. To keep this volume comprehensible, DTBIA organizes exploration into three hierarchical levels—brain regions, voxels, and slice sections—and uses two VR brain models, one at real scale and one enlarged for flying inside the data. Two case studies with brain researchers are offered as evidence that experts can spot differences between simulated and biological BOLD signals, find regions of interest, and analyze the Default Mode Network. If the claim holds, DTBIA is a working example of immersive analytics applied to brain-inspired AI, where simulation fidelity itself becomes something a researcher can see and inspect.","feed_headline":"Brain researchers can fly from whole brain to single voxels in VR","feed_subtitle":"Experts compare simulated and real brain activity and trace pathways from regions to voxels to slices.","key_machinery":"The central mechanism is a two-scale immersive brain model combined with a three-level exploration workflow. A Real-Scale brain model, positioned near the user for walking and overview, carries functional BOLD-encoded spheres and voxel cubes; a Large-Scale brain model supports egocentric flying for navigation inside the data. The workflow follows the overview-first principle as it moves from region-level to voxel-level to slice-section-level analysis, letting users teleport from a selected region into the large-scale model. The 3D Force-Directed Edge Bundling (3D-FDEB) algorithm carries the structural-analysis load: it takes the top 10% of connections by weight, computes endpoint coordinates, and clusters spatially similar edges using spring and electrostatic-repulsion forces so DTI pathways read as coherent bundles rather than overlapping lines.","core_discovery":"On the paper's own terms, the discovery is a system-design result: a VR environment organized around a hierarchical, top-down workflow enables domain experts to perform the three analytical tasks they said existing tools did not support. DTBIA's workflow moves from region-level comparison of BOLD signals to voxel-level inspection of time-series line charts and DTI connections, and finally to sagittal, horizontal, and coronal slice views; at each stage a Real-Scale brain supports overview and a Large-Scale brain supports immersive flying navigation. A 3D force-directed edge bundling algorithm filters the DTI data to its top 10% of connections, 38,036 links, so structural pathways do not collapse into visual clutter. In the case studies, experts observed that DTB-simulated BOLD signals peaked about three time points later and were weaker than biological signals, and a neuroscience expert used the system to examine the seven-region Default Mode Network, finding high hippocampal BOLD activity consistent with its role in memory and imagination. The paper concludes that these observations and expert feedback provide strong evidence for the system's practicality and effectiveness.","pith_inferences":["The side-by-side comparison could be extended into a quantitative benchmark: automatically computing the three-time-point lag and amplitude ratio between simulated and biological BOLD signals would turn the observed delay into a measurable model-fidelity score.","The same hierarchical immersive workflow could transfer to other simulation-validation settings, such as climate or fluid-dynamics models, wherever a 3D spatial field has time series at many points and a dense network of connections.","A natural next test is a controlled study with external, task-naive participants comparing DTBIA against DTBVis on the three defined analytic tasks, to separate the contribution of the VR interface from the contribution of the hierarchical data design.","The paper's own acknowledged limitations—subjective bias from interactive exploration, cybersickness at low thresholds, and the absence of gesture controls—suggest that usability, not data capacity, is the current bottleneck."],"forward_implications":["Domain experts can visually compare simulated DTB BOLD signals against biological fMRI side by side, so discrepancies such as timing delays or weaker amplitudes become directly observable.","The region-to-voxel-to-slice workflow gives researchers a path from whole-brain overview to specific voxels and anatomical planes, supporting both functional and structural analysis in one environment.","The 3D-FDEB filtering to the top 10% of DTI connections makes large structural networks legible in VR, and thresholding lets users isolate the most significant pathways.","Because the same workflow handles human and macaque data, DTBIA supports cross-species comparisons of functional and structural brain organization.","If the case-study evidence is accepted, VR-based exploration is a viable alternative to 2D tools for DTB research, and the system can aid model validation rather than only post-hoc visualization."],"supporting_citations":[{"why":"Defines the DTB framework and supplies the simulated neural data that the system visualizes.","marker":"[7]"},{"why":"Describes the DTB simulation and assimilation platform whose BOLD and DTI outputs are used in the case studies.","marker":"[8]"},{"why":"Presents DTBVis, the 2D predecessor and baseline whose limitations motivate the immersive extension.","marker":"[10]"},{"why":"Provides the AAL atlas that defines the brain regions used as the top level of the hierarchy.","marker":"[59]"},{"why":"Supplies the overview-first interaction principle that the hierarchical workflow follows.","marker":"[67]"},{"why":"Provides the interactive 3D force-directed edge bundling adaptation used to cluster DTI edges.","marker":"[73]"},{"why":"Introduces the original force-directed edge bundling algorithm that the 3D version adapts.","marker":"[74]"},{"why":"Supports the side-by-side comparison design used for simulated versus biological BOLD signals.","marker":"[75]"}],"fun_headline_variants":["VR system lets brain experts fly from whole brain to voxels","Immersive VR reveals brain pathways with 3D edge bundling","New VR tool compares simulated and real brain signals in 3D","Fly through brain regions and voxels in an immersive VR world"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The effectiveness claim rests on the assumption that experts' observations and think-aloud feedback, gathered from a group that includes people who helped design DTBIA, fairly measure how well the system works.","fun_headline_variants_meta":{"raw":{"variants":["VR system lets brain experts fly from whole brain to voxels","Immersive VR reveals brain pathways with 3D edge bundling","New VR tool compares simulated and real brain signals in 3D","Fly through brain regions and voxels in an immersive VR world"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000303,"raw_usage":{"total_tokens":1750,"prompt_tokens":956,"completion_tokens":794,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":572,"completion_tokens_details":{"reasoning_tokens":720}},"tokens_in":572,"tokens_out":794,"duration_ms":7590,"temperature":1.0,"reasoning_tokens":720,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T12:38:22.598442+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A controlled study in which external, non-design users attempt the three defined analytic tasks with DTBIA and with the 2D DTBVis baseline, measuring time, error, and insight quality, would settle the claim: if DTBIA shows no measurable advantage, or if external users cannot reproduce the case-study findings such as the three-time-point BOLD delay, the effectiveness claim would be refuted.","supporting_citations":[{"cited_title":"Digital Twin Brain: a simulation and assimilation platform for whole human brain","cited_arxiv_id":"2308.01241","evidence_quote":"Describes the DTB simulation and assimilation platform whose BOLD and DTI outputs are used in the case studies."},{"cited_title":"Dtbvis: An interactive visual comparison system for digital twin brain and human brain","cited_arxiv_id":null,"evidence_quote":"Presents DTBVis, the 2D predecessor and baseline whose limitations motivate the immersive extension."},{"cited_title":"Tzourio-Mazoyer, B","cited_arxiv_id":null,"evidence_quote":"Provides the AAL atlas that defines the brain regions used as the top level of the hierarchy."},{"cited_title":"The craft of informa- tion visualization","cited_arxiv_id":null,"evidence_quote":"Supplies the overview-first interaction principle that the hierarchical workflow follows."},{"cited_title":"Interactive 3d force-directed edge bundling","cited_arxiv_id":null,"evidence_quote":"Provides the interactive 3D force-directed edge bundling adaptation used to cluster DTI edges."},{"cited_title":"Force-directed edge bundling for graph visualization","cited_arxiv_id":null,"evidence_quote":"Introduces the original force-directed edge bundling algorithm that the 3D version adapts."},{"cited_title":"Visual comparison for information visualization","cited_arxiv_id":null,"evidence_quote":"Supports the side-by-side comparison design used for simulated versus biological BOLD signals."}],"review_version":1}