{"id":"460160b7-ac5a-4d23-99fb-c67f3f1bf93f","arxiv_id":"2605.31360","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Introduces dashi, a Python library with unsupervised (information geometry-based) and supervised methods to quantify temporal and multi-source dataset shifts for trustworthy health AI.","lead":"dashi is a new open-source Python library for exploring, quantifying, and characterizing dataset shifts in AI models using unsupervised information geometry methods and supervised performance analysis, focused on health applications. A smart generalist might read it to see practical tools that help detect when training data differs from real-world data, which can break AI reliability in medical settings.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader correctly flagged that full-text access was required for evaluation. With the manuscript now available, the work is a software contribution whose primary assertions are descriptive (library exists, implements the stated methods, runs on the examples). No load-bearing technical assumption about correctness of the metrics or their causal link to safety outcomes is required for the paper's stated contribution to hold; therefore the UNVERDICTED verdict does not change.","tokens_in":1779,"tokens_out":285,"duration_ms":7253,"concrete_test":"Clone the dashi repository, execute the three provided case-study notebooks end-to-end on the released data, and verify that all reported Global Probabilistic Deviation and Source Probabilistic Outlyingness values are reproduced within 1 %; if any notebook fails or values diverge, the reproducibility claim is affected.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper introduces an open-source library with unsupervised (information-geometric) and supervised shift-characterization tools, demonstrated on three health-AI case studies. The central claim is that these tools enable trustworthy AI pipelines via interactive analytics and variability metrics. No internal inconsistency, missing derivation, or unsupported technical step is apparent from the described construction; the library's utility rests on its implementation and user adoption rather than a novel theoretical result that could be falsified by a single check.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript introduces dashi, an open-source Python library for the exploration, quantification, and characterization of dataset shifts. It offers a dual approach: an unsupervised method based on information geometry and non-parametric statistical manifolds (including Information Geometric Temporal plots and metrics such as Global Probabilistic Deviation and Source Probabilistic Outlyingness) plus a supervised method for quantifying model performance degradation. Both operate over user-defined temporal and domain/source batches. Utility is demonstrated via three simulated and real-world health AI case studies (gestational diabetes mellitus, COVID-19, and emergency medical dispatch). The central claim is that the library's interactive visual analytics and variability metrics support trustworthy AI development and deployment by assessing data coherence.","tokens_in":1852,"tokens_out":486,"duration_ms":17739,"significance":"If the implemented metrics and visualizations prove reliable and actionable, an open-source library providing both unsupervised geometric and supervised performance-based shift tools would address a genuine gap in accessible software for dataset shift analysis, particularly in high-stakes health AI applications where shifts can affect safety and rights. The dual unsupervised/supervised design and batch flexibility are explicit strengths that could facilitate reproducible pipelines.","major_comments":[{"comment":"Case Studies section: the three demonstrations are described at a high level but the manuscript provides no quantitative validation metrics, error analysis, or comparison against existing shift-detection baselines for the unsupervised metrics (Global Probabilistic Deviation, Source Probabilistic Outlyingness). This is load-bearing for the claim that the tools meaningfully support AI trustworthiness and safety.","section":"Case Studies"},{"comment":"Methods / Unsupervised Approach: the information-geometric and non-parametric manifold constructions are referenced but lack explicit algorithmic pseudocode, parameter settings, or sensitivity analysis, preventing independent assessment of whether the metrics are robust or merely descriptive.","section":"Methods"}],"minor_comments":[{"comment":"The manuscript would benefit from a dedicated 'Availability and Installation' subsection that includes the exact GitHub or PyPI link, license, and minimum Python/dependency versions.","section":null},{"comment":"Notation for the variability metrics is introduced in prose; adding a short mathematical definitions table or appendix would improve clarity for readers unfamiliar with information geometry.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their constructive comments, which highlight important areas for strengthening the manuscript. We agree that both major points require attention and will revise the paper to address them directly.","responses":[{"response":"We agree that the case studies, as currently presented, are primarily illustrative and do not include the requested quantitative validation, error analysis, or baseline comparisons. This limits the strength of the claims regarding support for AI trustworthiness. In the revised manuscript we will add quantitative evaluations of the unsupervised metrics (e.g., correlation with known distribution changes, comparison of Global Probabilistic Deviation and Source Probabilistic Outlyingness against baselines such as Kolmogorov-Smirnov tests and other shift detectors), along with error analysis and discussion of how these metrics relate to downstream model performance degradation.","revision_made":"yes","referee_comment":"[Case Studies] Case Studies section: the three demonstrations are described at a high level but the manuscript provides no quantitative validation metrics, error analysis, or comparison against existing shift-detection baselines for the unsupervised metrics (Global Probabilistic Deviation, Source Probabilistic Outlyingness). This is load-bearing for the claim that the tools meaningfully support AI trustworthiness and safety."},{"response":"We concur that the absence of pseudocode, explicit parameter settings, and sensitivity analysis hinders reproducibility and independent evaluation. The revised Methods section will include algorithmic pseudocode for the core unsupervised procedures (information-geometric manifold construction, Global Probabilistic Deviation, and Source Probabilistic Outlyingness), the specific parameter values used in the library implementation, and a sensitivity analysis examining robustness to key hyperparameters and data characteristics.","revision_made":"yes","referee_comment":"[Methods] Methods / Unsupervised Approach: the information-geometric and non-parametric manifold constructions are referenced but lack explicit algorithmic pseudocode, parameter settings, or sensitivity analysis, preventing independent assessment of whether the metrics are robust or merely descriptive."}],"tokens_in":1436,"tokens_out":405,"duration_ms":16609,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main thing to know is that this paper ships an open-source Python library called dashi that combines unsupervised methods based on information geometry with supervised checks for model performance drop, all aimed at temporal and multi-source shifts in health data.\n\nWhat it actually does is pull together existing ideas—non-parametric statistical manifolds, metrics like Global Probabilistic Deviation and Source Probabilistic Outlyingness, plus interactive plots—into one package that works across user-defined batches. The three case studies on gestational diabetes, COVID-19, and emergency dispatch show the intended workflow. For someone building or deploying health models, having a single tool that handles both unsupervised variability and supervised degradation in the same framework could save time.\n\nThe soft spot is that the write-up stays at the feature-list level. There are no reported benchmarks against other shift detectors, no error analysis on the new metrics, and no code-level description of how the manifolds or outlyingness scores are computed. Without those, it is hard to judge whether the library adds more than a convenient wrapper around standard statistical tests.\n\nThis is for practitioners who need practical monitoring tools rather than new theory. A reader already working on trustworthy health AI pipelines might find the batch handling and dual approach worth trying.\n\nIt is worth sending to peer review so referees can look at the actual implementation and any reproducibility materials.","headline":"dashi is a new Python library packaging info-geometry and supervised checks for dataset shifts in health AI, but the abstract gives almost no implementation or validation detail.","tokens_in":2350,"tokens_out":347,"would_cite":false,"duration_ms":6486,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"The dashi Python library quantifies dataset shifts with unsupervised information geometry metrics and supervised performance checks to support trustworthy AI.","keywords":["dataset shift","Python library","trustworthy AI","information geometry","health AI","covariate shift","temporal shift","multi-source variability"],"falsifier":"A controlled test on health data in which shifts detected and measured by dashi show no consistent correlation with actual drops in model accuracy or increases in safety risks.","tokens_in":2706,"feed_emoji":"📊","tokens_out":633,"duration_ms":17789,"temperature":0.7,"pith_summary":"The paper introduces dashi, an open-source Python library for exploring, quantifying, and characterizing dataset shifts where train and test data distributions differ. It offers an unsupervised approach based on information geometry and non-parametric statistical manifolds, plus a supervised approach that measures model performance degradation, both operating on user-defined temporal and multi-source batches. The library is demonstrated on health AI case studies involving gestational diabetes, COVID-19, and emergency medical dispatch. A sympathetic reader would care because uncontrolled shifts can degrade AI models and compromise safety, especially in healthcare where patient rights are at stake. By supplying visual analytics and variability metrics, the work aims to enable more coherent data assessment throughout the AI lifecycle.","feed_headline":"Library quantifies dataset shifts to support safe AI in health","feed_subtitle":"Unsupervised geometry metrics and performance checks track changes over time and across sources in health applications.","key_machinery":"The dual unsupervised-supervised framework that applies information geometry and non-parametric statistical manifolds for variability metrics alongside performance degradation analysis on temporal and multi-source batches.","core_discovery":"dashi is a Python library providing a dual approach to dataset shift analysis: an unsupervised method that uses information geometry and non-parametric statistical manifolds to characterize data variability through metrics such as Global Probabilistic Deviation and Source Probabilistic Outlyingness, together with Information Geometric Temporal plots, and a supervised method that quantifies model performance degradation, with both methods applicable across user-defined temporal and domain or source batches.","pith_inferences":["The same metrics could support ongoing monitoring once a model is deployed rather than only during development.","Integration with existing training workflows might allow automatic retraining triggers when certain shift thresholds are crossed.","The library's structure could be tested on non-health domains such as financial or sensor data where distribution changes are also common."],"forward_implications":["Shifts can be quantified and visualized across both temporal and multi-source batches using the supplied metrics.","Model performance changes due to shifts can be tracked through the supervised component.","Interactive analytics enable assessment of data coherence to guide AI pipeline decisions.","The tools apply to training and operational stages to help maintain reliability in health AI systems."],"fun_headline_variants":["dashi Python library quantifies health AI dataset shifts","dashi uses geometry metrics to analyze data shifts","Library dashi checks model degradation from shifts","dashi dual approach for shift characterization in health","Python dashi supports AI with shift variability metrics"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That the unsupervised metrics derived from information geometry and non-parametric manifolds deliver actionable characterization of shifts that meaningfully supports AI trustworthiness and safety.","fun_headline_variants_meta":{"raw":{"variants":["dashi Python library quantifies health AI dataset shifts","dashi uses geometry metrics to analyze data shifts","Library dashi checks model degradation from shifts","dashi dual approach for shift characterization in health","Python dashi supports AI with shift variability metrics"]},"model":"grok-4.3","cost_usd":0.004729,"raw_usage":{"total_tokens":2359,"prompt_tokens":720,"num_sources_used":0,"completion_tokens":69,"cost_in_usd_ticks":47287000,"prompt_tokens_details":{"text_tokens":720,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1570,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":720,"tokens_out":69,"duration_ms":9474,"temperature":1.0,"reasoning_tokens":1570,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T22:56:46.941805+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A controlled test on health data in which shifts detected and measured by dashi show no consistent correlation with actual drops in model accuracy or increases in safety risks.","supporting_citations":[],"review_version":1}