REVIEW 5 major objections 6 minor 1 cited by
A High-Quality Thermoelectric Material Database with Self-Consistent ZT Filtering
T0 review · 5 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Comparing a paper's reported ZT against the ZT recomputed from its own Seebeck, resistivity, and conductivity curves yields a curated database of 272 internally consistent thermoelectric samples.
desk verdict Sc-ZT filtering is a genuinely useful curation protocol, but the dataset's claim to 'high quality' rests on internal consistency, not on verified absolute accuracy. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the self-consistent ZT (Sc-ZT) filter, a six-threshold screening protocol whose core quantity is the ZT error $\delta(\mathrm{ZT}) = \mathrm{ZT}_{\mathrm{fig}} - \mathrm{ZT}_{\mathrm{TEP}}$, evaluated on curves collocated to a common 2 K temperature grid. The filter combines average-ZT, peak-ZT, maximum-error, root-mean-square-error, and two normalised maximum-error statistics, with default thresholds $(0.1, 0.1, 0.1, 0.1, 0.2, 0.2)$ that are user-tunable. This construction turns the vague notion of 'suspicious data' into a quantitative, repeatable test and simultaneously produces a taxonomy of the error types that make thermoelectric figures of merit unreliable.
What would settle it
Take the samples that failed the Sc-ZT filter, obtain the original authors' raw tabulated measurements, and recompute ZT from those tables. If a large fraction of the rejected samples reproduce the reported ZT when computed from the original tables, then the filter is primarily removing accurate entries whose plotted curves are at fault, and $\delta(\mathrm{ZT})$ is not a reliable error metric.
Extended reading notes
Core claim
The central claim is that the ZT error, defined as $\delta(\mathrm{ZT}) = \mathrm{ZT}_{\mathrm{fig}} - \mathrm{ZT}_{\mathrm{TEP}}$ with $\mathrm{ZT}_{\mathrm{TEP}} = \alpha^2 \rho^{-1} \kappa^{-1} T$, is a reliable diagnostic for whether a thermoelectric publication's performance claims are internally consistent. Applying six threshold filters built on this error, the authors reduce 355 candidate samples to 272, and the distribution of residual errors shifts from strongly non-normal (Q-Q $R^2 = 0.6864$) to essentially normal ($R^2 = 0.9324$). The same protocol applied to a large open digitised dataset removes thousands of self-inconsistent entries, including unit-label mistakes that can inflate ZT by a factor of $10^6$. The paper identifies six concrete error mechanisms that produce this inconsistency: resolution error, publication bias, ZT overestimation from curve fitting, extrapolation beyond the measured temperature range, interpolation across phase transitions, and digitisation noise.
Load-bearing premise
The filter treats the ZT recomputed from the digitised Seebeck, resistivity, and thermal-conductivity curves as the trustworthy reference, so if those digitised curves are themselves systematically wrong, or if the reported ZT was computed from a different measurement set than the plotted curves, the filter's verdict will be biased.
Editorial extensions
If this is right
- Any publication reporting a ZT curve alongside its three constituent property curves can be checked automatically for internal consistency, so the protocol serves as a general data-quality screen rather than a one-off curation exercise.
- The curated 272-sample database is internally consistent enough to act as a benchmark for machine-learning predictions of thermoelectric property curves, with residual scatter dominated by digitisation noise rather than systematic publication bias.
- When the same filter is applied to a large community dataset, it reduces the usable sample count from roughly 15,500 to about 10,800, showing that many self-inconsistent entries are detectable and removable at scale.
- Users can trade dataset size against strictness: the default thresholds keep 272 samples, stricter thresholds keep 187, and the strictest tested thresholds keep 71, with the Q-Q $R^2$ of the error distribution rising to about 0.985 when only the most self-consistent data remain.
Reading between the lines
- If adopted as a reporting standard, the same comparison could be run by journals before publication: authors would submit the collocated curves and the $\delta(\mathrm{ZT})$ plot as supplementary material, preventing inflated figures of merit from entering the literature in the first place.
- The discrepancy logic generalises to any derived material property that is a deterministic function of independently plotted curves, such as power factor, Lorenz number, or lattice thermal conductivity, so the protocol could become a template for other functional-materials databases.
- The paper's validation assumes the digitised TEP curves are closer to ground truth than the reported ZT values; a targeted test against original experimental tables for a handful of rejected samples would settle whether the filter is removing publication errors or digitisation artefacts.
- Because the same filter caught order-of-magnitude errors from unit confusion, it could also serve as an automated unit-sanity check in crowdsourced or legacy datasets where label mistakes are common.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript introduces teMatDb, a database of temperature-dependent thermoelectric properties (Seebeck coefficient, electrical resistivity, thermal conductivity, and ZT) digitized from published figures. The central methodological contribution is a self-consistent ZT (Sc-ZT) filter that compares the reported ZT (ZTfig) with the ZT recalculated from the digitized TEPs (ZTTEP), defining δ(ZT) = ZTfig − ZTTEP. Six filters (Avg ZT, Peak ZT, Max error, RMS error, and two normalized versions) with default thresholds (0.1, 0.1, 0.1, 0.1, 0.2, 0.2) are applied to a parent database teMatDb v1.1.6, reducing 355 samples to teMatDb272 (272 samples from 262 publications). The same protocol is applied to Starrydata2 to generate starryz10840. The paper presents Q-Q plots and ZT-ZT comparisons as evidence of improved consistency, and provides code and data on GitHub/Zenodo. The authors claim the filtered dataset is a robust, high-quality experimental benchmark for machine learning and materials design.
Significance. If the dataset were validated for absolute accuracy, teMatDb272 would be a valuable benchmark: it covers a broad compositional space, includes self-consistency checks that catch unit errors and extrapolation artifacts, and ships with open code and data, which is a strength. The Sc-ZT filtering framework is a useful, physically motivated quality-control tool that could be applied to other experimental databases. However, the central claim of 'high-quality' is currently overreaching: the validation in Section 4 demonstrates only that the filter enforces internal consistency between digitized TEPs and reported ZT, not that the values are accurate. Because the paper emphasizes the dataset as a reliable experimental benchmark, the lack of external validation is a load-bearing issue that should be addressed before publication.
major comments (5)
- [Section 4, Table 4] The validation of the Sc-ZT filter is self-referential: Table 4 reports that the Q-Q R2 of δ(ZT) improves from 0.6864 to 0.9324 after filtering, but this improvement is expected because the filter removes samples with large δ(ZT), truncating the distribution that the R2 is computed on. No independent validation is provided: the digitized TEPs are never compared against original data tables, authoritative measurements, or remeasured values for any sample in teMatDb272. Consequently, the claim that teMatDb272 constitutes a 'high-quality' experimental benchmark is not supported; the filter demonstrates internal consistency, not absolute accuracy. I recommend either adding an external validation of a subset of samples against the original publications' tables or tempering the claims to 'internally consistent' throughout the abstract and main text.
- [Table 1, Section 2] The default Sc-ZT filter thresholds (0.1, 0.1, 0.1, 0.1, 0.2, 0.2) in Table 1 are presented without derivation or sensitivity analysis. The digitization noise analysis in Tables S1-S2 shows that relative errors in reciprocals can reach 8% and the mean relative error in reciprocal values can be as large as 6.26%, yet the thresholds are stated in absolute ZT units (0.1) and in normalized units (0.2). For a low-ZT sample (e.g., ZT ~ 0.3), an absolute error of 0.1 is a 33% error, which may be acceptable or not depending on the application. The authors should justify the thresholds using the digitization uncertainty or provide a systematic sensitivity analysis showing how the retained dataset changes as thresholds vary, and discuss how threshold choices affect downstream machine-learning use.
- [Section 2, Tables S1-S2] The paper quantifies digitization noise on a synthetic test figure, but does not propagate this uncertainty into the individual δ(ZT) values used for filtering. Each sample is accepted or rejected based on point estimates without per-sample error bars. Given that reciprocal amplification can cause up to 8% error (Table S2), some borderline samples in teMatDb272 may have been erroneously retained or rejected. The authors should provide a per-sample uncertainty estimate for ZTTEP and δ(ZT) (e.g., via bootstrap or repeated digitization) and report how many samples lie near the filtering boundaries.
- [Section 2, Figure 3] The Q-Q plot R2 against a normal distribution is used as a data-quality metric, but the paper does not justify why δ(ZT) should be normally distributed, nor does it test this assumption (e.g., with Shapiro-Wilk or Anderson-Darling tests). The increase in R2 after filtering (from 0.6864 to 0.9324) is partly a mechanical consequence of removing extreme values, which makes the remaining distribution more compact and closer to normal. The authors should either provide a physical argument for normality of δ(ZT) or replace the Q-Q R2 with a less assumption-dependent metric, such as the mean absolute error or the proportion of samples within a tolerance band.
- [Table 1, Section 2] The Avg ZT and Peak ZT filters in Table 1 compare averages and peaks computed over potentially different temperature ranges: Avg(ZTfig) is integrated over ΔTfig while Avg(ZTTEP) is integrated over ΔTTEP, and the manuscript states that these are computed over the ZT or TEP temperature ranges, respectively. If ZTfig extends beyond the TEP measurement range (as in the extrapolation error discussed for sample_id = 113), the average over a wider interval will differ even when the curves coincide over the overlap, leading to spurious filter outcomes. The authors should clarify whether Avg and Peak filters are evaluated over the overlapping temperature range, as done for the Max and RMS filters, and if not, justify the use of different integration intervals.
minor comments (6)
- [Abstract] The abstract uses 'tMatDb272' while the rest of the text uses 'teMatDb272'; please make the name consistent throughout.
- [Table 3] The row 'Entries for κ in rawTEPs' appears to be incomplete; the text states 3,422 entries, but the value is missing from the table.
- [Section 2, Sc-ZT filtering protocol] The phrase 'temperature rannges' should be corrected to 'temperature ranges'.
- [Section 4, Limitations] The limitations paragraph would benefit from an explicit statement that Sc-ZT filtering checks self-consistency and cannot detect systematic measurement errors common to all three TEPs (for example, a biased thermal conductivity measurement that leaves ZTTEP and ZTfig both shifted).
- [Figure 5] The logarithmic-scale plot in (a) and the linear-scale zoom in (b) would be easier to interpret if the axes were labeled with units (ZT is dimensionless) and if the three datasets were identified in a legend consistent with the text.
- [Table 2] The field 'Composition_detailed' is described as 'Full stoichiometric composition'; consider providing an example of its format in the Usage Notes to help users parse the metadata.
Circularity Check
Sc-ZT filtering is a self-consistency screen, not an independent accuracy check; the reported quality metric is the filter output by construction.
-
self definitional
[Section 2, Sc-ZT filtering protocol (page 13) and Table 4]
"After filtering, the δ(ZT) distribution approximates a normal shape with a high R2 value of 0.9324. This demonstrates that filtering effectively removes anomalies such as bias, extrapolation, and interpolation error, leaving primarily digitisation noise."
The filtering criteria in Table 1 are direct caps on Avg(δ), Peak(δ), Max(δ), Rms(δ), and normalized Max(δ). Applying these caps to δ(ZT) necessarily truncates the very distribution whose Q-Q R² is then reported as 'data quality' in Table 4. The improvement from R²=0.6864 to 0.9324 is an expected consequence of removing high-deviation points, not an independent demonstration that the surviving α, ρ, κ values are accurate. Any external error common to both ZTfig and ZTTEP (e.g., a biased κ measurement or the same wrong curve digitized) would pass the filter unchanged.
-
self definitional
[Section 1, Background and Summary (page 4-5)]
"By applying this protocol, we systematically identified and excluded erroneous data entries, resulting in a high-quality dataset named teMatDb272, which exhibits improved internal consistency between TEPs and ZT values."
The dataset is called 'high-quality' because it passes the Sc-ZT protocol, and the Sc-ZT protocol is presented as effective because it produces a high-quality dataset. The only property optimized is internal consistency between digitized ZTfig and ZTTEP computed from the same source figures; the protocol does not compare digitized TEPs against original data tables, raw measurements, or independent remeasurement. Thus the claim that teMatDb272 is a 'robust dataset for data-driven and machine-learning-based materials design' rests on a definitional equivalence between 'passes the filter' and 'high-quality,' rather than on external validation of the experimental values.
full rationale
The construction uses the physically exact, parameter-free relation ZT = α²ρ⁻¹κ⁻¹T, so the comparison itself is not fitted. The external application to Starrydata2 is genuine independent evidence that the filter catches gross unit/label errors, and there is no load-bearing self-citation chain. However, for teMatDb272 itself the 'high-quality' claim is self-referential: the data are selected because their self-consistency error is small, and then the same error metric is reported as validation. This does not confirm the absolute accuracy of the digitized TEPs, which is the paper's implied benchmark value. The circularity is partial rather than total, so a score of 4 is appropriate.
Assumptions & free parameters
free parameters (7)
- Avg ZT filter threshold =
0.1
- Peak ZT filter threshold =
0.1
- Max ZT error threshold =
0.1
- RMS ZT error threshold =
0.1
- Normalized Max/Avg ZT error threshold =
0.2
- Normalized Max/Peak ZT error threshold =
0.2
- Temperature collocation interval =
2 K
assumptions (5)
- domain assumption ZT is defined as alpha^2 T / (rho kappa).
- domain assumption Digitized TEPs are accurate enough to serve as the reference for judging ZTfig.
- ad hoc to paper The distribution of delta(ZT) in a high-quality dataset should be approximately normal; a high Q-Q R2 indicates quality.
- domain assumption Piecewise linear interpolation at 2 K intervals adequately represents TEP curves between measured points.
- ad hoc to paper The default filter thresholds (0.1, 0.1, 0.1, 0.1, 0.2, 0.2) define an acceptable level of self-consistency.
Cite this review
Pith. "Pith review of A High-Quality Thermoelectric Material Database with Self-Consistent ZT Filtering." pith.science (2026). https://pith.science/paper/4JKGKT4E
@misc{pith2026250519150,
author = {Pith},
title = {Pith review of: A High-Quality Thermoelectric Material Database with Self-Consistent ZT Filtering},
year = {2026},
howpublished = {\url{https://pith.science/paper/4JKGKT4E}},
note = {Machine review of arXiv:2505.19150}
}
read the original abstract
This study presents a curated thermoelectric material database, teMatDb, constructed by digitizing literature-reported data. It includes temperature-dependent thermoelectric properties (TEPs), Seebeck coefficient, electrical resistivity, thermal conductivity, and figure of merit (ZT), along with metadata on materials and their corresponding publications. A self-consistent ZT (Sc-ZT) filter set was developed to measure ZT errors by comparing reported ZT's from figures with ZT's recalculated from digitized TEPs. Using this Sc-ZT protocol, we generated tMatDb272, comprising 14,717 temperature-property pairs from 272 high-quality TEP sets across 262 publications. The method identifies various types of ZT errors, such as resolution error, publication bias, ZT overestimation, interpolation and extrapolation error, and digitization noise, and excludes inconsistent samples from the dataset. teMatDb272 and the Sc-ZT filtering framework offer a robust dataset for data-driven and machine-learning-based materials design, device modeling, and future thermoelectric research.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
Physics-Informed Neural Operators for Generalizable and Label-Free Inference of Temperature-Dependent Thermoelectric Properties
A physics-informed neural operator trained on 20 simulated thermoelectric materials infers thermal conductivity and Seebeck coefficient for 60 unseen materials from six sparse measurements, with test R-squared above 0.97.
Reference graph
Works this paper leans on
-
[1]
Agrawal, A. & Choudhary, A. Perspective: Materials informatics and big data: Realization of the “fourth paradigm” of science in materials science. APL Mater. 4, 053208 (2016)
work page 2016
-
[2]
Jain, A. et al. Commentary: The Materials Project: A materials genome approach to accelerating materials innovation. APL Mater. 1, 011002 (2013)
work page 2013
- [3]
-
[4]
Artrith, N. et al. Best practices in machine learning for chemistry. Nat. Chem. 13, 505–508 (2021)
work page 2021
-
[5]
Katsura, Y ., Akiyama, M., Morito, H., Fujioka, M. & Sugahara, T. Systematic searches for new inorganic materials assisted by materials informatics. Sci. Technol. Adv. Mater. 26, 2428154 (2024)
work page 2024
-
[6]
Katsura, Y . et al. Starrydata: from published plots to shared materials data. Sci. Technol. Adv. Mater. Methods 2506976 (2025) doi:10.1080/27660400.2025.2506976
arXiv 2025
-
[7]
Barroso-Luque, L. et al. Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models. Preprint at https://doi.org/10.48550/arXiv.2410.12771 (2024)
-
[8]
& Ong, S
Chen, C. & Ong, S. P. A universal graph deep learning interatomic potential for the periodic table. Nat. Comput. Sci. 2, 718–728 (2022)
2022
Show all 29 references
-
[9]
& Han, S
Park, Y ., Kim, J., Hwang, S. & Han, S. Scalable Parallel Algorithm for Graph Neural Network Interatomic Potentials in Molecular Dynamics Simulations. J. Chem. Theory Comput. 20, 4857–4868 (2024)
2024
- [10]
-
[11]
Jumper, J. et al. Highly accurate protein structure prediction with AlphaFold. Nature 596, 583–589 (2021)
2021
-
[12]
& Rehme, S
Zagorac, D., Müller, H., Ruehl, S., Zagorac, J. & Rehme, S. Recent developments in the Inorganic Crystal Structure Database: theoretical crystal structure data and related features. J. Appl. Crystallogr. 52, 918–925 (2019)
2019
-
[13]
Ryu, B. et al. Best thermoelectric efficiency of ever-explored materials. iScience 26, 106494 (2023)
2023
-
[14]
T., Davies, D
Butler, K. T., Davies, D. W., Cartwright, H., Isayev, O. & Walsh, A. Machine learning for molecular and materials science. Nature 559, 547–555 (2018)
2018
-
[16]
J., Pereyra, A
Snyder, G. J., Pereyra, A. & Gurunathan, R. Effective Mass from Seebeck Coefficient. Adv. Funct. Mater. 32, 2112772 (2022)
2022
-
[17]
Goldsmid, H. J. Introduction to Thermoelectricity. vol. 121 (Springer, Berlin, Heidelberg, 2016)
2016
-
[18]
M., Kim, H
Gibbs, Z. M., Kim, H. -S., Wang, H. & Snyder, G. J. Band gap estimation from temperature dependent Seebeck measurement—Deviations from the 2e|S|maxTmax relation. Appl. Phys. Lett. 106, 022112 (2015)
2015
-
[19]
-S., Gibbs, Z
Kim, H. -S., Gibbs, Z. M., Tang, Y ., Wang, H. & Snyder, G. J. Characterization of Lorenz number with Seebeck coefficient measurement. APL Mater. 3, 041506 (2015)
2015
-
[20]
Katsura, Y . et al. Data-driven analysis of electron relaxation times in PbTe-type thermoelectric materials. Sci. Technol. Adv. Mater. 20, 511–520 (2019)
2019
-
[21]
WebPlotDigitizer
Rohatgi, A. WebPlotDigitizer
-
[23]
Lee, J. K. et al. Control of thermoelectric properties through the addition of Ag in the Bi0.5Sb1.5Te3Alloy. Electron. Mater. Lett. 6, 201–207 (2010)
2010
-
[24]
Ioffe, A. F. Semiconductor Thermoelements and Thermoelectric Cooling. (Infosearch, 1957)
1957
-
[25]
& Park, S
Ryu, B., Chung, J. & Park, S. Thermoelectric degrees of freedom determining thermoelectric efficiency. iScience 24, 102934 (2021)
2021
-
[26]
& Seo, H
Chung, J., Ryu, B. & Seo, H. Unique temperature distribution and explicit efficiency formula for one - dimensional thermoelectric generators under constant Seebeck coefficients. Nonlinear Anal. Real World Appl. 68, 103649 (2022)
2022
-
[27]
& Park, S
Ryu, B., Chung, J. & Park, S. Thermoelectric algebra made simple for thermoelectric generator module performance prediction under constant Seebeck coefficient approximation. J. Appl. Phys. 137, 055001 (2025)
2025
-
[28]
Zhao, L.-D. et al. Ultrahigh power factor and thermoelectric performance in hole-doped single-crystal SnSe. Science 351, 141–144 (2016)
2016
-
[29]
Lee, J. K. et al. Improvement of thermoelectric properties through controlling the carrier concentration of AgPb18SbTe20 alloys by Sb addition. Electron. Mater. Lett. 8, 659–663 (2012)
2012
-
[31]
M., Takagiwa, Y
Wang, H., Gibbs, Z. M., Takagiwa, Y . & Snyder, G. J. Tuning bands of PbSe for better thermoelectric efficiency. Energy Environ. Sci. 7, 804–811 (2014)
2014
-
[32]
& Park, S
Chung, J., Ryu, B. & Park, S. Dimension reduction of thermoelectric properties using barycentric polynomial interpolation at Chebyshev nodes. Sci. Rep. 10, 13456 (2020). arXiv:2505.19150 Page 25 / 45 Table 1. Sc -ZT filters and description. Summary of the Sc -ZT filtering sche...
2020 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.