REVIEW 5 major objections 5 minor 44 references
LLM-based Evaluation Policy Extraction for Ecological Modeling
T0 review · 5 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper claims that an LLM, guided by a tunable base metric, can extract expert-level evaluation policies for ecological time series from pairwise annotations.
desk verdict Novel and promising approach to interpretable time-series evaluation, but the headline claim about capturing human expert criteria rests on an underpowered five-sample test and a likely bug in the base metric. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central mechanism is the paired weight-optimization and policy-extraction loop. The base metric scores a prediction $P_i$ against observations $Y_i$ as $S(Y_i,P_i)$, a period-weighted sum of peak-alignment similarity and derivative (slope/curvature) differences, segmented into before/in/after the dominant rise-fall period. At step $d$ the LLM updates weights via $w_{d+1}=\mathrm{LLM}(w_d,H_d,P_i^{a,b},C)$ from the current weights, optimization history $H_d$, one pairwise expert annotation, and constraints (bounds, smoothness, normalization). The history is also transcribed into a structured policy $\pi_d=\{M_d,F_d,S_d,R_d\}$—metric names, formulas, scores summing to $K$, and a decision rule—and a candidate policy is retained only if it beats the incumbent on a validation set in at least 70% of repeated runs.
What would settle it
Run the same APEF training on each expert's annotations separately and on a bootstrapped majority vote, then recompute the test-set Spearman correlations; if the spread of these values overlaps the range of the baseline metrics (e.g., near-zero for R2 in several settings), the central claim that APEF captures expert criteria rather than the specific majority vote would be falsified.
Extended reading notes
Core claim
The paper's central claim is that APEF captures complex assessment criteria provided by human expert annotators, evidenced by high correlation with target scores. Concretely, APEF learns weights plus natural-language policies from pairwise preferences and reproduces three kinds of target rankings: rankings generated by preset base-metric weights on synthetic data, rankings from majority-voted expert annotations for GPP and CO2 flux, and rankings from a land-model benchmarking score system. In the expert-annotation experiment, APEF's Spearman correlations with expert rankings reach 0.785 for CO2 flux, 0.752 for GPP, and 0.417 for the combined two-variable task, above the reported baselines; in the benchmark experiment the correlations remain comparable even though the target score is built from standard metrics. The extracted policies include human-readable metrics such as a Peak Period Consistency Score and a proportion-of-large-error-time-steps rule.
Load-bearing premise
The argument assumes that the majority-voted expert pairwise annotations used to train and validate APEF are a trustworthy target; with Fleiss kappa values of 0.69, 0.58, and 0.63 and test rankings built from only five model predictions, annotation noise alone could change the reported correlations.
Editorial extensions
If this is right
- If APEF is right, expert visual inspection of ecological model outputs can be scaled up: pairwise annotations from a handful of experts are converted into reusable scoring rules that run automatically on new model runs.
- The extracted policies expose which temporal features matter (peak timing, amplitude, slope, curvature, period consistency), making evaluation criteria inspectable by other scientists.
- Evaluation can be adapted per community: changing the training annotations shifts the learned policy, so agronomists and climatologists can maintain distinct, standardized criteria from the same base metric.
- Multi-variable assessments can include inter-series consistency (e.g., GPP and CO2 flux together), so models are judged not only on each output but on how well correlated variables move together.
- The framework extends beyond ecology to any scientific domain where model outputs are time series and experts can give pairwise preferences.
Reading between the lines
- A cleaner test of the LLM's contribution would be to compare APEF against a non-LLM optimizer (e.g., grid search or Bayesian optimization) over the same base metric; if those match APEF's correlations, the natural-language policy layer is adding interpretability, not ranking power.
- A useful stress test is to train separate policies on each expert's judgments and measure agreement between policies; if policies diverge, the majority-vote policy should be interpreted as one consensus view, not a hidden ground truth.
- The policy-validation rule (keep a policy only if it beats the previous one on a validation set in 70% of runs) raises the question of how much of the final policy's performance is inherited from the base metric's weight optimization; ablating the policy layer would separate those contributions.
- If the method generalizes, LLM-based policy extraction could be applied to other environmental time series such as hydrological forecasts or remote-sensing products, where pairwise expert preference data are easier to collect than global scores.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes APEF, a framework that combines a modular base metric for ecological time series (peak alignment, derivative, amplitude) with LLM-based weight optimization and natural-language policy extraction to produce interpretable, adaptable evaluation policies. The authors validate APEF on three settings: synthetic rankings generated by the base metric with preset weights, rankings from three human experts, and the ILAMB benchmarking scores. They report Spearman correlations between APEF and target rankings and compare against traditional metrics, TILDE-Q, and PRP-Rank. The central claim is that APEF effectively captures complex assessment criteria, including human expert preferences.
Significance. The problem is important: ecological model evaluation often requires expert visual inspection of time series, and standard numeric metrics miss domain-specific temporal patterns. If the approach worked as claimed, learnable, interpretable evaluation policies would be a meaningful advance, with clear practical value for benchmarking in carbon-cycle and agroecosystem modeling. The interpretable policy output (e.g., the extracted PPCS) and the multivariate extension are strengths, and the paper ships an open-source implementation. However, the current validation is not yet convincing enough to establish the central claim: the strongest independent evidence is underpowered, and the synthetic experiment is self-referential.
major comments (5)
- [Section 5.1.2, Table 2 (left)] The expert-alignment experiment is underpowered and the reported statistics are not internally consistent. The test set contains only 5 model predictions, so with n=5 Spearman's rho can only take the discrete values 1 - Σd²/20, which do not include 0.785, 0.752, or 0.417. The paper provides no confidence intervals, permutation tests, or p-values, so the conclusion that APEF 'effectively captures' expert criteria is not statistically supported. The GPP+CO2 value of 0.417 is not 'high' by any standard, and the majority-voted labels come from only three experts with Fleiss kappa 0.58–0.69, i.e., moderate agreement.
- [Section 5.1.1, Table 1] The synthetic experiment is self-referential: the target ranking is generated by applying the base metric (Eq. 5) with preset weights, and APEF optimizes the same base metric's weights to match that ranking. This is a recovery check, not evidence that APEF captures assessment criteria beyond the base metric. Indeed, Section 5.1.1 states the dataset is for evaluating whether the framework can 'recover the assessment using the base metric.' The conclusion should not cite this experiment as evidence for capturing human expert criteria.
- [Section 3.1, Eqs. (3)–(5)] The base metric S(Y,P) in Eq. 5 adds a similarity term (S_Peak, higher is better) to a distance term (S_Deriv, higher is worse), and Eq. 6 then treats higher S as better. As written, a larger derivative distance would increase S and make a model appear better, which contradicts the intended preference direction. The subsequent conversion 'S(Y,P)=1/(1+S(Y,P))' reuses the same symbol S without clarifying whether S is a distance or a similarity. Since all weight optimization and ranking depend on this score, the inconsistency needs to be resolved and the experiments re-run with a well-defined score.
- [Section 5.2.3, Table 2 (right)] On the ILAMB target, APEF achieves Spearman correlations of 0.750 (CO2) and 0.821 (GPP), which are below the TILDE-Q baseline of 0.857 for both. The text describes this as 'comparable,' but it is strictly worse. Since ILAMB is the only independent, non-circular target in the paper, this result does not support the claim of adaptability to different evaluation settings.
- [Section 4.3, Eq. (13)] The policy-level validation criterion is ambiguous: Eq. (13) says the new policy is accepted 'if ρ_val(π_{d+1}) > ρ_val(π_d) for θ_LLM runs,' but the text states 'the threshold θ is set to be 70%.' If θ is a fraction of repeated runs, the equation should say so explicitly; if it is a number of runs, 70% is meaningless. This matters because the policy acceptance rule controls which extracted policies are evaluated and reported.
minor comments (5)
- [Throughout] There are several typos: 'Amplitutde' in Table 1's header, 'CIMP6' in Section 5.2.3, and 'experinment' in the paragraph before Section 5.2.1.
- [Section 5.2.1] The text refers to 'Table 4' but the correlation table that appears is labeled 'Table 1'; the numbering should be corrected.
- [Section 4.2, Eqs. (8)–(9)] The policy component is written as 'M_D' in Eq. (8) but 'M_d' in Eq. (9); the subscript notation should be consistent.
- [Section 4.2, Decision Rule] In item (ii), the text says 'multiple series (e.g., P_{i,(a)} and P_{i,(a)})' where the second index should presumably be (b); as written, the example is self-comparison.
- [Section 5.2.3] The assertion that 'most traditional metrics tend to produce high correlation performance (>0.85)' on the ILAMB target is not backed by any reported numbers in the paper, so the claim cannot be verified.
Circularity Check
Synthetic benchmark is a self-referential recovery check; expert and ILAMB targets are independent, so circularity is partial.
-
self definitional
[Section 5.1.1 (Synthetic Dataset) and Section 4.4 (Eq. 14)]
"We use each preset weight combination to compute the base metric and then create the target ranking. We use Eq. 14 to compute the score for two variables GPP+CO2 and create the target ranking."
Eq. 14 is also APEF's internal combined score for multivariate series: Section 4.4 says 'the combined score is computed as the average of the individual base metric scores, scaled by their correlation difference, as follows: Eq. 14.' The synthetic GPP+CO2 target is generated by the very scoring formula APEF optimizes, so the high Spearman correlations in Table 1 measure recovery or fitting of the target-generating formula rather than acquisition of independently defined expert preferences.
-
fitted input called prediction
[Section 5.1.1 / 5.2.1, Eqs. 5-6]
"For this experiment, we use each of the metrics to rank the samples in testing set and calculate the Spearman's correlation between the generated ranking and the ground truth ranking. Here the ground truth ranking is generated based on the base metric score (Section 3.1) with the preset weights."
For the single-variable synthetic settings, the ground truth is Eq. 5 with preset weights, and APEF's weight optimizer adjusts the same Eq. 5 weights so that the estimated base metric scores align with those preset-weight rankings (Eq. 6). Because the test set is a holdout from that same generated ranking, the correlations in Figure 4 and Table 1 are a consistency check on fitting the same parametric family, not independent validation of evaluation-policy extraction.
full rationale
Most of the paper's validation chain is self-contained and not circular: the human-expert experiment uses an external target (majority-voted expert pairwise comparisons) and the ILAMB experiment uses an external formulaic score, so neither reduces to APEF's own inputs. The circular component is confined to the synthetic benchmark: target rankings are produced by the same Eq. 5 / Eq. 14 base-metric family that APEF's weight optimizer is trained to fit (Eq. 6 and Section 4.4), so high correlations in Table 1 and Figure 4 show in-family recovery rather than independent preference capture. The central claim about capturing expert criteria therefore rests primarily on the human-expert experiment, which is independent but statistically weak: only 5 test predictions and Fleiss kappa 0.58-0.69, with no significance tests or confidence intervals. That weakness is an evidence-quality concern, not circularity. No load-bearing self-citation or imported uniqueness theorem is present, so the overall score reflects partial, experiment-specific circularity rather than a derivation that is equivalent to its inputs.
Assumptions & free parameters
free parameters (3)
- Rise/fall segmentation thresholds theta_rise, theta_fall, theta_period =
0.01, -0.01, 5
- Base metric component weights (w_peak, w_der, w_amp, tolerance) =
Not reported for real experiments; preset (0.8,0.1,0.1), (0.1,0.8,0.1), (0.1,0.1,0.8) in synthetic experiments
- Policy validation threshold theta_LLM and score normalization K =
70% and 10
assumptions (4)
- domain assumption Each time series can be represented by a single rise-then-fall pattern, or segmented into such patterns.
- domain assumption Pairwise expert annotations are sufficient to reconstruct an overall ranking of all samples.
- domain assumption The LLM (o3-mini) produces valid weight updates and mathematically grounded policy formulas, and the 70% validation gate removes unstable outputs.
- domain assumption Majority-voted expert annotations, with Fleiss kappa 0.58 to 0.69, are reliable ground truth for training and testing.
Cite this review
Pith. "Pith review of LLM-based Evaluation Policy Extraction for Ecological Modeling." pith.science (2026). https://pith.science/paper/JRUPMVKW
@misc{pith2026250513794,
author = {Pith},
title = {Pith review of: LLM-based Evaluation Policy Extraction for Ecological Modeling},
year = {2026},
howpublished = {\url{https://pith.science/paper/JRUPMVKW}},
note = {Machine review of arXiv:2505.13794}
}
read the original abstract
Evaluating ecological time series is critical for benchmarking model performance in many important applications, including predicting greenhouse gas fluxes, capturing carbon-nitrogen dynamics, and monitoring hydrological cycles. Traditional numerical metrics (e.g., R-squared, root mean square error) have been widely used to quantify the similarity between modeled and observed ecosystem variables, but they often fail to capture domain-specific temporal patterns critical to ecological processes. As a result, these methods are often accompanied by expert visual inspection, which requires substantial human labor and limits the applicability to large-scale evaluation. To address these challenges, we propose a novel framework that integrates metric learning with large language model (LLM)-based natural language policy extraction to develop interpretable evaluation criteria. The proposed method processes pairwise annotations and implements a policy optimization mechanism to generate and combine different assessment metrics. The results obtained on multiple datasets for evaluating the predictions of crop gross primary production and carbon dioxide flux have confirmed the effectiveness of the proposed method in capturing target assessment preferences, including both synthetically generated and expert-annotated model comparisons. The proposed framework bridges the gap between numerical metrics and expert knowledge while providing interpretable evaluation policies that accommodate the diverse needs of different ecosystem modeling studies.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
V. A. Boiko, R. MacKnight, and B. Kline. 2023. Emergent autonomous scientific research capabilities of large language models.Nature620 (2023), 47–55. https: //doi.org/10.1038/s41586-023-06404-x
-
[2]
Olivier Boucher, Jérôme Servonnat, Anna Lea Albright, Olivier Aumont, Yves Balkanski, Vladislav Bastrikov, Slimane Bekki, Rémy Bonnet, Sandrine Bony, Laurent Bopp, Pascale Braconnot, Patrick Brockmann, Patricia Cadule, Arnaud Caubel, Frederique Cheruy, Francis Codron, Anne Cozic, David Cugnet, Fabio D’Andrea, Paolo Davini, Casimir de Lavergne, Sébastien D...
2020
-
[3]
Cheng-Han Chiang and Hung-yi Lee. 2023. Can large language models be an alternative to human evaluations?arXiv preprint arXiv:2305.01937(2023)
arXiv 2023
-
[4]
Hannah L Cloke and Florian Pappenberger. 2008. Evaluating forecasts of extreme events for hydrological applications: an approach for screening unfamiliar per- formance measures.Meteorological Applications: A journal of forecasting, practical applications, training techniques and modelling15, 1 (2008), 181–197
work page 2008
-
[5]
Nathan Collier, Forrest M Hoffman, David M Lawrence, Gretchen Keppel-Aleks, Charles D Koven, William J Riley, Mingquan Mu, and James T Randerson. 2018. The International Land Model Benchmarking (ILAMB) system: design, theory, and implementation.Journal of Advances in Modeling Earth Systems10, 11 (2018), 2731–2754
work page 2018
-
[6]
K Cuddington, M-J Fortin, LR Gerber, Alan Hastings, A Liebhold, M O’connor, and C Ray. 2013. Process-based models are required to manage ecological systems in a changing world.Ecosphere4, 2 (2013), 1–12
work page 2013
-
[7]
Danabasoglu, J.-F
G. Danabasoglu, J.-F. Lamarque, J. Bacmeister, D. A. Bailey, A. K. DuVivier, J. Edwards, L. K. Emmons, J. Fasullo, R. Garcia, A. Gettelman, C. Hannay, M. M. Holland, W. G. Large, P. H. Lauritzen, D. M. Lawrence, J. T. M. Lenaerts, K. Lindsay, W. H. Lipscomb, M. J. Mills, R. Neale, K. W. Oleson, B. Otto-Bliesner, A. S. Phillips, W. Sacks, S. Tilmes, L. van...
2020
-
[8]
David R Easterling, Kenneth E Kunkel, Michael F Wehner, and Liqiang Sun. 2016. Detection and attribution of climate extremes in the observed record.Weather and Climate Extremes11 (2016), 17–27
work page 2016
Show all 44 references
-
[9]
Eyring, S
V. Eyring, S. Bony, G. A. Meehl, C. A. Senior, B. Stevens, R. J. Stouffer, and K. E. Taylor. 2016. Overview of the Coupled Model Intercomparison Project Phase 6 (CMIP6) experimental design and organization.Geoscientific Model Development 9, 5 (2016), 1937–1958. https://doi.org...
2016 doi
-
[10]
Simone Fatichi, Enrique R Vivoni, Fred L Ogden, Valeriy Y Ivanov, Benjamin Mirus, David Gochis, Charles W Downer, Matteo Camporese, Jason H Davison, Brian Ebel, et al. 2016. An overview of current applications, challenges, and future trends in distributed process-based models ...
2016
-
[11]
Gregory Flato, Jochem Marotzke, Babatunde Abiodun, Pascale Braconnot, Sin Chan Chou, William Collins, Peter Cox, Fatima Driouech, Seita Emori, Veronika Eyring, et al. 2014. Evaluation of climate models. InClimate change 2013: the physical science basis. Contribution of Working...
2014
-
[12]
Robert C Garrett, Trevor Harris, Bo Li, and Zhuo Wang. 2024. Validating Cli- mate Models with Spherical Convolutional Wasserstein Distance.arXiv preprint arXiv:2401.14657(2024)
2024 arXiv
-
[13]
Martin Gauch, Frederik Kratzert, Oren Gilon, Hoshin Gupta, Juliane Mai, Grey Nearing, Bryan Tolson, Sepp Hochreiter, and Daniel Klotz. 2023. In defense of met- rics: Metrics sufficiently encode typical human preferences regarding hydrologi- cal model performance.Water Resource...
2023
-
[14]
Gutjahr, D
O. Gutjahr, D. Putrasahan, K. Lohmann, J. H. Jungclaus, J.-S. von Storch, N. Brüggemann, H. Haak, and A. Stössel. 2019. Max Planck Institute Earth System Model (MPI-ESM1.2) for the High-Resolution Model Intercomparison Project (HighResMIP).Geoscientific Model Development12, 7 ...
2019 doi
-
[15]
Matthew R Hipsey, Louise C Bruce, Casper Boon, Brendan Busch, Cayelan C Carey, David P Hamilton, Paul C Hanson, Jordan S Read, Eduardo De Sousa, Michael Weber, et al. 2019. A General Lake Model (GLM 3.0) for linking with high- frequency sensor data from the Global Lake Ecologi...
2019
-
[16]
Matthew R Hipsey, Louise C Bruce, Casper Boon, Brendan Busch, Cayelan C Carey, David P Hamilton, Paul C Hanson, Jordan S Read, Eduardo de Sousa, Michael Weber, et al. 2019. A General Lake Model (GLM 3.0) for linking with high- frequency sensor data from the Global Lake Ecologi...
2019
-
[17]
IPCC. 2022. IPCC Sixth Assessment Report. (2022)
2022
-
[18]
Xiaowei Jia, Jacob Zwart, Jeffrey Sadler, Alison Appling, Samantha Oliver, Steven Markstrom, Jared Willard, Shaoming Xu, Michael Steinbach, Jordan Read, et al
-
[19]
David Katzin, Eldert J Van Henten, and Simon Van Mourik. 2022. Process-based greenhouse climate models: Genealogy, current status, and future directions. Agricultural Systems198 (2022), 103388
2022
-
[20]
Seyed Mehran Kazemi, Rishab Goel, Sepehr Eghbali, Janahan Ramanan, Jaspreet Sahota, Sanjay Thakur, Stella Wu, Cathal Smyth, Pascal Poupart, and Marcus Brubaker. 2019. Time2vec: Learning a vector representation of time.arXiv preprint arXiv:1907.05321(2019)
2019 arXiv
-
[21]
Popi Konidari and Dimitrios Mavrakis. 2007. A multi-criteria evaluation method for climate change mitigation policy instruments.Energy Policy35, 12 (2007), 6235–6257
2007
-
[22]
Hyunwook Lee, Chunggi Lee, Hongkyu Lim, and Sungahn Ko. 2024. TILDE- Q: A Transformation Invariant Loss Function for Time-Series Forecasting. arXiv:2210.15050 [cs.LG] https://arxiv.org/abs/2210.15050
2024 arXiv
-
[23]
Yen-Ting Lin and Yun-Nung Chen. 2023. Llm-eval: Unified multi-dimensional automatic evaluation for open-domain conversations with large language models. Conference’17, July 2017, Washington, DC, USA Qi Cheng, Licheng Liu, Qing Zhu, Runlong Yu, Zhenong Jin, Yiqun Xie, and Xiaow...
2023 arXiv
-
[24]
Licheng Liu, Wang Zhou, Kaiyu Guan, Bin Peng, Shaoming Xu, Jinyun Tang, Qing Zhu, Jessica Till, Xiaowei Jia, Chongya Jiang, et al. 2024. Knowledge-guided machine learning can improve carbon cycle quantification in agroecosystems. Nature communications15, 1 (2024), 357
2024
-
[25]
Licheng Liu, Wang Zhou, Kaiyu Guan, Bin Peng, Shaoming Xu, Jinyun Tang, Qing Zhu, Jessica Till, Xiaowei Jia, Chongya Jiang, Sheng Wang, Ziqi Qin, Hui Kong, Robert Grant, Symon Mezbahuddin, Vipin Kumar, and Zhenong Jin. 2024. Knowledge-guided machine learning can improve carbon...
2024
-
[26]
2015.PRMS-IV, the precipitation- runoff modeling system, version 4
Steven L Markstrom, R Steve Regan, Lauren E Hay, Roland J Viger, Richard M Webb, Robert A Payn, and Jacob H LaFontaine. 2015.PRMS-IV, the precipitation- runoff modeling system, version 4. Technical Report. US Geological Survey
2015
-
[27]
Tung Nguyen, Johannes Brandstetter, Ashish Kapoor, Jayesh K Gupta, and Aditya Grover. 2023. ClimaX: A foundation model for weather and climate.arXiv preprint arXiv:2301.10343(2023)
2023 arXiv
-
[28]
Alexander Nikitin, Letizia Iannucci, and Samuel Kaski. 2023. TSGM: A Flexible Framework for Generative Modeling of Synthetic Time Series.arXiv preprint arXiv:2305.11567(2023)
2023 arXiv
-
[29]
NOAA. 2024. Monthly Averages of Carbon Dioxide Flask measurements at Trinidad Head, California, United States. https://gml.noaa.gov/data/dataset.php? item=thd-co2-flask-month
2024
-
[30]
Arjun Panickssery, Samuel R Bowman, and Shi Feng. 2024. Llm evaluators recognize and favor their own generations.arXiv preprint arXiv:2404.13076 (2024)
2024 arXiv
-
[31]
Gilberto Pastorello, Carlo Trotta, Eleonora Canfora, Housen Chu, Danielle Chris- tianson, You-Wei Cheah, Cristina Poindexter, Jiquan Chen, Abdelrahman El- bashandy, Marty Humphrey, et al. 2020. The FLUXNET2015 dataset and the ONEFlux processing pipeline for eddy covariance dat...
2020
-
[32]
Zhen Qin, Rolf Jagerman, Kai Hui, Honglei Zhuang, Junru Wu, Le Yan, Jiaming Shen, Tianqi Liu, Jialu Liu, Donald Metzler, Xuanhui Wang, and Michael Ben- dersky. 2024. Large Language Models are Effective Text Rankers with Pairwise Ranking Prompting. arXiv:2306.17563 [cs.IR] http...
2024 arXiv
-
[33]
Markus Reichstein, Gustau Camps-Valls, Bjorn Stevens, Martin Jung, Joachim Denzler, Nuno Carvalhais, and F Prabhat. 2019. Deep learning and process understanding for data-driven Earth system science.Nature566, 7743 (2019), 195–204
2019
-
[34]
Johannes Schmude, Sujit Roy, Will Trojak, Johannes Jakubik, Daniel Salles Civ- itarese, Shraddha Singh, Julian Kuehnert, Kumar Ankur, Aman Gupta, Christo- pher E Phillips, et al . [n. d.]. Prithvi wxc: Foundation model for weather and climate, 2024.URL https://arxiv. org/abs/2...
2024 arXiv
-
[35]
Seland, M
Ø. Seland, M. Bentsen, D. Olivié, T. Toniazzo, A. Gjermundsen, L. S. Graff, J. B. Debernard, A. K. Gupta, Y.-C. He, A. Kirkevåg, J. Schwinger, J. Tjiputra, K. S. Aas, I. Bethke, Y. Fan, J. Griesfeller, A. Grini, C. Guo, M. Ilicak, I. H. H. Karset, O. Landgren, J. Liakka, K. O....
2020
-
[36]
Sellar, Colin G
Alistair A. Sellar, Colin G. Jones, Jane P. Mulcahy, Yongming Tang, Andrew Yool, Andy Wiltshire, Fiona M. O’Connor, Marc Stringer, Richard Hill, Julien Palmieri, Stephanie Woodward, Lee de Mora, Till Kuhlbrodt, Steven T. Rumbold, Douglas I. Kelley, Rich Ellis, Colin E. Johnson...
2019
-
[37]
N. C. Swart, J. N. S. Cole, V. V. Kharin, M. Lazare, J. F. Scinocca, N. P. Gillett, J. Anstey, V. Arora, J. R. Christian, S. Hanna, Y. Jiao, W. G. Lee, F. Majaess, O. A. Saenko, C. Seiler, C. Seinen, A. Shao, M. Sigmond, L. Solheim, K. von Salzen, D. Yang, and B. Winter. 2019....
2019 doi
-
[38]
Yidong Wang, Zhuohao Yu, Zhengran Zeng, Linyi Yang, Cunxiang Wang, Hao Chen, Chaoya Jiang, Rui Xie, Jindong Wang, Xing Xie, et al. 2023. Pandalm: An automatic evaluation benchmark for llm instruction tuning optimization.arXiv preprint arXiv:2306.05087(2023)
2023 arXiv
-
[39]
Wikipedia contributors. 2021. Nash–Sutcliffe model efficiency coefficient — Wikipedia, The Free Encyclopedia. https://en.wikipedia.org/w/index.php?title= Nash%E2%80%93Sutcliffe_model_efficiency_coefficient&oldid=1055432315 [On- line; accessed 6-February-2022]
2021
-
[40]
Zhihan Yue, Yujing Wang, Juanyong Duan, Tianmeng Yang, Congrui Huang, Yunhai Tong, and Bixiong Xu. 2022. Ts2vec: Towards universal representation of time series. InProceedings of the AAAI Conference on Artificial Intelligence, Vol. 36. 8980–8987
2022
-
[41]
Wang Zhou, Kaiyu Guan, Bin Peng, Jinyun Tang, Zhenong Jin, Chongya Jiang, Robert Grant, and Symon Mezbahuddin. 2021. Quantifying carbon budget, crop yields and their responses to environmental variability using the ecosys model for US Midwestern agroecosystems.Agricultural and...
2021
-
[42]
Hanwei Zhu, Haoning Wu, Yixuan Li, Zicheng Zhang, Baoliang Chen, Lingyu Zhu, Yuming Fang, Guangtao Zhai, Weisi Lin, and Shiqi Wang. 2024. Adaptive Image Quality Assessment via Teaching Large Multimodal Model to Compare. arXiv preprint arXiv:2405.19298(2024)
2024 arXiv
-
[43]
Tilo Ziehn, Matthew Chamberlain, Rachel Law, Andrew Lenton, Roger Bodman, Martin Dix, Lauren Stevens, Yingping Wang, and Jhan Srbinovsky. 2020. The Australian Earth System Model: ACCESS-ESM1.5.Journal of Southern Hemisphere Earth Systems Science70 (08 2020). https://doi.org/10...
2020 doi
-
[2021]
InProceedings of the 2021 SIAM International Conference on Data Mining (SDM)
Physics-guided recurrent graph model for predicting flow and temperature in river networks. InProceedings of the 2021 SIAM International Conference on Data Mining (SDM). SIAM, 612–620
2021
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.