Pith. sign in

REVIEW 2 cited by

forester: A Tree-Based AutoML Tool in R

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2409.04789 v1 pith:AMNWK7TB submitted 2024-09-07 cs.LG cs.MSstat.ME

classification cs.LGcs.MSstat.ME
keywords automldataforesterlearningmachinetree-basedanalysismodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The majority of automated machine learning (AutoML) solutions are developed in Python, however a large percentage of data scientists are associated with the R language. Unfortunately, there are limited R solutions available. Moreover high entry level means they are not accessible to everyone, due to required knowledge about machine learning (ML). To fill this gap, we present the forester package, which offers ease of use regardless of the user's proficiency in the area of machine learning. The forester is an open-source AutoML package implemented in R designed for training high-quality tree-based models on tabular data. It fully supports binary and multiclass classification, regression, and partially survival analysis tasks. With just a few functions, the user is capable of detecting issues regarding the data quality, preparing the preprocessing pipeline, training and tuning tree-based models, evaluating the results, and creating the report for further analysis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Investigating the Impact of Balancing, Filtering, and Complexity on Predictive Multiplicity: A Data-Centric Perspective

    stat.ML 2024-12 reject novelty 4.0 of 10

    Balancing methods increase predictive multiplicity on imbalanced datasets, while filtering effects are inconsistent and largely non-significant in the authors' own tests.

  2. Rashomon effect in Educational Research: Why More is Better Than One for Measuring the Importance of the Variables?

    cs.CY 2024-12 reject novelty 4.0 of 10

    On OULAD demographic data, variable importance rankings differ across equally accurate tree models, and the reported accuracy improvement of the Rashomon set is largely a consequence of how the set is defined.

Pith tools