{"work":{"id":"db5dbf68-75ca-4123-ac22-d1b5b056831d","openalex_id":"https://openalex.org/W1554944419","doi":"10.1007/978-0-387-84858-7","arxiv_id":"recordID/1416361","raw_key":null,"title":"The Elements of Statistical Learning: Data Mining, Inference, and Prediction","authors":[{"given":"Trevor","family":"Hastie","sequence":"first","affiliation":[]},{"given":"Robert","family":"Tibshirani","sequence":"additional","affiliation":[]},{"given":"Jerome","family":"Friedman","sequence":"additional","affiliation":[]}],"authors_text":"doi:10","year":2009,"venue":"Springer Series in Statistics","abstract":null,"external_url":"https://doi.org/10.1007/978-0-387-84858-7","cited_by_count":21949,"metadata_source":"doi_reference","metadata_fetched_at":"2026-07-10T05:16:47.784177+00:00","pith_arxiv_id":null,"created_at":"2026-05-08T21:14:13.358160+00:00","updated_at":"2026-07-11T11:50:26.030339+00:00","title_quality_ok":false,"display_title":"Data Mining, Inference, and Prediction","render_title":"Data Mining, Inference, and Prediction"},"hub":{"state":{"work_id":"db5dbf68-75ca-4123-ac22-d1b5b056831d","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":38,"external_cited_by_count":21949,"distinct_field_count":18,"first_pith_cited_at":"2023-05-10T16:16:24+00:00","last_pith_cited_at":"2026-07-09T15:03:24+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T08:29:37.845705+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":3},{"context_role":"method","n":1}],"polarity_counts":[{"context_polarity":"background","n":3},{"context_polarity":"use_method","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"Data Mining, Inference, and Prediction","claims":[{"claim_text":"deployment because the evaluation protocol does not reflect real-world conditions or because issues such as data leakage and distribution shift are overlooked [1, 2]. This gap between apparent validation success and operational performance highlights the need for more rigorous and context-aware evaluation methods. The primary objective of model evaluation is to estimate how well a learned model generalizes to unseen data [3]. However, generalization cannot be reduced to a single universal criter","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"One way to explain these results is the basic fact that BSRBF-KAN, FastKAN, FasterKAN and MLP have more parameters than the LTBs-KAN architecture. In addition, FastKAN,and FasterKANuse a linear function as spline which allows to reduce complexity during training improving the learnability of the information in the data set. This is a phenomenon that happens in learning known as the \"bias-variance dilemma\" [43] which points out to the need to increase the size of the dataset when using more compl","claim_type":"background","confidence":0.75,"evidence_strength":"citation_context"},{"claim_text":"For each outer foldk= 1, . . . , K out, the training data are passed to the inner loop, where cross-validation is performed to select optimal hyper- parameters ˆθk. The model trained on the corresponding inner training set is then evaluated on the outer test fold. The overall error estimate is obtained as: ˆE= 1 Kout KoutX k=1 L y(k), f(x(k); ˆθk) \u0001 ,(20) whereL(·,·) denotes the loss function. In this study, the root mean squared error (RMSE) was used. As Eq. (20) indicates, the error estimate r","claim_type":"method","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"Figure 9: Hyperparameter analysis for adaptive Material Fingerprinting applied to skin data. Figure 10: Stress-stretch data and discovered model for skin with hyperparametersn a =1 ands=0.7, averageR 2 =0.3777. framework adds one anisotropic term and discovers the following transversely isotropic 2-term model ˜W= +2.3271·10 −3 3X j=2 X k<j h exp (9.08[λjλk −1]) i +7.4216·10 2 log(cosh(0.60 [λa −1])) 2. (15) This model achieves a substantially improved average accuracy ofR2 =0.8441, see Fig. 11, ","claim_type":"background","confidence":0.5,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Data Mining, Inference, and Prediction because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (3 contexts).","role_counts":[{"n":3,"context_role":"background"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-06-05T22:10:22.353489+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"0cdf7ecb-6a89-4080-b8ba-a88f33bbd6f8","orcid":null,"display_name":"Trevor Hastie"},{"id":"98228047-727b-48f3-bf88-19ac20e639a4","orcid":null,"display_name":"Robert Tibshirani"},{"id":"6d6aa47c-0a15-4e23-930e-ea9eb4ef526a","orcid":null,"display_name":"Jerome Friedman"}]},"error":null,"updated_at":"2026-06-05T22:10:22.818104+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-06-05T22:10:21.432843+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"Masset, R","work_id":"238df2e4-a3e5-46f3-860e-3ae2b0094b97","shared_citers":3},{"title":"Arlot, A","work_id":"d68fb59d-d676-45da-8797-d71516451377","shared_citers":2},{"title":"Emre Celebi, Hassan A","work_id":"6a633aeb-c538-417e-b44f-00ba234707dd","shared_citers":2},{"title":"Fr ¨anti, S","work_id":"fed59631-8880-4011-9ae1-388a67b90f05","shared_citers":2},{"title":"Gradient-based learning applied to document recognition","work_id":"0a3595ca-57f9-43f8-8e2f-aface7154b99","shared_citers":2},{"title":"Layer Normalization","work_id":"20a2d720-0046-4c7c-bcd6-327ec8143f69","shared_citers":2},{"title":"$\\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains","work_id":"6a8d8dc4-0cc0-4052-8109-abbcdcd4a962","shared_citers":1},{"title":"$\\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment","work_id":"3a498b1a-455f-4667-b572-c5216c99a89c","shared_citers":1},{"title":"10.1198/jcgs.2011.10139","work_id":"b136f4f7-48c3-4b67-96b4-5f71034f71e6","shared_citers":1},{"title":"12\", number=","work_id":"f5a8511e-8e15-43be-8411-b0fd84dfe52f","shared_citers":1},{"title":"1971 , edition =","work_id":"dfacd2f0-e4d7-42ec-9402-879fa4849a13","shared_citers":1},{"title":"1976 , institution =","work_id":"57e97e41-1960-45bf-9afb-4f7abc365a16","shared_citers":1},{"title":"1988 , pages =","work_id":"687bdc88-a2cf-4026-b0fc-2b2dd7c6dd7f","shared_citers":1},{"title":"1988 , publisher =","work_id":"57619237-9cda-42ae-97f6-68b4ff3398c5","shared_citers":1},{"title":"1998 , isbn =","work_id":"8d97e741-ec48-4237-ac26-76985a68b195","shared_citers":1},{"title":"1998 , issn =","work_id":"37048c60-7ed8-4247-85e4-87e984957bbf","shared_citers":1},{"title":"2000 , journal =","work_id":"4bc35884-ada3-4c0f-8df4-cdc204b5595c","shared_citers":1},{"title":"2002 , issue_date =","work_id":"f721f5bc-0f50-4f61-acfd-c3d75b6f3285","shared_citers":1},{"title":"2002 , publisher =","work_id":"2f38236e-daf2-4a44-bcce-e1d109763864","shared_citers":1},{"title":"2003 , issue_date =","work_id":"5c49c420-c1e3-4977-a4cd-ba631ff59c6f","shared_citers":1},{"title":"2005 , keywords =","work_id":"78b2abf3-ce5f-48e9-b2a0-67223d124e97","shared_citers":1},{"title":"2006 , publisher=","work_id":"eb90b940-666b-47bb-bf52-e3f57107b00a","shared_citers":1},{"title":"2006 , publisher=","work_id":"974cb542-0353-4696-bd7a-0d55e49e722f","shared_citers":1},{"title":"(2007).An introduction to categorical data analysis(2nd ed.)","work_id":"b81af1c4-00c0-4dc2-83a6-87421addbce6","shared_citers":1}],"time_series":[{"n":1,"year":2023},{"n":2,"year":2024},{"n":2,"year":2025},{"n":16,"year":2026}],"dependency_candidates":[{"n":1,"role":"method","polarity":"use_method","paper_title":"Smart Ensemble Learning Framework for Predicting Groundwater Heavy Metal Pollution","primary_cat":"cs.LG","context_text":"For each outer foldk= 1, . . . , K out, the training data are passed to the inner loop, where cross-validation is performed to select optimal hyper- parameters ˆθk. The model trained on the corresponding inner training set is then evaluated on the outer test fold. The overall error estimate is obtained as: ˆE= 1 Kout KoutX k=1 L y(k), f(x(k); ˆθk) \u0001 ,(20) whereL(·,·) denotes the loss function. In this study, the root mean squared error (RMSE) was used. As Eq. (20) indicates, the error estimate represents the average performance across outer folds, thus reflecting generalisation capacity. A configuration ofK out = 5 andK in = 5 was implemented, balancing bias and variance while maintaining computational feasibility.","citing_arxiv_id":"2605.00056"}]},"error":null,"updated_at":"2026-06-05T22:10:22.855224+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-06-05T22:09:59.548096+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"Data Mining, Inference, and Prediction","claims":[{"claim_text":"deployment because the evaluation protocol does not reflect real-world conditions or because issues such as data leakage and distribution shift are overlooked [1, 2]. This gap between apparent validation success and operational performance highlights the need for more rigorous and context-aware evaluation methods. The primary objective of model evaluation is to estimate how well a learned model generalizes to unseen data [3]. However, generalization cannot be reduced to a single universal criter","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"One way to explain these results is the basic fact that BSRBF-KAN, FastKAN, FasterKAN and MLP have more parameters than the LTBs-KAN architecture. In addition, FastKAN,and FasterKANuse a linear function as spline which allows to reduce complexity during training improving the learnability of the information in the data set. This is a phenomenon that happens in learning known as the \"bias-variance dilemma\" [43] which points out to the need to increase the size of the dataset when using more compl","claim_type":"background","confidence":0.75,"evidence_strength":"citation_context"},{"claim_text":"For each outer foldk= 1, . . . , K out, the training data are passed to the inner loop, where cross-validation is performed to select optimal hyper- parameters ˆθk. The model trained on the corresponding inner training set is then evaluated on the outer test fold. The overall error estimate is obtained as: ˆE= 1 Kout KoutX k=1 L y(k), f(x(k); ˆθk) \u0001 ,(20) whereL(·,·) denotes the loss function. In this study, the root mean squared error (RMSE) was used. As Eq. (20) indicates, the error estimate r","claim_type":"method","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"Figure 9: Hyperparameter analysis for adaptive Material Fingerprinting applied to skin data. Figure 10: Stress-stretch data and discovered model for skin with hyperparametersn a =1 ands=0.7, averageR 2 =0.3777. framework adds one anisotropic term and discovers the following transversely isotropic 2-term model ˜W= +2.3271·10 −3 3X j=2 X k<j h exp (9.08[λjλk −1]) i +7.4216·10 2 log(cosh(0.60 [λa −1])) 2. (15) This model achieves a substantially improved average accuracy ofR2 =0.8441, see Fig. 11, ","claim_type":"background","confidence":0.5,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Data Mining, Inference, and Prediction because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (3 contexts).","role_counts":[{"n":3,"context_role":"background"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-06-05T22:10:22.858625+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"Data Mining, Inference, and Prediction","claims":[{"claim_text":"deployment because the evaluation protocol does not reflect real-world conditions or because issues such as data leakage and distribution shift are overlooked [1, 2]. This gap between apparent validation success and operational performance highlights the need for more rigorous and context-aware evaluation methods. The primary objective of model evaluation is to estimate how well a learned model generalizes to unseen data [3]. However, generalization cannot be reduced to a single universal criter","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"One way to explain these results is the basic fact that BSRBF-KAN, FastKAN, FasterKAN and MLP have more parameters than the LTBs-KAN architecture. In addition, FastKAN,and FasterKANuse a linear function as spline which allows to reduce complexity during training improving the learnability of the information in the data set. This is a phenomenon that happens in learning known as the \"bias-variance dilemma\" [43] which points out to the need to increase the size of the dataset when using more compl","claim_type":"background","confidence":0.75,"evidence_strength":"citation_context"},{"claim_text":"For each outer foldk= 1, . . . , K out, the training data are passed to the inner loop, where cross-validation is performed to select optimal hyper- parameters ˆθk. The model trained on the corresponding inner training set is then evaluated on the outer test fold. The overall error estimate is obtained as: ˆE= 1 Kout KoutX k=1 L y(k), f(x(k); ˆθk) \u0001 ,(20) whereL(·,·) denotes the loss function. In this study, the root mean squared error (RMSE) was used. As Eq. (20) indicates, the error estimate r","claim_type":"method","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"Figure 9: Hyperparameter analysis for adaptive Material Fingerprinting applied to skin data. Figure 10: Stress-stretch data and discovered model for skin with hyperparametersn a =1 ands=0.7, averageR 2 =0.3777. framework adds one anisotropic term and discovers the following transversely isotropic 2-term model ˜W= +2.3271·10 −3 3X j=2 X k<j h exp (9.08[λjλk −1]) i +7.4216·10 2 log(cosh(0.60 [λa −1])) 2. (15) This model achieves a substantially improved average accuracy ofR2 =0.8441, see Fig. 11, ","claim_type":"background","confidence":0.5,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Data Mining, Inference, and Prediction because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (3 contexts).","role_counts":[{"n":3,"context_role":"background"},{"n":1,"context_role":"method"}]},"error":null,"updated_at":"2026-06-05T22:10:21.440536+00:00"}},"summary":{"title":"Data Mining, Inference, and Prediction","claims":[{"claim_text":"deployment because the evaluation protocol does not reflect real-world conditions or because issues such as data leakage and distribution shift are overlooked [1, 2]. This gap between apparent validation success and operational performance highlights the need for more rigorous and context-aware evaluation methods. The primary objective of model evaluation is to estimate how well a learned model generalizes to unseen data [3]. However, generalization cannot be reduced to a single universal criter","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"One way to explain these results is the basic fact that BSRBF-KAN, FastKAN, FasterKAN and MLP have more parameters than the LTBs-KAN architecture. In addition, FastKAN,and FasterKANuse a linear function as spline which allows to reduce complexity during training improving the learnability of the information in the data set. This is a phenomenon that happens in learning known as the \"bias-variance dilemma\" [43] which points out to the need to increase the size of the dataset when using more compl","claim_type":"background","confidence":0.75,"evidence_strength":"citation_context"},{"claim_text":"For each outer foldk= 1, . . . , K out, the training data are passed to the inner loop, where cross-validation is performed to select optimal hyper- parameters ˆθk. The model trained on the corresponding inner training set is then evaluated on the outer test fold. The overall error estimate is obtained as: ˆE= 1 Kout KoutX k=1 L y(k), f(x(k); ˆθk) \u0001 ,(20) whereL(·,·) denotes the loss function. In this study, the root mean squared error (RMSE) was used. As Eq. (20) indicates, the error estimate r","claim_type":"method","confidence":0.7,"evidence_strength":"citation_context"},{"claim_text":"Figure 9: Hyperparameter analysis for adaptive Material Fingerprinting applied to skin data. Figure 10: Stress-stretch data and discovered model for skin with hyperparametersn a =1 ands=0.7, averageR 2 =0.3777. framework adds one anisotropic term and discovers the following transversely isotropic 2-term model ˜W= +2.3271·10 −3 3X j=2 X k<j h exp (9.08[λjλk −1]) i +7.4216·10 2 log(cosh(0.60 [λa −1])) 2. (15) This model achieves a substantially improved average accuracy ofR2 =0.8441, see Fig. 11, ","claim_type":"background","confidence":0.5,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks Data Mining, Inference, and Prediction because it crossed a citation-hub threshold. Current citing contexts most often use it as background evidence (3 contexts).","role_counts":[{"n":3,"context_role":"background"},{"n":1,"context_role":"method"}]},"graph":{"co_cited":[{"title":"Masset, R","work_id":"238df2e4-a3e5-46f3-860e-3ae2b0094b97","shared_citers":3},{"title":"Arlot, A","work_id":"d68fb59d-d676-45da-8797-d71516451377","shared_citers":2},{"title":"Emre Celebi, Hassan A","work_id":"6a633aeb-c538-417e-b44f-00ba234707dd","shared_citers":2},{"title":"Fr ¨anti, S","work_id":"fed59631-8880-4011-9ae1-388a67b90f05","shared_citers":2},{"title":"Gradient-based learning applied to document recognition","work_id":"0a3595ca-57f9-43f8-8e2f-aface7154b99","shared_citers":2},{"title":"Layer Normalization","work_id":"20a2d720-0046-4c7c-bcd6-327ec8143f69","shared_citers":2},{"title":"$\\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains","work_id":"6a8d8dc4-0cc0-4052-8109-abbcdcd4a962","shared_citers":1},{"title":"$\\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment","work_id":"3a498b1a-455f-4667-b572-c5216c99a89c","shared_citers":1},{"title":"10.1198/jcgs.2011.10139","work_id":"b136f4f7-48c3-4b67-96b4-5f71034f71e6","shared_citers":1},{"title":"12\", number=","work_id":"f5a8511e-8e15-43be-8411-b0fd84dfe52f","shared_citers":1},{"title":"1971 , edition =","work_id":"dfacd2f0-e4d7-42ec-9402-879fa4849a13","shared_citers":1},{"title":"1976 , institution =","work_id":"57e97e41-1960-45bf-9afb-4f7abc365a16","shared_citers":1},{"title":"1988 , pages =","work_id":"687bdc88-a2cf-4026-b0fc-2b2dd7c6dd7f","shared_citers":1},{"title":"1988 , publisher =","work_id":"57619237-9cda-42ae-97f6-68b4ff3398c5","shared_citers":1},{"title":"1998 , isbn =","work_id":"8d97e741-ec48-4237-ac26-76985a68b195","shared_citers":1},{"title":"1998 , issn =","work_id":"37048c60-7ed8-4247-85e4-87e984957bbf","shared_citers":1},{"title":"2000 , journal =","work_id":"4bc35884-ada3-4c0f-8df4-cdc204b5595c","shared_citers":1},{"title":"2002 , issue_date =","work_id":"f721f5bc-0f50-4f61-acfd-c3d75b6f3285","shared_citers":1},{"title":"2002 , publisher =","work_id":"2f38236e-daf2-4a44-bcce-e1d109763864","shared_citers":1},{"title":"2003 , issue_date =","work_id":"5c49c420-c1e3-4977-a4cd-ba631ff59c6f","shared_citers":1},{"title":"2005 , keywords =","work_id":"78b2abf3-ce5f-48e9-b2a0-67223d124e97","shared_citers":1},{"title":"2006 , publisher=","work_id":"eb90b940-666b-47bb-bf52-e3f57107b00a","shared_citers":1},{"title":"2006 , publisher=","work_id":"974cb542-0353-4696-bd7a-0d55e49e722f","shared_citers":1},{"title":"(2007).An introduction to categorical data analysis(2nd ed.)","work_id":"b81af1c4-00c0-4dc2-83a6-87421addbce6","shared_citers":1}],"time_series":[{"n":1,"year":2023},{"n":2,"year":2024},{"n":2,"year":2025},{"n":16,"year":2026}],"dependency_candidates":[{"n":1,"role":"method","polarity":"use_method","paper_title":"Smart Ensemble Learning Framework for Predicting Groundwater Heavy Metal Pollution","primary_cat":"cs.LG","context_text":"For each outer foldk= 1, . . . , K out, the training data are passed to the inner loop, where cross-validation is performed to select optimal hyper- parameters ˆθk. The model trained on the corresponding inner training set is then evaluated on the outer test fold. The overall error estimate is obtained as: ˆE= 1 Kout KoutX k=1 L y(k), f(x(k); ˆθk) \u0001 ,(20) whereL(·,·) denotes the loss function. In this study, the root mean squared error (RMSE) was used. As Eq. (20) indicates, the error estimate represents the average performance across outer folds, thus reflecting generalisation capacity. A configuration ofK out = 5 andK in = 5 was implemented, balancing bias and variance while maintaining computational feasibility.","citing_arxiv_id":"2605.00056"}]},"authors":[{"id":"6d6aa47c-0a15-4e23-930e-ea9eb4ef526a","orcid":null,"display_name":"Jerome Friedman","source":"manual","import_confidence":0.72},{"id":"98228047-727b-48f3-bf88-19ac20e639a4","orcid":null,"display_name":"Robert Tibshirani","source":"manual","import_confidence":0.72},{"id":"0cdf7ecb-6a89-4080-b8ba-a88f33bbd6f8","orcid":null,"display_name":"Trevor Hastie","source":"manual","import_confidence":0.72}]}}