Analysis and correction of bias in Total Decrease in Node Impurity measures for tree-based algorithms

Sandri Marco; Zuccolotto Paola<sup>*</sup>

doi:10.1007/s11222-009-9132-0

摘要

Variable selection is one of the main problems faced by data mining and machine learning techniques. These techniques are often, more or less explicitly, based on some measure of variable importance. This paper considers Total Decrease in Node Impurity (TDNI) measures, a popular class of variable importance measures defined in the field of decision trees and tree-based ensemble methods, like Random Forests and Gradient Boosting Machines. In spite of their wide use, some measures of this class are known to be biased and some correction strategies have been proposed. The aim of this paper is twofold. Firstly, to investigate the source and the characteristics of bias in TDNI measures using the notions of informative and uninformative splits. Secondly, a bias-correction algorithm, recently proposed for the Gini measure in the context of classification, is extended to the entire class of TDNI measures and its performance is investigated in the regression framework using simulated and real data.

出版日期2010-10

全文

访问全文

收藏分享被引(34) 浏览

更新时间：2024-04-10 19:32

Analysis and correction of bias in Total Decrease in Node Impurity measures for tree-based algorithms

摘要

全文

产品服务

站内浏览

服务支持

联系方式

科研之友