☰
机器学习6:集成学习:集成学习思想、随机森林算法、Adaboost算法、GBDT、XGBoost
2026/10/5 10:23:33 网站建设 项目流程

集成学习思想

随机森林算法

''' 案例:集成学习算法之Bagging思想 随机森林算法演示 集成学习: 概述:把多个弱学习器组成一个强学习器的过程->集成学习 思想: Bagging思想: 1.有放回的随机抽取样本 2.平权投票 3.可以并行执行 Boosting思想: 1.每次训练都会使用全部样本 2.加权投票->预测正确:权重降低,预测错误:权重增加 3.只能串行执行 Bagging思想代表: 随机森林算法 随机森林算法: 1.每个弱学习器都是CART树,必须是二叉树 2.有放回的随机抽样,平权投票,并行执行 ''' #导包 import numpy as np import pandas as pd from sklearn.model_selection import train_test_split#划分数据集 from sklearn.ensemble import RandomForestClassifier#随机森林分类器 from sklearn.tree import DecisionTreeClassifier#决策树分类器 from sklearn.model_selection import GridSearchCV#网格搜索 #1.加载数据 data=pd.read_csv('./datas/train.csv') data.info() #2.数据预处理 #2.1抽取特征和标签 x=data[['Pclass','Sex','Age']].copy() y=data['Survived'] #2.2空值处理,用Age列的平均值填充缺失值 x['Age']=x['Age'].fillna(x['Age'].mean()) #2.3热编码处理 x=pd.get_dummies(x) #2.4划分训练集和测试集 x_train,x_test,y_train,y_test=train_test_split(x,y,test_size=0.2,random_state=10) #3.特征工程 #4.模型训练,预测,评估 #场景1:单一决策树 #4.1创建决策树对象,演示单一的决策树效果 estimator1=DecisionTreeClassifier() #4.2模型训练 estimator1.fit(x_train,y_train) #4.3模型预测 y_pre1=estimator1.predict(x_test) print(f'模型1预测结果:{y_pre1}') #4.4模型评估 print(f'决策树模型的评估准确率为:{estimator1.score(x_test,y_test)}') print('-'*23) #场景2:随机森林算法->采用默认参数 #4.1创建随机森林对象,演示多个的决策树(Bagging思想)效果 estimator2=RandomForestClassifier()#n_estimators=100表示有100个决策树,max_depth=10表示最大深度为10,绘制决策树时,最多10层 #4.2模型训练 estimator2.fit(x_train,y_train) #4.3模型预测 y_pre2=estimator2.predict(x_test) print(f'模型2预测结果:{y_pre2}') #4.4模型评估 print(f'随机森林模型的评估准确率为:{estimator2.score(x_test,y_test)}') print('-'*23) #场景3:随机森林算法->网格搜索 #4.1创建随机森林对象,演示多个的决策树(Bagging思想)效果 estimator3=RandomForestClassifier()#n_estimators=100表示有100个决策树,max_depth=10表示最大深度为10,绘制决策树时,最多10层 #4.2参数准备 params={'n_estimators':[30,50,60,90,100],'max_depth':[2,3,5,7]} #4.3创建网格搜索对象结合交叉验证 gs_estimator=GridSearchCV(estimator3,param_grid=params,cv=5) #4.4模型训练 gs_estimator.fit(x_train,y_train) #4.5模型预测 y_pre3=gs_estimator.predict(x_test) print(f'模型3预测结果:{y_pre3}') #4.6模型评估 print(f'随机森林模型的评估准确率为:{gs_estimator.score(x_test,y_test)}') #4.7查看最佳参数 print(f'最佳参数为:{gs_estimator.best_params_}')

Adaboost算法

PS:下面的以5.5来分时,有6个样本分类错误,下面写错了

案例AdaBoost实战葡萄酒数据

Adaptive Boost自适应提升:逐步调整权重->对(下降),错(上升)

''' 案例: 演示AdaBoost算法之葡萄酒案例 AdaBoost算法介绍: 它属于Boosting思想,即串行执行,每次使用全部样本,最后加权投票 原理: 1.使用全部样本,通过决策树模型(第一个弱分类器)进行训练,获取结果; 思路:预测正确->权重下降;预测错误->权重上升 2.把第1个弱分类器的处理结果,交给第2个弱分类器进行训练,获取结果; 思路:预测正确->权重下降;预测错误->权重上升 3.重复以上步骤,直到所有弱分类器训练完成 4.最后,根据所有弱分类器的投票结果,得到最终分类结果 思路:投票数最多的类别,就是最终分类结果 ''' #导包 import pandas as pd from sklearn.preprocessing import LabelEncoder#标签编码器 from sklearn.model_selection import train_test_split#训练集、测试集分割 from sklearn.tree import DecisionTreeClassifier#决策树分类器 from sklearn.ensemble import AdaBoostClassifier#AdaBoost分类器,集成Boosting思想 from sklearn.metrics import accuracy_score#模型评估->正确率 #1.获取数据集 df_wine=pd.read_csv('./datas/wine0501.csv') # df_wine.info() # print(df_wine['Class label'].unique())#[1 2 3]有3种类别,但是决策树只能识别二叉树 #2.数据预处理 #2.1从标签列(Class label)中,过滤掉1类别,剩下2,3,类别 df_wine=df_wine[df_wine['Class label']!=1] # print(df_wine['Class label'].unique())#[2 3]只有2,3,类别 #2.2获取特征列和标签列 x=df_wine[['Alcohol','Hue']]#酒精和色泽 y=df_wine['Class label']#标签列 #2.3打印数据 # print(x[:5]) # print(y[:5]) #2.4通过标签编码器,把标签列转换为数值列 le=LabelEncoder() y=le.fit_transform(y) # print(y)#[2,3]->[0,1] #2.5分割数据集 #参数1:特征列,参数2:标签列,参数3:测试集占比,参数4:随机种子,参数5:按类别比例分割 x_train,x_test,y_train,y_test=train_test_split(x,y,test_size=0.2,random_state=42,stratify=y) #3.特征工程 #4.模型训练 #场景1:单一决策树->充当弱分类器 #4.1创建模型对象 estimator1=DecisionTreeClassifier() #4.2训练模型 estimator1.fit(x_train,y_train) #4.3模型预测 y_pred1=estimator1.predict(x_test) print(f'单一决策树模型预测结果:{y_pred1}') #4.4模型评估 print(f'单一决策树模型正确率:{accuracy_score(y_test,y_pred1)}') #场景2:AdaBoost分类器->集成Boosting思想,CART树,200棵 #4.1创建模型对象 estimator2=AdaBoostClassifier(estimator=estimator1,n_estimators=200,learning_rate=0.5,algorithm='SAMME') #4.2训练模型 estimator2.fit(x_train,y_train) #4.3模型预测 y_pred2=estimator2.predict(x_test) print(f'AdaBoost分类器模型预测结果:{y_pred2}') #4.4模型评估 print(f'AdaBoost分类器模型正确率:{accuracy_score(y_test,y_pred2)}')

GBDT

根据划分左右子树后,左右的目标值均值作为预测值,进行迭代计算,依次类推

这里算出来的负梯度,会被当成下一轮的目标值进行迭代计算

这个预测值是根据上次划分来计算目标值的均值作为预测值

用平方损失找本棵树的切分点,需要把负梯度传给下一个作为真实值,把切分点传给下一个作为划分点求均值作为预测值,相减得到残差(负梯度、下棵树的真实值),还需要找第二棵树的切分点,用于下棵树求均值

''' 案例: 演示Boosting思想之GBDT(Gradient Boosting Decision Tree,梯度提升树)处理泰坦尼克号数据集 GBDT梯度提升树解释: 概述:通过拟合付梯度来获取一个强学习器 流程: 1.采用所有目标值的均值,作为第一个弱学习器的预测值 2.目标值-预测值=负梯度(残差),该列的值作为第2个弱学习器的目标值 3.针对第一个弱学习器,一次计算每个分割点的最小平方和,找到最佳分割点,至此:第一个弱学习器搭建完毕 4.把上述的分割点带入第2个弱学习器,计算它的预测值=以此分割点为界,目标值的均值,即为该部分数据的预测值 5.计算第2个弱学习器的付梯度,最佳分割点,至此:第二个弱学习器搭建完毕 6.以此类推,直至程序结束 ''' #导包 import numpy as np import pandas as pd from sklearn.model_selection import train_test_split#分割数据集为训练集和测试集 from sklearn.tree import DecisionTreeClassifier#决策树分类器 from sklearn.ensemble import GradientBoostingClassifier#梯度提升树分类器 from sklearn.metrics import accuracy_score#准确率评估 from sklearn.model_selection import GridSearchCV # 网格搜索 #1.读取数据集 df=pd.read_csv('./datas/train.csv') # df.info() #2.数据预处理 #2.1提取特征和标签 x=df[['Pclass','Sex','Age']].copy() y=df['Survived'].copy() #2.2处理Age列的缺失值,用该列的均值填充 # x['Age'].fillna(x['Age'].mean(),inplace=True) x['Age']=x['Age'].fillna(x['Age'].mean()) #2.3热编码处理字符串类型 x=pd.get_dummies(x) #2.4分割数据集为训练集和测试集 x_train,x_test,y_train,y_test=train_test_split(x,y,test_size=0.2,random_state=22) #3.特征工程 #4.模型训练,预测,评估 #场景1:决策树分类器 #4.1创建模型对象 estimator=DecisionTreeClassifier() #4.2训练模型 estimator.fit(x_train,y_train) #4.3模型预测 y_pred=estimator.predict(x_test) print(f'单个决策树对象的预测结果{y_pred}') #4.4模型评估 print(f'决策树分类器的准确率为:{accuracy_score(y_test,y_pred)}')#0.7821229050279329 #场景2:梯度提升树分类器 #4.1创建模型对象 estimator2=GradientBoostingClassifier() #4.2训练模型 estimator2.fit(x_train,y_train) #4.3模型预测 y_pred=estimator2.predict(x_test) #4.4模型评估 print(f'梯度提升树分类器的准确率为:{accuracy_score(y_test,y_pred)}')#0.7541899441340782 #场景3:网格搜索和交叉验证 #4.1创建模型对象 estimator3=GradientBoostingClassifier() #4.2定义参数网格 params={ 'n_estimators':[100,200,300],#弱学习器数量 'learning_rate':[0.1,0.2,0.3],#学习率 'max_depth': [3, 5] # 树最大深度 } #4.3创建网格搜索对象 grid_search=GridSearchCV(estimator3,params,cv=5) #4.4训练模型 grid_search.fit(x_train,y_train) #4.5模型评估 print(f'网格搜索的准确率为:{grid_search.best_score_}')#0.827331823106471 print(f'网格搜索后的模型为:{grid_search.best_estimator_}')

XGBoost

最重要的是将这个打分函数记住;

并且上面的4句话,可以陈述出,具体是干什么的;

PS:最后一个判断一个树是否要进行分类,是依据打分函数的。

XGBoost算法API

bst = XGBClassifier(n_estimators, max_depth, learning_rate, objective)

xgb案例:红酒品质分类

''' 案例:通过XGBoost极限梯度提升树 完成 红酒品质分类案例 回顾:XGBoost极限梯度提升树 概述: Extreme Gradient Boosting Tree,底层采用打分函数决定是否分支 原理: Gain值=分枝前的打分-(分枝后的左子树的打分+分枝后的右子树的打分) 如果Gain值大于0,说明分枝有收益,否则不分枝 ''' #导包 import xgboost as xgb import joblib#保存和加载模型 import numpy as np import pandas as pd from collections import Counter#统计元素出现次数 from sklearn.model_selection import train_test_split, GridSearchCV # 划分数据集 from sklearn.metrics import accuracy_score,classification_report#评估模型,分类报告 from sklearn.model_selection import StratifiedKFold#分层K折交叉验证,类似网格搜索时cv=k折数 from sklearn.utils import class_weight#计算类别权重 #1.定义函数,对红酒品质分类元数据->拆分成训练集和测试集,并存储到csv文件中 def dm01_data_split(): #1.加载数据集 df=pd.read_csv('./datas/红酒品质分类.csv') #2.查看数据集 # df.info() #3.抽取特征数据和标签数据 x=df.iloc[:,:-1] y=df.iloc[:,-1]-3 #最后一列是品质(标签),品质从3开始,所以减3 #4.查看数据集 # print(x[:5]) # print(y[:5]) # print(f'查看标签结果的分布情况:{Counter(y)}') #5.切分数据集为训练集和测试集 #参1:特征数据,参2:标签数据,参3:测试集占比,参4:随机种子,参5:参考数据集的标签分布 x_train,x_test,y_train,y_test=train_test_split(x,y,test_size=0.2,random_state=42,stratify=y) #6.把上述的训练集特征和标签数据拼接到一起,测试集特征和标签数据也拼接到一起,最后写到csv文件中 pd.concat([x_train,y_train],axis=1).to_csv('./datas/红酒品质分类_训练集.csv',index=False) pd.concat([x_test,y_test],axis=1).to_csv('./datas/红酒品质分类_测试集.csv',index=False)#忽略索引 #2.定义函数,训练模型,并保存模型 def dm02_train_model(): # 1.读取训练集和测试集 train_data=pd.read_csv('./datas/红酒品质分类_训练集.csv') test_data=pd.read_csv('./datas/红酒品质分类_测试集.csv') #2.提取训练集和测试集的特征数据和标签数据 x_train=train_data.iloc[:,:-1]#除了最后一列,都是特征 y_train=train_data.iloc[:,-1]#最后一列是标签 x_test=test_data.iloc[:,:-1] y_test=test_data.iloc[:,-1] #3.创建模型对象 estimator=xgb.XGBClassifier( n_estimators=100, learning_rate=0.1, max_depth=3, random_state=42, objective='multi:softmax',#多分类问题,使用多分类模型 ) #加入平衡权重,因为数据集是样本不均衡的, #参1:平衡权重,参2:标签数据(即:参考标签数据分布,平衡权重) # class_weight.compute_class_weight('balanced',y_train) weights = class_weight.compute_class_weight('balanced', classes=np.unique(y_train), y=y_train) #4.模型训练 estimator.fit(x_train,y_train) #5.模型评估 print(f'准确率:{estimator.score(x_test,y_test)}') #6.保存模型 joblib.dump(estimator,'./model/红酒品质分类_模型.pkl')#后缀名也可以为.pth,都是pickle文件格式 print('模型保存成功') #3.定义函数,测试模型 def dm03_use_model(): # 1.读取训练集和测试集 train_data = pd.read_csv('./datas/红酒品质分类_训练集.csv') test_data = pd.read_csv('./datas/红酒品质分类_测试集.csv') # 2.提取训练集和测试集的特征数据和标签数据 x_train = train_data.iloc[:, :-1] # 除了最后一列,都是特征 y_train = train_data.iloc[:, -1] # 最后一列是标签 x_test = test_data.iloc[:, :-1] y_test = test_data.iloc[:, -1] #3.加载模型 estimator=joblib.load('./model/红酒品质分类_模型.pkl') #4.创建网格搜索+交叉验证(结合分层采样数据),找模型最优参数组合 #4.1定义变量,记录:参数组合 param_dict={'max_depth':[3,4,5,6,7],'learning_rate':[0.2,0.6,1.0,1.3],'n_estimators':[70,100,200,250]} #4.2创建分层采样对象 #参1:K折数,参2:随机种子,参3:是否打乱数据 kf=StratifiedKFold(n_splits=5,random_state=42,shuffle=True) #4.3创建网格搜索+交叉验证对象, #参1:模型对象,参2:参数组合,参3:分层采样对象 gs_estimator=GridSearchCV(estimator,param_grid=param_dict,cv=kf) #5.模型训练 gs_estimator.fit(x_train,y_train) #6.模型预测 y_pred=gs_estimator.predict(x_test) print(f'预测值为:{y_pred}') #7.打印模型评估系数 print(f'最优估计器对象组合:{gs_estimator.best_params_}') print(f'最优评分:{gs_estimator.best_score_}') print(f'准确率:{accuracy_score(y_test,y_pred)}') #4.测试 if __name__=='__main__': # dm01_data_split() # dm02_train_model() dm03_use_model()

需要专业的网站建设服务?

联系我们获取免费的网站建设咨询和方案报价,让我们帮助您实现业务目标

立即咨询