Glossary

Random forests and boosting

Definition

Random forests and boosting are ensemble methods that combine hundreds of decision trees into one more accurate model. A random forest builds trees independently on random subsets of data and variables and averages (votes) the result; boosting builds trees sequentially, each correcting the errors of the previous ones. Both rank among the most successful methods for tabular data.

Data mining & machine learning

Random forests are robust to overfitting and noise and give good results almost without tuning; boosting (stochastic gradient boosting) is often even more accurate but more sensitive to settings (number of trees, learning rate, depth). Both provide predictor importance, which partly replaces the lost interpretability.

They are used for churn prediction, risk scoring, predictive maintenance or defect detection. Always verify accuracy on test data or by cross-validation, and with imbalanced classes watch AUC and sensitivity rather than overall accuracy.

In Statistica

The Data Mining menu offers Random Forests and Boosted Trees (stochastic gradient boosting) for classification and regression, with automatic stopping by test error, predictor importance plots, a confusion matrix, ROC curve and model saving for deployment; in the Workspace environment they are easily compared with other models.

Related terms

Knowledgebase guides

FAQ

When a random forest and when boosting?
A random forest as a reliable first model without tuning; boosting when you want to squeeze out maximum accuracy and have time to tune parameters with cross-validation.
How do I explain a "black box" model?
Predictor importance plots and partial dependence plots show which variables affect the outcome and in which direction.

Try it on your own data

Statistica free for 30 days

Full version, no credit card. Or get a pricing estimate in a minute.

We'll prepare a tailored quote. Free and with no obligation.

Tell us how many users and what analyses you need — we'll get back to you with a concrete license quote, usually within a few business days.

Try for free first
Statistica.pro

Official Statistica partner for the European Union, based in Prague, Czech Republic. Operated by DataBon s.r.o.

On the market since 1995, formerly as StatSoft CR s.r.o.

Contact

DataBon s.r.o.
Korunní 2569/108, 101 00 Prague 10, Czech Republic
+420 602 284 038
Company ID 09743804 · VAT ID CZ09743804

University licence users: please contact the licence administrator at your institution first.

© 2026 DataBon · statistica.pro · Personal data processing (GDPR) ·