Master global and local explainability for decision trees, random forests, and gradient boosting machines.

This course provides a comprehensive framework for interpreting tree-based machine learning models, moving from simple decision trees to complex ensemble methods. It covers the mechanics of tree induction, feature importance, and impurity-based metrics for both global and local model explanations. Participants will learn how to extract and visualize insights from random forests and gradient boosting machines, including practical implementations using scikit-learn, XGBoost, and LightGBM. The content emphasizes the mathematical intuition behind model predictions and provides strategies for assessing feature contributions in both regression and classification tasks.
Topics Covered:

Data Scientist | Python Developer | Author and Instructor
I'm a data scientist, machine learning educator and open-source developer. I've built machine learning models for credit risk, insurance claims and fraud prevention, and I am passionate about helping data scientists build models that hold up in real-world projects. My courses are designed for intermediate and advanced practitioners. They cover feature engineering, feature selection, hyperparameter optimization, imbalanced data and the design of robust machine learning pipelines, with a strong emphasis on techniques you can apply straight away in your own work. In the age of generative AI, when a working model can be coded in minutes, the real skill lies in understanding the methods deeply enough to review that output with rigor and a critical eye, and that's exactly what these courses are built to develop. I'm the creator and maintainer of Feature-engine, an open-source Python library for feature engineering and feature selection used by data scientists worldwide. I'm also the author of three books published by Packt: Python Feature Engineering Cookbook, Feature Selection in Machine Learning, and Imbalanced Data: Myths, Mistakes and Modern Solutions. I speak regularly at conferences and meetups, and I enjoy connecting technical communities with the tools and knowledge they need to succeed. In 2018 I received a Data Science Leaders Award, and in 2019 LinkedIn recognized me as one of its voices in data science and analytics. Before moving into data science, I earned an MSc in Biology and a PhD in Biochemistry, then spent more than eight years as a research scientist at institutions including University College London and the Max Planck Institute. That scientific training still shapes how I teach: rigorous, evidence-based, and focused on understanding why a method works before reaching for it.
Provincial regulators of CPAs in Canada do not require that independent providers of CPD be approved to offer courses. Instead, individual CPAs are responsible for assessing whether a CPD activity meets their requirements, and may take activities from any source provided those requirements are met.
Every course offered on LearnFormula is delivered by a qualified subject matter expert or learning organization, and advances learning objectives that are relevant to the responsibilities or professional competencies of Canadian CPAs. All activities on LearnFormula are quantifiable in terms of hours, and are also verifiable, in that users receive documented evidence of their attendance via a certificate of completion after finishing a course (and this certificate is stored by LearnFormula indefinitely). Nearly 100,000 Canadian CPAs successfully satisfy their CPD requirements via LearnFormula on an annual basis.