Build, train and validate high-performing machine learning models
A machine learning model must be evaluated against the problem it aims to solve. Structure data preparation, training and validation to make your choices explicit. Develop a technical approach that enables you to compare results and identify model limitations.
- Duration
- 2 days 14 hours
- Code
- IA034FR Code
Presentation
The value of an AI project depends on producing reliable predictive models that generalise. This 2-day course takes you beyond theory to develop the operational skills needed to build robust machine learning models. Learn to navigate every critical stage, from cleaning raw data to fine-tuning hyperparameters.
The programme emphasises experimental methodology: choosing the right regression, classification or clustering algorithm, preparing data effectively and rigorously validating results to avoid overfitting. Use leading libraries such as Scikit-learn to put these concepts into practice.
Through practical workshops using real datasets, develop your judgement in interpreting performance metrics and making sound modelling decisions. Leave with a comprehensive toolkit for designing high-performing, auditable AI solutions.
Objectives
By the end of this course, you will be able to:
- describe the stages of a machine learning model's lifecycle, from design to deployment;
- select and configure algorithms best suited to a given business problem;
- apply validation and performance measurement techniques to assess a model;
- detect and address overfitting risks and potential biases;
- optimise hyperparameters to maximise model robustness and generalisation.
Program
Module 1: Understanding the predictive modelling approach
- Defining objectives and problem types: classification, regression and clustering.
- The complete machine learning model lifecycle.
Hands-on exercises
- Identify the ML problem type from different business use cases.
Module 2: Preparing and structuring data
- Cleaning data and handling missing values.
- Encoding categorical variables for algorithmic processing.
- Feature normalisation and standardisation techniques.
Hands-on exercises
- Prepare and clean a raw dataset using Pandas and Scikit-learn.
Module 3: Training models
- Overview of classic algorithms: linear regression, decision trees, SVM and KNN.
- Configuration strategies and train, test and validation data splits.
Hands-on exercises
- Train several competing models on the same real dataset.
Module 4: Evaluating performance
- Analysing classification metrics: accuracy, precision, recall and F1-score.
- Analysing regression metrics: RMSE, MAE and R².
- Using confusion matrices and ROC/AUC curves.
Hands-on exercises
- Compare several models and interpret the results.
Module 5: Validating models and improving reliability
- Implementing cross-validation and K-fold validation.
- Detecting and addressing overfitting and underfitting.
- Applying regularisation techniques to improve generalisation.
Hands-on exercises
- Implement cross-validation and analyse performance differences.
Module 6: Optimising and selecting the final solution
- Automating hyperparameter searches through Grid Search and Random Search.
- Selecting relevant features and analysing variable importance.
- Model interpretability tools: SHAP and LIME.
Hands-on exercises
- Optimise a complex model with GridSearchCV and interpret key variables.
Audience
This course is intended for technical professionals seeking to specialise, including:
- data analysts and junior data scientists seeking a structured approach;
- developers acquiring specialist machine learning skills;
- AI project managers and technical Product Owners seeking to understand models' inner workings;
- anyone involved in predictive model design.
Prerequisites
The following prerequisites apply:
- Professional experience: initial experience in data manipulation.
- Basic knowledge:
- proficiency in Python fundamentals and its data manipulation libraries;
- fundamental statistics and machine learning knowledge.
Teaching and assessment methods
- Initial skills assessment
- Training materials provided to participants
- Continuous assessment throughout the course
- End-of-course feedback questionnaire
- Combination of theory and practical application
- Attendance records
- Post-course follow-up evaluation
- Practical exercises
Course highlights
- Complete approach: master the entire value chain, from raw data to an optimised, validated model.
- Quality focus: avoid common pitfalls such as overfitting to ensure reliable production models.
- Intensive practice: consolidate learning through 6 technical workshops covering preparation, training and optimisation.
- Practical toolkit: leave proficient in standard libraries such as Scikit-learn and Pandas, and professional evaluation methods.
Dates and sessions
Choose the date and delivery format that suit you.
No upcoming sessions are currently available.
Session alerts
Brand names and logos mentioned in this course description, such as Python, Scikit-learn and Pandas, belong to their respective owners. Their use for educational purposes does not constitute a commitment or partnership.
fr
en