Web scraping: data analysis and extraction with Python
Your analyses need web data that has been collected and structured methodically. Use Python to prepare scraping scripts and turn information into usable datasets. Build your ability to develop reproducible data collection processes independently and verify the quality of the results.
- Duration
- 3 days 21 hours
- Code
- WSAEP Code
Presentation
In today's digital world, data is an extremely valuable source of information. Scraping involves extracting this data from websites. There are several ways to do this, and using Python with tools such as the Beautiful Soup 4 library is one of the most effective approaches.
Designed primarily for developers, data analysts, integration specialists and business intelligence consultants, this 3-day course provides in-depth knowledge and skills in data scraping with Python. You will learn to collect relevant, usable data by applying modern analysis techniques. The 8 Python scraping modules explore these processes in detail and introduce widely used scraping tools and strategies.
With an engaging, participatory approach, this course is intended for beginners who want to learn essential data analysis tools in a short time. Numerous practical exercises use real business scenarios so that you can practise and apply your new skills directly in an appropriate environment.
Objectives
The web scraping with Python course enables you to develop the following skills:
- scrape, isolate, modify and delete data using an appropriate, systematic method;
- understand and apply modern Python techniques to convert data into usable datasets;
- implement an effective scraping strategy across different sources to collect relevant data;
- code a script with a loop to scrape efficiently;
- perform collection operations to supply structured volumes of data to a data lake.
Program
Module 1: understanding Python data structures
- The concept of retrieving data with Python.
- The basic components of a data structure: lists, tuples, sequences, sets and dictionaries.
Module 2: applying advanced techniques to built-in data structures
- The specific characteristics of Python's built-in data structures.
- Common operations on data files.
Module 3: using the NumPy and Pandas libraries
- Steps for creating arrays with NumPy.
- Steps for creating DataFrames with Pandas.
- Data visualisation and statistics with NumPy and Pandas.
- Calculating general statistics on DataFrames using the NumPy and Pandas modules.
Module 4: preparing data with Python (data wrangling)
- What is data wrangling?
- What processes are involved in data wrangling?
- Data subsets, filtering and partitioning.
- Finding outliers and handling incorrect data values.
- Concatenating, joining and merging data.
- Data wrangling techniques with Pandas.
- Advanced use of lists and the zip command.
- Data formatting techniques.
Module 5: basic scraping with Python
- What is scraping?
- Levels of complexity depending on the chosen source.
- Recognising data extracted from multiple text-based and non-text-based web pages.
- The different tools available for scraping.
- Introduction to the BeautifulSoup library.
- Using CSS Select.
Module 6: expert-level scraping with Python
- Fundamentals of web scraping.
- The importance of working with the BeautifulSoup library.
- Using Python as an extraction, transformation and loading (ETL) tool.
- Scraping structured .csv, .xml and .json files.
- File reading and writing processes.
- Analysing file data from multiple sources.
- Features for accessing and processing data in linear blocks.
Module 7: implementing a Python scraper
- Common scraping techniques, such as GET requests and sequential pages.
- Establishing an on-page search strategy to identify data quickly.
- Creating an algorithm for automated scraping.
- Sending data to a web page.
- Methods for retrieving specific data.
- Using POST and GET requests.
- Navigation techniques for identifying website data.
- Establishing a navigation strategy.
- Developing the scraper.
Module 8: using scraping in everyday work
- Cross-disciplinary data wrangling and data scraping knowledge applicable to real-world situations.
Audience
This course is intended for:
- professionals and businesses in any sector interested in Big Data, particularly data scientists, data analysts, business analysts, web developers and marketers.
Prerequisites
The web scraping with Python course requires the following prerequisites:
- basic programming and algorithm knowledge and skills;
- a good understanding of Python fundamentals.
Teaching and assessment methods
- Initial skills assessment
- Training materials provided to participants
- Continuous assessment throughout the course
- End-of-course feedback questionnaire
- Combination of theory and practical application
- Attendance records
- Post-course follow-up evaluation
- Quiz / multiple-choice questions
- Practical exercises
- Case study
Course highlights
Dates and sessions
Choose the date and delivery format that suit you.
No upcoming sessions are currently available.
Session alerts
Training content provided in partnership with Softeam Institute (in French)
Python is a registered trademark of the Python Software Foundation
fr
en