Perform a search on the site.

Your currency

Big Data with Spark

Your analytics depend on available, well-organised data. Use Apache Spark to connect collection, transformation and delivery, giving structure to your data flows. Develop your ability to design processing that meets application and business needs.

Duration
5 days 35 hours
Code
BDT01FR Code

Presentation

Big Data is regarded as one of the greatest computing challenges of our time. For organisations, it represents a turning point at least as significant as the Internet in its day. Its strongest advocates even see an industrial revolution comparable to the discovery of electricity in the 19th century or computing in the late 20th century.

Big Data and analytics are used in almost every field and by organisations of all sizes. Over the years, they have become a major economic and strategic issue for businesses, generally supporting objectives such as improving customer experience, optimising processes and operational performance, and strengthening or diversifying the business model.

As a result, more companies are looking for people who can analyse and manage the volumes of data generated by their network-based activities.

Apache Spark is regarded as the world's most mature and widely used framework for large-scale data analysis. It is, in essence, a successor to MapReduce, with the added advantage of bringing together many of the tools required in a Hadoop cluster.

Objectives

By the end of the course, participants will be able to:

  • master Spark's fundamental concepts;
  • develop applications with Spark;
  • discover and understand RDDs;
  • explore and manipulate data using Zeppelin;
  • work with data using Spark SQL;
  • understand how Spark MLlib works;
  • build predictive models with Spark ML;
  • set up a Spark cluster.
Last update: 23/09/2026