Big Data with Spark
Your analytics depend on available, well-organised data. Use Apache Spark to connect collection, transformation and delivery, giving structure to your data flows. Develop your ability to design processing that meets application and business needs.
- Duration
- 5 days 35 hours
- Code
- BDT01FR Code
Presentation
Big Data is regarded as one of the greatest computing challenges of our time. For organisations, it represents a turning point at least as significant as the Internet in its day. Its strongest advocates even see an industrial revolution comparable to the discovery of electricity in the 19th century or computing in the late 20th century.
Big Data and analytics are used in almost every field and by organisations of all sizes. Over the years, they have become a major economic and strategic issue for businesses, generally supporting objectives such as improving customer experience, optimising processes and operational performance, and strengthening or diversifying the business model.
As a result, more companies are looking for people who can analyse and manage the volumes of data generated by their network-based activities.
Apache Spark is regarded as the world's most mature and widely used framework for large-scale data analysis. It is, in essence, a successor to MapReduce, with the added advantage of bringing together many of the tools required in a Hadoop cluster.
Objectives
By the end of the course, participants will be able to:
- master Spark's fundamental concepts;
- develop applications with Spark;
- discover and understand RDDs;
- explore and manipulate data using Zeppelin;
- work with data using Spark SQL;
- understand how Spark MLlib works;
- build predictive models with Spark ML;
- set up a Spark cluster.
Program
Why Spark?
- Introduction.
- Big Data challenges.
- The MapReduce revolution.
- MapReduce versus Spark.
- Apache Hadoop and its ecosystem.
- HDFS and file storage in Hadoop: NameNode and DataNode.
- YARN and data processing in a Hadoop cluster.
- Sizing and configuring a cluster.
- Sqoop and importing data into Hadoop.
Understanding and developing with Spark
- The basics.
- Functional programming with Scala.
- Parallel programming with Scala.
- Discovering and understanding RDDs.
- Aggregating data with paired RDDs.
- Writing and running Spark applications.
- Transformations and actions.
- Configuring Spark applications.
- Running processing in a distributed environment.
- The RDD lifecycle.
- Data processing with Spark.
- DataFrames and Spark SQL.
- Working with Zeppelin.
- Advanced features and performance improvements.
- A brief introduction to Apache Flume and Apache Kafka.
Spark for data science
- Introduction to machine learning.
- Supervised and unsupervised learning.
- Testing and evaluation.
- Algorithm classes.
- Introduction to Spark ML and MLlib.
- Implementing algorithms in MLlib.
- Conclusion.
Practical exercises
- Installing and configuring Spark.
- Getting started with Spark.
- Working with different datasets using RDDs.
- Working with datasets through SQL queries.
- Connecting to an external database through JDBC.
- Setting up a Spark cluster.
- Using Spark ML and MLlib.
- HDFS commands.
- Using HDFS storage.
- Parallel programming with Spark: caching and persisting data.
- Using accumulators to check data quality and using broadcast variables.
- Advanced partitioning and operations as a starting point for optimisation.
- Working with data using Zeppelin.
- Spark SQL with UDFs.
- Spark SQL with Hive.
- Spark SQL and queries.
Audience
This course is for developers, data scientists, systems architects and technical managers who want to deploy Spark solutions in their organisations.
Prerequisites
No prior knowledge of Spark is required.
A good knowledge of Java is required.
Teaching and assessment methods
- Initial skills assessment
- Training materials provided to participants
- Continuous assessment throughout the course
- End-of-course feedback questionnaire
- Combination of theory and practical application
- Attendance records
- Post-course follow-up evaluation
Dates and sessions
Choose the date and delivery format that suit you.
No upcoming sessions are currently available.
fr
en