Building batch data analytics solutions on Amazon EMR
Your analyses depend on available, properly organised data. With AWS, connect collection, transformation and delivery to structure data flows. Strengthen your ability to design processing that meets application and business needs.
- Duration
- 1 day 20 hours
- Code
- AWS26FR Code
Presentation
Amazon EMR (Elastic Map Reduce) simplifies running big data frameworks such as Spark and Hadoop for batch analytics. It enables efficient processing of large data volumes for ETL, log analysis and other uses. EMR offers flexibility, cost optimisation and integration with other AWS services.
This data analytics course provides the skills to master these concepts and build effective solutions. You will explore designing, deploying and optimising EMR analytics solutions on AWS using industry best practices. You will be able to implement efficient data pipelines, choose suitable AWS services, ensure data security and optimise costs.
By the end of this one-day program, you will also master the key skills for designing and implementing batch data analytics solutions and developing modern AWS data architectures.

As an authorised Amazon Web Services premium training partner (ATP), Oo2 offers skills development and certification training that meets the organisation's rigorous quality standards.
Objectives
By the end of this AWS Amazon EMR course, you will be able to:
- compare data warehouse, data lake and modern data architecture features and benefits;
- design and implement a batch data analytics solution;
identify and apply appropriate techniques, including compression, to optimise data storage; - select and deploy appropriate data ingestion, transformation and storage options;
- choose suitable instance and node types, clusters, automatic scaling and network topology for a specific business use case;
- understand how data storage and processing affect the analysis and visualisation mechanisms needed to obtain actionable insights;
- secure data at rest and in transit;
- monitor analytics workloads to identify and resolve problems;
- apply cost management best practices.
Program
Module 1: understanding data analytics and the data pipeline
- Data analytics use cases.
- Using the data pipeline for analytics.
Module 2: discovering Amazon EMR
- Using Amazon EMR in analytics solutions.
- Amazon EMR cluster architecture.
- Cost management strategies.
Labs:
- Launch an Amazon EMR cluster.
Module 3: exploring the data analytics pipeline
- Ingestion and storage.
- Storage optimisation with Amazon EMR.
- Data ingestion techniques.
Module 4: high-performance batch analytics with Apache Spark
- Apache Spark use cases on Amazon EMR.
- Why use Apache Spark on Amazon EMR?
- Spark concepts.
- Transformation, processing and analysis.
- Using notebooks with Amazon EMR.
Labs:
- Connect to an EMR cluster and run Scala commands using the Spark shell.
- Perform low-latency data analysis with Apache Spark on Amazon EMR.
Module 5: processing and analysing batch data with Amazon EMR and Apache Hive
- Using Amazon EMR with Hive for batch processing.
- Transformation, processing and analysis.
- Introduction to Apache HBase on Amazon EMR.
Labs:
- Process batch data using Amazon EMR with Hive.
Module 6: serverless data processing
- Serverless data processing, transformation and analysis.
- Using AWS Glue with Amazon EMR workloads.
Labs:
- Orchestrate Spark data processing using AWS Step Functions.
Module 7: cluster security and monitoring
- Introduction to securing EMR clusters.
Labs :
- Client-side encryption with EMRFS.
- Review Apache Spark cluster history.
Module 8: designing batch data analytics solutions
- Batch data analytics use cases.
Practical work:
- Design a batch data analytics workflow.
Module 9: developing modern data architectures on AWS
- Introduction to modern data architectures.
Audience
This course is intended for:
- data platform engineers designing and deploying data analytics solutions;
- data architects and operators responsible for building and managing analytics pipelines.
Prerequisites
This AWS course requires the following prerequisite:
- At least 1 year of experience managing open-source data frameworks, such as Apache Spark or Apache Hadoop.
Teaching and assessment methods
- Initial skills assessment
- Training materials provided to participants
- Continuous assessment throughout the course
- End-of-course feedback questionnaire
- Combination of theory and practical application
- Attendance records
- Post-course follow-up evaluation
- Practical exercises
- Case study
Course highlights
- AWS-certified expert trainers: benefit from recognised, AWS-certified data analytics expertise.
- Interactive practical learning: master AWS analytics tools and techniques through demonstrations and exercises, preparing for real solution development challenges.
- Key skills development: content closely aligned with designing and implementing batch analytics solutions and developing modern data architectures on AWS.
Dates and sessions
Choose the date and delivery format that suit you.
No upcoming sessions are currently available.
Session alerts
AWS and Amazon EMR are registered trademarks of Amazon.com, Inc. or its affiliates.
fr
en