DevOps SRE Foundation℠: DevOps Engineering Fundamentals
Service reliability should be something you can discuss and monitor with development teams. Connect practices, automation and feedback through DevOps to organise a more consistent way of working. Develop a framework for reducing friction and supporting service improvement.
- Duration
- 2 days 16 hours
- Code
- DEVOPSEF Code
- Certification
- Site Reliability Engineering Foundation℠ Certification
Accredited training for the Site Reliability Engineering Foundation℠ certification.
Presentation
In a constantly evolving IT landscape, businesses face increasing risks. In this context, the site reliability engineer (SRE) plays a crucial role in improving system availability and resilience. Working with development teams within a DevOps approach, the SRE automates operational tasks, establishes monitoring systems and manages incidents. This rapidly growing role offers opportunities for IT professionals.
This DevOps SRE course introduces essential concepts and practices. You will explore SRE foundations and their relationship with DevOps and other methods, then learn to manage service level objectives (SLOs) and error budgets. You will also understand how to reduce toil, establish effective monitoring with SLIs and automate tasks using SRE tools. Finally, you will examine antifragility, SRE's organisational impact and future trends.
This 2-day program also prepares you to take PeopleCert's DevOps SRE Foundation certification exam, included in our offer (see the Certification tab for details). You will gain the skills and knowledge to advance your career in the expanding DevOps and SRE field.
Objectives
By the end of this DevOps SRE course, you will be able to:
- understand the origins of site reliability engineering (SRE) and its emergence at Google LLC;
- explain the relationship between SRE, DevOps and other IT engineering methodologies;
- master site reliability engineering fundamentals;
- understand service level objectives and their value to customers;
- identify service level indicators;
- implement a modern monitoring system;
- establish error budgets and define error-related strategies;
- apply good practices to reduce the toil budget;
- assess the impact of complex tasks on business performance;
- demonstrate that observability is a determining factor in service quality;
- use modern site reliability engineering tools and automation good practices;
- incorporate site antifragility engineering principles;
- explain the organisational benefits of implementing SRE in a business;
- take the exam and earn SRE Foundation℠ certification.
Program
Module 1: Understanding SRE fundamentals and practices
- What is site reliability engineering?
- The emergence of SRE at Google and its evolution.
- How do SRE and DevOps approaches differ?
- Positioning SRE alongside other frameworks such as ITIL and Agile.
- Core SRE principles: embracing risk, teamwork, automation, monitoring and more.
Module 2: Managing service level objectives and budgets
- Using service level objectives (SLOs).
- Establishing an error budget.
- Implementing an error budget policy.
- Introduction to SLI measurement methods and tools.
Module 3: Reducing the toil budget
- What does toil mean?
- Types of toil and their impact.
- Why toil is undesirable.
- The link between toil and team burnout.
- Balancing the toil budget.
Module 4: Monitoring with service level indicators
- Using service level indicators (SLIs).
- SLI types: latency, error rate, throughput and more.
- Establishing monitoring and observation.
- Introduction to monitoring and observability tools such as Prometheus, Grafana and the ELK stack.
- The difference between monitoring, logging and tracing.
Module 5: Automating with SRE software
- What does automation mean?
- Designing automation experiments.
- Classifying automation systems.
- Automation security.
- Automation software.
- Test and deployment automation (CI/CD) and its integration with SRE.
Module 6: Developing antifragility and learning from failure
- Learning from failure: examples of post-mortems and blameless post-mortems.
- Benefits of antifragility.
- Chaos Engineering and GameDay concepts.
- Rebalancing organisational hierarchies.
Module 7: Analysing SRE's organisational impact
- Identifying why businesses choose site reliability engineering:
- discussion of SRE adoption's cultural aspects: collaboration, communication and trust.
- Applicable adoption models.
- SRE organisational models: centralised, decentralised, hybrid and others.
- Support service needs.
- Thorough reviews.
- Site reliability engineering at different scales.
Module 8: Exploring developments and other approaches
- Discussion and comparison of other approaches:
- comparing SRE with ITIL, Agile and Lean.
- Anticipating SRE developments:
- discussion of future SRE trends, including AI/ML integration.
Module 9: Preparing for the DevOps SRE exam
- Taking a self-marked mock exam.
- Questions and answers.
- Guidance for the official exam.
Audience
This course is designed for:
- system engineers and administrators seeking a deeper understanding of SRE practices to improve infrastructure reliability and performance;
- DevOps engineers seeking stronger skills in integrating SRE principles into CI/CD pipelines and automation practices;
- developers wishing to understand SRE principles to design more robust applications that are easier to operate in production;
- IT and operations managers seeking a strategic view of implementing SRE;
- project managers and Scrum Masters wishing to integrate SRE practices into agile development cycles and manage projects with a reliability-focused approach;
- anyone wishing to understand and apply SRE principles.
Prerequisites
The following prerequisites apply:
- ability to read and understand English for the official exam;
- understanding of basic operating system, networking, database and software architecture concepts;
- basic knowledge of DevOps principles and practices, including continuous integration, continuous delivery and infrastructure as code;
- experience as a system administrator, system engineer, developer or operations manager (recommended).
Teaching and assessment methods
- Initial skills assessment
- Training materials provided to participants
- Continuous assessment throughout the course
- End-of-course feedback questionnaire
- Combination of theory and practical application
- Attendance records
- Post-course follow-up evaluation
- Practical exercises
- Case study
- Mock exam
Course highlights
• Exam readiness: maximise your preparation with a mock exam and the official exam included in our offer.
• Experiential learning: apply your knowledge to real situations through hands-on exercises and case studies.
Dates and sessions
Choose the date and delivery format that suit you.
No upcoming sessions are currently available.
Session alerts
Site Reliability Engineering is a service mark of the DevOps Institute.
fr
en
