Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps
- Defining AIOps and its significance in modern IT
- Comparing traditional monitoring with AIOps-driven observability
- Key components and architecture of AIOps systems
Collection and Normalization of Operational Data
- Types of observability data: metrics, logs, and traces
- Ingesting data from diverse sources such as servers, containers, and cloud environments
- Utilizing agents and exporters (Prometheus, Beats, Fluentd)
Data Correlation and Anomaly Detection
- Time series correlation techniques and statistical methods
- Applying ML models for effective anomaly detection
- Identifying incidents within distributed systems
Alerting and Noise Reduction
- Crafting intelligent alert rules and appropriate thresholds
- Implementing suppression, deduplication, and alert grouping strategies
- Integration with Alertmanager, Slack, PagerDuty, or Opsgenie
Root Cause Analysis and Visualization
- Leveraging dashboards to visualize metrics and identify trends
- Analyzing events and timelines to support Root Cause Analysis (RCA)
- Tracking issues across layers using distributed tracing tools
Automation and Remediation
- Initiating automated scripts or workflows triggered by incidents
- Connecting with ITSM systems like ServiceNow and Jira
- Key use cases: self-healing, auto-scaling, and traffic rerouting
Open Source and Commercial AIOps Platforms
- Overview of key tools: Prometheus, Grafana, ELK, Moogsoft, and Dynatrace
- Criteria for evaluating and selecting an AIOps platform
- Demonstrations and hands-on practice with a chosen stack
Summary and Next Steps
Requirements
- A solid grasp of IT operations and system monitoring concepts
- Practical experience with monitoring tools or dashboards
- Familiarity with basic log and metric formats
Target Audience
- Operations teams managing infrastructure and applications
- Site Reliability Engineers (SREs)
- IT monitoring and observability teams