Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps Using Open Source Tools
- Overview of AIOps concepts and their advantages.
- The role of Prometheus and Grafana within the observability stack.
- Positioning ML in AIOps: contrasting predictive and reactive analytics.
Configuring Prometheus and Grafana
- Installation and configuration of Prometheus for time series data collection.
- Building Grafana dashboards using real-time metrics.
- Examining exporters, relabeling rules, and service discovery mechanisms.
Data Preprocessing for Machine Learning
- Extraction and transformation of Prometheus metrics.
- Preparing datasets optimized for anomaly detection and forecasting.
- Utilizing Grafana transformations or Python-based pipelines.
Machine Learning for Anomaly Detection
- Implementation of basic ML models for outlier detection, such as Isolation Forest and One-Class SVM.
- Training and evaluating models using time series data.
- Visualizing detected anomalies within Grafana dashboards.
Metric Forecasting with Machine Learning
- Developing forecasting models, including ARIMA, Prophet, and an introduction to LSTM.
- Predicting system load and resource utilization.
- Leveraging predictions for early warning systems and scaling decisions.
Integrating ML into Alerting and Automation
- Formulating alert rules based on ML outputs or defined thresholds.
- Managing notifications using Alertmanager and routing configurations.
- Executing scripts or automation workflows triggered by anomaly detection.
Scaling and Operationalizing AIOps
- Integration with external observability platforms like the ELK stack, Moogsoft, or Dynatrace.
- Operationalizing ML models within observability pipelines.
- Best practices for implementing AIOps at scale.
Recap and Future Directions
Requirements
- A solid grasp of system monitoring and observability fundamentals.
- Practical experience with Grafana or Prometheus.
- Proficiency in Python and a foundational understanding of machine learning principles.
Target Audience
- Observability engineers.
- Infrastructure and DevOps teams.
- Monitoring platform architects and Site Reliability Engineers (SREs).