Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Foundations of Agentic AI in IT Operations
- The shift from static runbooks to reasoning agents: the evolution of IT automation
- Agent components: reasoning loops, tool utilization, memory management, and planning strategies
- Deciding when to automate tasks versus when human oversight is essential
Agent Frameworks and System Architectures
- Single-agent methodologies: ReAct, Plan-and-Execute, and tool-calling loops
- Multi-agent structures: supervisor, hierarchical, and swarm models
- Comparing frameworks: LangGraph, CrewAI, AutoGen, and custom agent solutions
- Constructing your initial operational agent: querying monitoring systems, diagnosing issues, and proposing solutions
Integrating Tools for IT Operations
- Linking agents to Prometheus, Grafana, Datadog, and PagerDuty interfaces
- Agent-driven log analysis: integrating Elasticsearch, Loki, and Splunk
- Leveraging infrastructure tools: utilizing kubectl, Terraform, and Ansible through agent actions
- Creating secure tool interfaces with parameter validation and idempotency
Automating Incident Response
- Automated incident triage: classifying severity and routing alerts
- Generating root cause hypotheses and collecting supporting evidence
- Automated remediation: executing restart, scaling, rollback, and failover actions
- Developing an incident runbook agent with progressive autonomy levels
Safety, Guardrails, and Human-in-the-Loop
- Categorizing actions: read-only, low-risk, high-risk, and destructive operations
- Establishing approval gates and escalation policies for critical tasks
- Implementing guardrail patterns: action allowlists, blast radius constraints, and rollback assurances
- Maintaining audit trails and decision provenance for compliance
Orchestrating Multi-Agent Systems for Complex Incidents
- Coordinating specialized agents: triage, diagnosis, and remediation agents
- Managing inter-agent communication and shared context
- Resolving conflicts when agents propose opposing actions
- Simulating end-to-end major incidents with multi-agent responses
Observability and Performance Evaluation
- Tracing agent reasoning chains for debugging and auditing purposes
- Assessing agent decision quality: precision, recall, and time-to-resolution metrics
- Creating feedback loops: learning from operator overrides and final outcomes
- Monitoring costs and token economics for operational agents
Production Deployment and Operational Management
- Deploying agents as services: using APIs, webhooks, and scheduled jobs
- Implementing gradual autonomy rollout: transitioning from shadow mode to full auto-remediation
- Handling agent failures: procedures for when the agent itself encounters issues
- Building the business case and measuring ROI for autonomous operations
Requirements
- Experience with IT operations, DevOps, or SRE practices.
- Proficiency with Python scripting and REST APIs.
- A foundational understanding of LLM capabilities and prompt engineering.
Target Audience
- SRE and DevOps engineers exploring AI-driven automation.
- Platform engineers developing self-healing infrastructure.
- IT operations leaders evaluating agentic AI for incident management.
14 Hours