Get in Touch

Course Outline

Foundations of Agentic AI in IT Operations

  • The shift from static runbooks to reasoning agents: the evolution of IT automation
  • Agent components: reasoning loops, tool utilization, memory management, and planning strategies
  • Deciding when to automate tasks versus when human oversight is essential

Agent Frameworks and System Architectures

  • Single-agent methodologies: ReAct, Plan-and-Execute, and tool-calling loops
  • Multi-agent structures: supervisor, hierarchical, and swarm models
  • Comparing frameworks: LangGraph, CrewAI, AutoGen, and custom agent solutions
  • Constructing your initial operational agent: querying monitoring systems, diagnosing issues, and proposing solutions

Integrating Tools for IT Operations

  • Linking agents to Prometheus, Grafana, Datadog, and PagerDuty interfaces
  • Agent-driven log analysis: integrating Elasticsearch, Loki, and Splunk
  • Leveraging infrastructure tools: utilizing kubectl, Terraform, and Ansible through agent actions
  • Creating secure tool interfaces with parameter validation and idempotency

Automating Incident Response

  • Automated incident triage: classifying severity and routing alerts
  • Generating root cause hypotheses and collecting supporting evidence
  • Automated remediation: executing restart, scaling, rollback, and failover actions
  • Developing an incident runbook agent with progressive autonomy levels

Safety, Guardrails, and Human-in-the-Loop

  • Categorizing actions: read-only, low-risk, high-risk, and destructive operations
  • Establishing approval gates and escalation policies for critical tasks
  • Implementing guardrail patterns: action allowlists, blast radius constraints, and rollback assurances
  • Maintaining audit trails and decision provenance for compliance

Orchestrating Multi-Agent Systems for Complex Incidents

  • Coordinating specialized agents: triage, diagnosis, and remediation agents
  • Managing inter-agent communication and shared context
  • Resolving conflicts when agents propose opposing actions
  • Simulating end-to-end major incidents with multi-agent responses

Observability and Performance Evaluation

  • Tracing agent reasoning chains for debugging and auditing purposes
  • Assessing agent decision quality: precision, recall, and time-to-resolution metrics
  • Creating feedback loops: learning from operator overrides and final outcomes
  • Monitoring costs and token economics for operational agents

Production Deployment and Operational Management

  • Deploying agents as services: using APIs, webhooks, and scheduled jobs
  • Implementing gradual autonomy rollout: transitioning from shadow mode to full auto-remediation
  • Handling agent failures: procedures for when the agent itself encounters issues
  • Building the business case and measuring ROI for autonomous operations

Requirements

  • Experience with IT operations, DevOps, or SRE practices.
  • Proficiency with Python scripting and REST APIs.
  • A foundational understanding of LLM capabilities and prompt engineering.

Target Audience

  • SRE and DevOps engineers exploring AI-driven automation.
  • Platform engineers developing self-healing infrastructure.
  • IT operations leaders evaluating agentic AI for incident management.
 14 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories