Computer Pathshalaकंप्यूटर पाठशाला
← All courses
AI/ML OpsFoundationSoon · pre-enroll

AIOps Fundamentals

Observability + AI for incident response

Weeks
6
Lessons
28
Browser labs
6
Students
Rating
Included in Pro
₹1,499/mo
all 14 courses · or ₹14,990/yr · cancel anytime
Free preview · 2 lessons

Try before you pay.

Two full lessons from AIOps Fundamentals — exact topics, hands-on lab pairings, same depth as the paid course. Watch the videos free; sign up to access labs + the rest of the curriculum.

Video coming soon · subscribe on YouTube to be notified
Lesson 0119 min

What AIOps actually means (and what it doesn't)

AIOps is more often a vendor pitch than a working pattern. We separate the genuinely useful applications (anomaly detection, alert correlation, root-cause hints) from the marketing.

What this lesson teaches
  • · The three real AIOps wins: noise reduction, anomaly detection, RCA hints
  • · Anti-patterns: 'AI-driven self-healing', 'predictive maintenance' for stateless web apps
  • · The data prerequisite: structured logs + metrics + traces with consistent IDs
  • · Build vs buy: when Datadog Watchdog / Splunk ITSI / New Relic AI is enough
Sign up free to access lab + sandbox →
Video coming soon · subscribe on YouTube to be notified
Lesson 0223 min

Anomaly detection on metrics — the practical version

Hands-on walkthrough of building a working anomaly detector on real Prometheus metrics. We use Prophet + a simple confidence-interval threshold and compare to Datadog's offering.

What this lesson teaches
  • · Time-series decomposition: trend + seasonality + residual
  • · Prophet vs ARIMA vs simple z-score for ops metrics
  • · Alert tuning: balancing precision vs recall (and why recall usually wins ops)
  • · Wiring anomaly alerts into PagerDuty / Slack with severity-aware routing
  • · Avoiding the 'alert fatigue → ignored alerts → outage' loop
● Paired lab

Train a Prophet model on Prometheus latency metrics from a sample app and trigger Slack alerts.

Sign up free to access lab + sandbox →

These are 2 of 28 lessons. Subscribe to @PathshalaDevOps for new lessons + course launches. The full 26 remaining lessons are included with cohort enrolment, with a 7-day money-back guarantee.

What you’ll build

10 capstones. Reviewed by senior engineers.

01Build a VPC + EC2 from scratch● Lab
02Containerize a Node app & push to ECR● Lab
03Deploy to EKS with Helm● Lab
04Terraform a 3-tier app● Lab
05GitHub Actions: build → push → deploy● Lab
06Blue-green deploy with Route53● Lab
07Set up Prometheus + Grafana on EKS● Lab
08SLO-based alerting● Lab
09Chaos test with AWS FIS● Lab
10Cost-optimize an EC2 fleet● Lab
Curriculum

28 lessons across 6 weeks

Week 01AIOps fundamentals4 lessons
Week 02Log signal mining4 lessons
  • 05Structured logs as a precondition for AIOpsvideo18m
  • 06Pattern extraction with Drain / LogPAIvideo22m
  • 07Wiring log anomalies into the alerting planevideo20m
  • 08Lab: extract templates + alert on rare-pattern emergencelab90m
Week 03Alert correlation + de-duplication4 lessons
  • 09Why one root cause fires 50 alerts (and what to do)video20m
  • 10AlertManager: grouping, inhibition, routing treesvideo24m
  • 11Graph-based clustering with service-dependency mapsvideo22m
  • 12Lab: collapse 50 alerts into 1 incident with AlertManagerlab90m
Week 04Root-cause hints from traces4 lessons
  • 13Distributed traces — the foundation for RCA hintsvideo20m
  • 14Critical-path analysis on a slow requestvideo22m
  • 15Correlating deploys with metric shiftsvideo20m
  • 16Lab: build a 'what changed' panel using Tempo + Grafanalab90m
Week 05Commercial AIOps comparison5 lessons
  • 17Datadog Watchdog under the hoodvideo18m
  • 18Splunk ITSI: when it shines, when it doesn'tvideo18m
  • 19New Relic AI + Dynatrace Davis — the contendersvideo20m
  • 20Build vs buy decision frameworkvideo22m
  • 21Lab: head-to-head bake-off on a sample incidentlab90m
Week 06Capstone — incident response with AIOps4 lessons
  • 22Pre-incident: signal hygiene + readiness checklistsvideo20m
  • 23During an incident: AIOps as a sidekick, not the decidervideo22m
  • 24Post-incident: feeding the model back from learningsvideo18m
  • 25Capstone lab: respond to a multi-service outage end-to-endlab150m
More from AI/ML Ops

Pair it with