Join Marcus Chen (Principal Platform Engineer · Databricks) for a free live session.

Marcus Chen
Principal Platform Engineer · Databricks
⭐ 4.9 / 5
In 6 weeks you'll write Terraform that provisions and fails over across AWS, Azure, and GCP, run chaos experiments with Gremlin and AWS FIS that actually break things, and ship a DR plan with RTO/RPO numbers you've proven — not estimated.
Design a DR architecture across AWS, Azure, and GCP using Terraform modules — with an ADR documenting trade-offs and failure modes per provider.
Derive RTO/RPO targets from real failure cost data and build a tiered DR plan — Backup/Restore through Active-Active — proven by simulation.
Run targeted compute, network, and dependency failure experiments with Gremlin and AWS FIS across all three clouds. Measure actual vs expected recovery.
Build a Grafana/Prometheus dashboard with SLO burn rate alerts and a quarterly resilience review format you run — not just reference.
This course is designed for:
Who manage multi-cloud environments but have separate codebases per provider with no shared DR strategy across them.
Who have a DR document no one has run — and aren't confident it would work if they needed it.
Who are accountable for uptime across AWS, Azure, and GCP but lack the observability and automated failover to back those commitments.
6 weeks · 3 sessions per week
Leave with real work to show, not just a certificate.
A detailed architecture document and Terraform codebase for a multi-cloud deployment, showcasing your ability to design resilient architectures across AWS, Azure, and GCP. This blueprint is a portfolio-ready artifact demonstrating your cross-cloud capabilities.
A comprehensive disaster recovery strategy document, including RTO/RPO definitions, failover procedures, and testing protocols. This plan is a valuable asset for showcasing your ability to manage and mitigate risks in multi-cloud environments.
A playbook detailing executed chaos experiments, findings, and resilience improvement strategies. This artifact demonstrates your practical expertise in chaos engineering and is suitable for showcasing your problem-solving skills in real-world scenarios.

Principal Platform Engineer · Databricks
⭐ 4.9 / 5
Marcus Chen is the Principal Platform Engineer for AI Infrastructure at Databricks, where he runs GPU cluster operations across 2,000+ nodes on AWS and Azure and owns the LLM inference platform serving production workloads. Before Databricks, he spent five years as a Senior SRE at Google Cloud. He teaches from the infra side — not the ML side.
⭐⭐⭐⭐⭐
"Our Week 4 chaos experiments revealed that our Azure failover had been misconfigured for months. We found it before a real outage did."
Avery Johnson
DevOps Engineer · Brex
⭐⭐⭐⭐⭐
"The cross-cloud Terraform module library replaced three separate codebases. We now provision the same stack on AWS and GCP from a single repo."
Riley Thompson
Cloud Engineer · Rippling
⭐⭐⭐⭐⭐
"First time we had RTO/RPO numbers backed by a tested plan. Our on-call rotation runs the DR playbook quarterly now."
Jordan Lee
SRE · Lattice
All sessions are instructor-led and live. Recordings available within 24 hours.
SUNDAY
9:00 AM PDT
Live ClassDeep dive into multi-cloud architecture patterns and Terraform script development.
WEDNESDAY
6:00 PM PDT
Lab SessionHands-on Terraform scripting and DR strategy design with instructor feedback.
THURSDAY
6:00 PM PDT
Build & ShipExecute chaos engineering experiments and automate failover processes.
The JD-backed research behind this course — from Dexity Intel.
The 2026 AI Infrastructure Stack: A Practical Guide →How far do you want to go?
Start free to experience our offering, choose the program length you would want to commit to.
You build. Nobody demos at you.
Every session is follow-along — you build the thing yourself while a practitioner works beside you. That is why the hours look long: they are yours to build in, with an expert on hand to guide you. None of it is a traditional lecture.
In the session, you'll: