I build autonomous AI agents for infrastructure — predictive failure detection, self-healing observability, and RAG-powered platform knowledge — across AWS, Azure, GCP & OCI.
I'm a Senior DevOps / Site Reliability & AI Engineer with 6+ years building and operating large-scale cloud platforms across AWS, Azure, GCP, and OCI.
My focus is Agentic DevOps — engineering autonomous AI agents for infrastructure operations, predictive failure detection, and intelligent observability using MCP servers, LangChain, and LangGraph. I design RAG pipelines, run LLM evaluation (RLHF, pass@k), and build embedding-based knowledge systems on top of a deep foundation in Kubernetes, DevSecOps, and Infrastructure-as-Code.
I ship production code in Python, Go, and TypeScript, and I care about systems that heal themselves before anyone gets paged.
An AI-engineering skill set layered on top of production-grade cloud & platform engineering.
Platforms and automation spanning agentic operations, security, observability, and MLOps.
Open to senior DevOps, SRE, Platform, and AI-engineering roles. Have an agentic infrastructure problem worth solving? Let's talk.