Site Reliability Engineering
- Production incident management
- Live-site operations and mitigation
- Root-cause and cross-incident analysis
- Post-incident learning
- Repair effectiveness
- Operational readiness

Amit Kumar · Microsoft · Site Reliability · Platform Engineering
I’m a Microsoft engineer working across Azure production reliability, live-site operations, incident engineering, automation, and AI-assisted reliability. I build systems that help engineers detect risk earlier, investigate production issues faster, and turn operational knowledge into scalable workflows.
9+ years · Scroll to explore ↓
Microsoft · Azure · SRE · AI for operations
Signal → context → decision
I use AI, telemetry, APIs, automation, and historical incident knowledge to help engineers see risk earlier and act with better context.
Quick profile
Deep enterprise cloud experience across infrastructure, Microsoft 365, identity, security, production reliability, and customer-facing engineering.
Experience
9+years in technology and cloudCurrent scope
Azurereliability and cloud operationsEngineering
AIautomation and agentic workflowsRecognition
2×Microsoft ACE AwardEducation
MBAIIM KozhikodePlatform
M365enterprise cloud expertiseAbout
I’m a technology professional with 9+ years of experience spanning Azure, Microsoft 365, cloud reliability, identity, security, automation, enterprise SaaS, and production engineering.
At Microsoft, my work focuses on Azure live-site reliability and engineering operations: production incidents, telemetry, cross-incident analysis, operational reviews, problem management, and automation.
I’m helping move reliability engineering from reactive response toward proactive intelligence by applying AI, APIs, telemetry, and agentic workflows to operational signals and historical knowledge.
Every repetitive operational process is an engineering opportunity.
Production systems should expect failures and recover intelligently.
Observability matters when it leads to decisions and preventative action.
What I work on
Engineering practices, platforms, and automation for complex cloud environments.
Illustrative operational model · Live
The engineering opportunity is to connect those signals across time, services, telemetry, and repairs—then turn them into earlier, better decisions.
AI × Reliability Engineering
AI should augment engineering judgment, not replace engineering ownership in critical production environments.
Using operational signals and historical patterns to surface potential reliability risks before customer impact expands.
Correlating telemetry, service context, incident history, and operational signals for faster engineering decisions.
Specialized agents collaborating across retrieval, investigation, analysis, and recommendation generation.
Transforming historical incident knowledge into systems that assist engineers during live-site events.
Connecting operational data sources with intelligent workflows to reduce manual investigation.
Keeping accountable engineering judgment at the center of AI-assisted production operations.
Featured engineering work · Drag to explore
01 · AI + Reliability
Evaluating operational signals and historical context to identify potential production risks earlier.
02 · SRE
Finding recurring failure modes, detection gaps, mitigation delays, dependencies, and repair effectiveness across incidents.
03 · DevOps
Reducing manual investigation and operational effort through API-driven automation and reusable engineering workflows.
04 · Enterprise Cloud
Deep enterprise experience across identity, messaging, security, compliance, and cloud administration.
Engineering stack
A reliability-first toolkit spanning production operations, cloud platforms, automation, observability, AI, and enterprise cloud.
Experience
A career built across the application, enterprise cloud, platform, and production layers.
Current focus: Azure production reliability, live-site engineering, problem management, operational intelligence, and AI-assisted reliability.
Earlier: Microsoft 365 / Identity / Security Engineering across Exchange Online, Entra ID, Graph, Defender, Purview, and complex production escalations.
Microsoft 365 Engineer across identity, messaging, security, automation, and platform administration.
Microsoft 365 Technical Consultant focused on Exchange Hybrid, migrations, disaster recovery, and PowerShell.
Software Engineer building and supporting web applications and production systems.
Credentials & recognition
Education & certifications
Recognition
Engineering in public · @amitkumarops
My GitHub is where I build around Python, DevOps automation, Azure, platform engineering, SRE, observability, infrastructure automation, AI agents, and reliability tooling.
SRE · Platform · DevOps · Azure · AI
Interested in technically challenging problems across cloud platforms, automation, and intelligent operations.