Amit Kumar, Microsoft site reliability, platform and cloud engineer

Amit Kumar · Microsoft · Site Reliability · Platform Engineering

Amit Kumar engineers systems that stay reliable.

I’m a Microsoft engineer working across Azure production reliability, live-site operations, incident engineering, automation, and AI-assisted reliability. I build systems that help engineers detect risk earlier, investigate production issues faster, and turn operational knowledge into scalable workflows.

Status OperationalFocus Azure SREMode Build / LearnProfile visit counter

9+ years · Scroll to explore ↓

Microsoft · Azure · SRE · AI for operations

Turning production signals, incidents, and telemetry into reliability at scale.

Signal → context → decision

From reactive response to proactive intelligence.

I use AI, telemetry, APIs, automation, and historical incident knowledge to help engineers see risk earlier and act with better context.

Scroll film

Quick profile

Depth meets direction.

Deep enterprise cloud experience across infrastructure, Microsoft 365, identity, security, production reliability, and customer-facing engineering.

Experience

9+years in technology and cloud

Current scope

Azurereliability and cloud operations

Engineering

AIautomation and agentic workflows

Recognition

Microsoft ACE Award

Education

MBAIIM Kozhikode

Platform

M365enterprise cloud expertise

About

Engineering reliability beyond incident response.

I’m a technology professional with 9+ years of experience spanning Azure, Microsoft 365, cloud reliability, identity, security, automation, enterprise SaaS, and production engineering.

At Microsoft, my work focuses on Azure live-site reliability and engineering operations: production incidents, telemetry, cross-incident analysis, operational reviews, problem management, and automation.

I’m helping move reliability engineering from reactive response toward proactive intelligence by applying AI, APIs, telemetry, and agentic workflows to operational signals and historical knowledge.

01
Automate what humans shouldn’t repeat.

Every repetitive operational process is an engineering opportunity.

02
Engineer for failure.

Production systems should expect failures and recover intelligently.

03
Make reliability measurable.

Observability matters when it leads to decisions and preventative action.

What I work on

The operating system behind reliability.

Engineering practices, platforms, and automation for complex cloud environments.

01

Site Reliability Engineering

  • Production incident management
  • Live-site operations and mitigation
  • Root-cause and cross-incident analysis
  • Post-incident learning
  • Repair effectiveness
  • Operational readiness
02

Platform Engineering

  • Cloud platform operations
  • Operational tooling
  • API integrations
  • Service telemetry
  • Standardized workflows
  • Self-service capabilities
03

DevOps & Automation

  • Python and PowerShell
  • CI/CD and GitHub Actions
  • Azure DevOps
  • KQL and data-driven troubleshooting
  • Azure Functions and Logic Apps
  • Workflow automation
Signal streams24 active
Correlation0.92 confidence
Decision modelHuman in loop
System stateLearning

Illustrative operational model · Live

Every incident leaves a signal.

The engineering opportunity is to connect those signals across time, services, telemetry, and repairs—then turn them into earlier, better decisions.

Reliability intelligence / conceptual visualization

AI × Reliability Engineering

Bringing intelligence into production operations.

AI should augment engineering judgment, not replace engineering ownership in critical production environments.

Pre-Outage Intelligence

Using operational signals and historical patterns to surface potential reliability risks before customer impact expands.

AI-Assisted Incident Analysis

Correlating telemetry, service context, incident history, and operational signals for faster engineering decisions.

Multi-Agent Workflows

Specialized agents collaborating across retrieval, investigation, analysis, and recommendation generation.

Operational Knowledge Automation

Transforming historical incident knowledge into systems that assist engineers during live-site events.

API + AI Integration

Connecting operational data sources with intelligent workflows to reduce manual investigation.

Human in the Loop

Keeping accountable engineering judgment at the center of AI-assisted production operations.

Featured engineering work · Drag to explore

Systems for seeing, learning, and acting earlier.

01 · AI + Reliability

AI-Assisted Reliability Intelligence

Evaluating operational signals and historical context to identify potential production risks earlier.

Agentic AIAzureAPIsReliability

02 · SRE

Cross-Incident Reliability Engineering

Finding recurring failure modes, detection gaps, mitigation delays, dependencies, and repair effectiveness across incidents.

SREProblem ManagementRCAObservability

03 · DevOps

Cloud Operations Automation

Reducing manual investigation and operational effort through API-driven automation and reusable engineering workflows.

PythonPowerShellAPIsDevOps

04 · Enterprise Cloud

Microsoft 365 & Identity Engineering

Deep enterprise experience across identity, messaging, security, compliance, and cloud administration.

Microsoft 365GraphEntra ID

Engineering stack

Tools are useful. Systems thinking is the multiplier.

A reliability-first toolkit spanning production operations, cloud platforms, automation, observability, AI, and enterprise cloud.

Reliability

SRELive-Site OperationsIncident EngineeringProblem ManagementRCAObservability

Cloud / Platform

Microsoft AzureCloud InfrastructurePlatform EngineeringAzure MonitorAzure Functions

DevOps / Automation

PythonPowerShellGitHub ActionsAzure DevOpsCI/CDREST APIs

AI Engineering

Agentic AIMulti-Agent SystemsAI-Assisted OperationsOperational IntelligenceLLM Workflows

Enterprise Cloud

Microsoft 365Entra IDExchange OnlineMicrosoft PurviewMicrosoft Defender

Engineering Data

KQLTelemetry AnalysisJSONMicrosoft GraphGit

Experience

From software to AI-augmented reliability.

A career built across the application, enterprise cloud, platform, and production layers.

Microsoft — Azure Reliability & Engineering Operations

Current focus: Azure production reliability, live-site engineering, problem management, operational intelligence, and AI-assisted reliability.

  • Engineer across large-scale Azure production environments and live-site operations
  • Analyze incidents and telemetry to identify recurring failure patterns and systemic risks
  • Drive cross-incident problem management and preventative engineering opportunities
  • Contribute to reliability-review tooling and operational intelligence workflows
  • Build AI-assisted capabilities combining operational data, APIs, automation, and agents
  • Apply structured incident learning to improve detection, mitigation, and repair effectiveness

Earlier: Microsoft 365 / Identity / Security Engineering across Exchange Online, Entra ID, Graph, Defender, Purview, and complex production escalations.

Navisite

Microsoft 365 Engineer across identity, messaging, security, automation, and platform administration.

Wipro

Microsoft 365 Technical Consultant focused on Exchange Hybrid, migrations, disaster recovery, and PowerShell.

Optimal Transnational

Software Engineer building and supporting web applications and production systems.

SoftwareEnterprise CloudAzure ProductionSRE + PlatformAI Reliability

Credentials & recognition

Engineering depth, business range.

Education & certifications

Executive MBAIIM Kozhikode · 2023–2025
AZ-104Azure Administrator
MS-500M365 Security
PL-900Power Platform

Recognition

2× ACE AwardMicrosoft
Star PerformerNavisite
Star AchieverWipro
Codex AwardOptimal Transnational

Engineering in public · @amitkumarops

Build. Document. Share the system.

My GitHub is where I build around Python, DevOps automation, Azure, platform engineering, SRE, observability, infrastructure automation, AI agents, and reliability tooling.

PythonAzureDevOpsAI agentsReliability tooling

SRE · Platform · DevOps · Azure · AI

Let’s build reliable systems.

Interested in technically challenging problems across cloud platforms, automation, and intelligent operations.