Site Reliability Engineer (SRE) - W2 - USC OR GC
Technology
Job Description
Job Title: Site Reliability Engineer (SRE)
\n
Location: Kansas City, MO (Onsite)
\n
GC or USC ONLY
\n
\n
Essential Job Functions
\n
- \n
- Deploy, configure, and manage Azure SRE Agent across production workloads to automate incident investigation, root cause analysis, and remediation workflows
- Monitor production systems and respond to incidents in partnership with Azure SRE Agent, reviewing agent-proposed diagnoses and mitigations and approving actions per governance policy
- Build and maintain runbooks, subagents, and agent hooks within Azure SRE Agent to automate common operational tasks and reduce manual toil
- Connect Azure SRE Agent to observability, incident management, and source control tooling (Azure Monitor, Application Insights, Log Analytics, GitHub, PagerDuty, or similar) to enable end-to-end automated investigations
- Define and maintain service level objectives (SLOs), service level indicators (SLIs), and error budgets for critical services, and use agent-driven insights to track and improve them
- Establish and enforce tool permissions, hooks, and governance controls for AI agent actions to ensure safe, auditable automation with appropriate human approval gates
- Analyze incident trends and agent-surfaced institutional knowledge to identify recurring issues and drive permanent fixes and reliability improvements
- Participate in on-call rotation, leveraging Azure SRE Agent to accelerate triage, reduce mean time to detect/resolve (MTTD/MTTR), and minimize after-hours disruptions
- Author and refine incident response plans, playbooks, and postmortem processes, incorporating agent-generated documentation and learnings
- Partner with DevOps, Platform Engineering, and development teams to instrument applications and infrastructure for effective AI-driven monitoring and diagnostics
- Continuously evaluate new Azure SRE Agent capabilities (connectors, private plugins, subagents) and pilot adoption to expand automated coverage
- Contribute to Infrastructure as Code and automation scripts that support reliable, repeatable, and agent-manageable environments
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
Experience in the following areas is required:
\n
- \n
- Bachelor’s degree in computer science, Information Technology, or a related field required, or equivalent professional experience
- 5+ years of experience in Site Reliability Engineering, DevOps, or a related production operations role
- Hands-on experience with Azure SRE Agent or comparable AI-powered incident response/observability tooling strongly preferred
- Strong working knowledge of Azure cloud services, including AKS, App Service, Azure Functions, and Azure networking
- Experience with observability and monitoring platforms (Azure Monitor, Application Insights, Log Analytics, or similar)
- Practical understanding of SRE fundamentals, including SLOs/SLIs, error budgets, incident management, and blameless postmortems
- Familiarity with CI/CD tooling such as Azure DevOps Pipelines and GitHub Actions
- Experience with Infrastructure as Code tools such as Bicep or Terraform
- Comfort working with AI agents and automation frameworks, including reviewing and approving AI-proposed remediations under governance controls
- Strong scripting and automation skills (PowerShell, Bash, or Python)
- Excellent troubleshooting skills and the ability to remain calm and methodical during production incidents
- Strong communication skills, with the ability to document findings clearly for both technical and non-technical audiences
- Willingness to participate in an on-call rotation.
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
