Search

Senior Reliability Engineer

PublishedPublished: 6/14/2022
Technology

Job Description

Job Description

Reliability Engineer – Azure / Terraform / Kubernetes

We are seeking an experienced Reliability Engineer to ensure the reliability, scalability, and operational health of our Azure cloud environment.

Terraform is central to this position. The ideal candidate will independently design, build, review, maintain, and troubleshoot Infrastructure-as-Code (IaC) for mission-critical environments. This role will collaborate closely with platform engineers, developers, security teams, and product stakeholders to automate cloud infrastructure, improve Kubernetes operations, and resolve issues affecting the availability and performance of data and analytics services.


Key Responsibilities

  • Design, implement, maintain, and troubleshoot production Azure infrastructure using Terraform.
  • Support the reliability, performance, scalability, and availability of workloads running in Azure Kubernetes Service (AKS).
  • Troubleshoot cloud infrastructure, networking, Kubernetes, and application reliability issues.
  • Automate cloud operations to reduce manual effort and improve consistency.
  • Collaborate with development and operations teams to improve deployment and incident-response practices.
  • Implement and enhance monitoring, alerting, dashboards, and operational reporting.
  • Identify reliability risks and recommend improvements to cloud architecture, infrastructure, and operational processes.
  • Create and maintain infrastructure documentation, procedures, troubleshooting guides, and operational runbooks.


Required Qualifications

  • 4+ years of experience in cloud infrastructure, systems engineering, DevOps, Site Reliability Engineering (SRE), or a similar role.
  • 2+ years of hands-on Microsoft Azure experience.
  • 2+ years of hands-on Terraform experience, including designing reusable modules, managing state, troubleshooting failures, and maintaining production infrastructure.
  • Experience administering and supporting Kubernetes environments; Azure Kubernetes Service (AKS) experience is preferred.
  • Experience supporting cloud-hosted systems and automating cloud operations.
  • Strong knowledge of cloud networking concepts and troubleshooting.
  • Strong analytical and problem-solving skills.
  • Ability to work on-site in Atlanta, GA, or the Washington, DC metro area.
  • Ability to obtain and maintain a Public Trust/Suitability clearance.
  • Bachelor’s degree or equivalent professional experience.


Preferred Qualifications

  • Experience with certificate lifecycle management, including maintaining, renewing, and rotating certificates.
  • Experience creating operational dashboards and reports; Power BI experience is preferred.
  • Experience with observability platforms such as Grafana, Prometheus, Elastic, or Splunk.
  • Experience integrating Azure resources with Active Directory.
  • Experience using AI-assisted or agentic coding tools such as Claude Code, Codex, or GitHub Copilot for application development or infrastructure automation.
  • Azure, Kubernetes, or Terraform certifications are a plus.
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...