Search

Site Reliability Engineer

PublishedPublished: 6/14/2022
Technology

Job Description

Role: Site Reliability Engineer with strong experience in AWS, Kubernetes and Terraform

\n

Location: Louisville, KY (Hybrid) (on site Tuesday and Thursday)

\n

Contract

\n


\n

Job Description

\n

Key Responsibilities

\n

Reliability & Observability

\n

    \n
  • Define and own Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for restaurant technology systems; use error budgets to balance reliability with velocity.
  • \n

  • Build and maintain monitoring, alerting, and observability platforms that provide meaningful signal — not noise.
  • \n

  • Lead blameless post-incident reviews; drive root cause analysis and ensure permanent corrective actions are implemented.
  • \n

  • Proactively identify reliability risks across the stack before they become incidents.
  • \n

  • Establish and track reliability metrics; report on system health to engineering leadership.
  • \n

\n

Architecture & Infrastructure

\n

    \n
  • Design and implement scalable, secure, and highly available cloud architectures primarily on AWS, with working knowledge of Azure.
  • \n

  • Architect and manage containerized workloads using Kubernetes (EKS, AKS), including edge Kubernetes deployments in restaurant environments.
  • \n

  • Design and implement serverless and container-based solutions, including AWS Fargate and other managed services.
  • \n

  • Develop and maintain Infrastructure as Code (IaC) using Terraform.
  • \n

  • Own architecture across the full restaurant technology stack — cloud, edge, networking, and device management — not just the cloud layer.
  • \n

\n

DevOps, Automation & Tooling

\n

    \n
  • Build and optimize CI/CD pipelines using GitLab CI/CD and modern DevOps practices.
  • \n

  • Build internal tools, automations, and middleware integrations that eliminate repetitive operational work.
  • \n

  • Use AI-assisted development to accelerate scripting, troubleshooting, and documentation.
  • \n

  • Champion a culture of engineering solutions over repeated manual fixes — if something is done twice, it should be automated.
  • \n

\n

Restaurant & Edge Technology

\n

    \n
  • Design and support Kubernetes-based edge systems deployed in restaurant locations.
  • \n

  • Support mobile application deployments and troubleshoot deployment issues across restaurant endpoints.
  • \n

  • Manage and optimize Mobile Device Management (MDM) platforms covering the restaurant device fleet.
  • \n

  • Configure and troubleshoot enterprise networking — primarily switches and restaurant-facing network infrastructure.
  • \n

  • Lead and participate in incident response for restaurant technology systems, including on-call coverage and post-incident review.
  • \n

  • Reduce mean time to detection (MTTD) and mean time to resolution (MTTR) through better tooling, runbooks, and automation.
  • \n

\n

Security & Governance

\n

    \n
  • Establish and enforce cloud governance, security policies, and architectural standards.
  • \n

  • Implement cloud security best practices: IAM strategy, network segmentation, encryption, and secrets management.
  • \n

  • Conduct security architecture reviews; identify vulnerabilities, misconfigurations, and compliance gaps.
  • \n

  • Integrate security into CI/CD pipelines (DevSecOps — Development, Security, and Operations), including automated scanning, policy validation, and vulnerability management.
  • \n

\n

Leadership & Collaboration

\n

    \n
  • Collaborate with engineering, DevOps, and security teams to ensure secure-by-design solutions across cloud and restaurant tech.
  • \n

  • Provide technical leadership and mentorship to engineering teams.
  • \n

  • Create and maintain documentation, runbooks, and architectural decision records.
  • \n

  • Continuously evaluate emerging technologies and recommend improvements.
  • \n

\n

Required Qualifications

\n

    \n
  • 6+ years in IT infrastructure, with 3+ years focused on site reliability engineering, cloud architecture, or platform engineering.
  • \n

  • Hands-on experience with AWS (VPC, EC2, ECS, EKS, Fargate, Lambda, IAM, RDS, S3).
  • \n

  • Working experience with Microsoft Azure.
  • \n

  • Strong expertise in Kubernetes and container orchestration, including edge or distributed deployments.
  • \n

  • Experience with GitLab CI/CD and CI/CD pipeline design.
  • \n

  • Solid experience with Terraform for infrastructure provisioning.
  • \n

  • Experience with enterprise networking — switch configuration, VLANs, network troubleshooting.
  • \n

  • Familiarity with Mobile Device Management (MDM) platforms.
  • \n

  • Experience with automation and scripting (Python, Bash, Go, or equivalent).
  • \n

  • Proven ability to build internal tooling and API integrations, not just configure managed services.
  • \n

  • Experience defining and operating against SLOs, SLIs, and error budgets.
  • \n

  • Comfortable working in Linux command-line environments; Windows familiarity a plus where restaurant endpoints require it.
  • \n

\n

Preferred Qualifications

\n

    \n
  • Experience designing serverless architectures (AWS Lambda, Fargate, API Gateway, EventBridge).
  • \n

  • Experience with DevSecOps tooling (SAST — Static Application Security Testing, DAST — Dynamic Application Security Testing, container scanning, IaC scanning).
  • \n

  • Familiarity with security frameworks (CIS, NIST, ISO 27001, SOC 2).
  • \n

  • AWS and/or Azure certifications.
  • \n

  • Experience with monitoring and observability tools (CloudWatch, Prometheus, Grafana, or SIEM solutions).
  • \n

  • Background in restaurant, retail, or distributed edge technology environments.
  • \n

  • Experience using AI-assisted development tools for scripting, troubleshooting, and documentation.
  • \n

\n

Key Competencies

\n

    \n
  • Generalist mindset — comfortable moving between cloud, edge, networking, and device management in the same week.
  • \n

  • Reliability-first thinking — treats toil reduction, error budgets, and post-incident learning as core engineering disciplines, not afterthoughts.
  • \n

  • Bias toward permanent fixes and automation over repeated manual intervention.
  • \n

  • Ability to balance strategic architecture with hands-on execution.
  • \n

  • Strong communication and cross-functional collaboration skills.
  • \n

  • Curious, proactive, and detail-oriented approach to systems design and operations.
  • \n

\n


Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...