Search

Director Platform Engineering - SRE / Observability

PublishedPublished: 6/14/2022
Technology

Job Description

Director, Platform Engineering

\n


\n

SALARY: $175k - $240k plus 30% bonus

\n


\n

LOCATION: CHICAGO, IL

\n

HYBRID 3 DAYS ONSITE

\n


\n

You will manage a team of SREs and managers. 7-10 people and managers. Keys SRE observability infrastructure kubernetes kafka aws terraform. j

\n


\n

The Director will lead and manage this team to drive reliability, observability, and cloud platform engineering excellence across a large, complex cloud-based computing environment. The ideal candidate is a hands-on, data-driven technical leader, who can personally raise the bar on SRE and observability practices, reduce waste, optimize cloud efficiency, and improve performance, in close partnership with a dedicated SRE/monitoring team and a centralized architecture function.

\n

\n

Qualifications:

\n


\n

    \n
  • [Required] 5+ years of demonstrated experience leading engineering teams, with an emphasis on developing key talent and cultivating positive, high-performing cultures
  • \n

  • [Required] 10+ years of progressive, hands-on experience in software engineering with an understanding of large-scale computing solutions (primarily AWS), including software design and development, database architectures, IP networking, security, cloud operations, and performance tuning
  • \n

  • [Required] Demonstrated, hands-on expertise building and maturing SRE and observability practices at scale, able to personally raise the technical bar for a dedicated SRE/monitoring team, not just consume their output
  • \n

  • [Required] Demonstrated track record owning the definition and governance of SLOs and SLAs for mission-critical systems, and personally driving resilience engineering practices (chaos engineering, load/performance testing) to validate reliability ahead of incidents
  • \n

  • [Required] Demonstrated track record driving reliability through a data-driven culture, waste/toil reduction, cloud efficiency measures, and performance optimization, treating metrics as the primary lens for every decision
  • \n

  • [Required] Experience defining, instrumenting, and acting on software delivery performance metrics (DORA metrics, cycle time, deployment frequency, lead time, etc.) to drive engineering improvement initiatives
  • \n

  • [Required] Strong consultative, communication, team player, and analytical skills, with the ability to regularly interact between various teams distributed across the US
  • \n

  • [Required] Strong technical team leadership and technical project management skills
  • \n

  • [Required] Relevant experience leading highly technical team members through adopting new technologies while maintaining highly available, mission-critical systems, with a proven track record of success
  • \n

  • [Required] Ability to clearly communicate verbally and in writing to business and technology leaders, architects, developers, and team members
  • \n

  • [Required] Must be able to collaborate effectively with a group of high-performing, technical individuals
  • \n

  • [Required] Experience acting as a product owner, defining roadmap, requirements, and priorities, for a platform capability such as observability, ideally in partnership with a separate team that owns the underlying tooling and operations
  • \n

  • [Required] Experience with architecting, implementing, and maintaining highly available mission-critical environments for 24x7 availability
  • \n

  • [Required] Demonstrated history of working within deadlines and ability to work well under pressure
  • \n

  • [Required] Experience managing work tasks using Agile methodology/scrum desired
  • \n

\n


\n

Technical Skills:

\n

    \n
  • [Required] Deep expertise in OpenTelemetry, including instrumentation standards, auto-instrumentation, semantic conventions, and the OTel Collector, as the foundation for a vendor-neutral, paved-road instrumentation strategy
  • \n

  • [Required] Hands-on experience with metrics engines and time-series databases, including Prometheus and PromQL, plus scale-out options such as Mimir, Thanos, VictoriaMetrics, or Amazon Managed Prometheus, including cardinality management
  • \n

  • [Required] Hands-on experience with tracing and logging backends such as Tempo/Jaeger, Loki/Elastic/Splunk (Splunk is common in financial services), and AWS X-Ray
  • \n

  • [Required] Experience with Kubernetes and Kafka observability specifically, including EKS metrics, kube-state-metrics, consumer lag, and broker health
  • \n

  • [Required] Experience with alerting and incident tooling such as PagerDuty or Opsgenie, including ServiceNow integration, alert routing, and noise reduction
  • \n

  • [Required] Hands-on experience with SLO-as-code frameworks such as OpenSLO, Sloth, or Nobl9, defining and governing SLOs in the repo rather than only in a vendor UI
  • \n

  • [Required] Experience delivering golden paths and templates (Terraform modules, Helm charts, pipeline templates) that ship with logging, metrics, tracing, dashboards, alerts, and SLO defaults out of the box
  • \n

  • [Required] Experience with resilience validation practices, including chaos engineering (e.g., AWS FIS, Gremlin) and load/performance testing
  • \n

  • [Required] Deep, hands-on mastery of observability tooling and practices (metrics, distributed tracing, centralized logging, dashboards, alerting), e.g., Datadog, Prometheus/Grafana, CloudWatch, or equivalent, sufficient to elevate, not just consume, a dedicated SRE team’s capability
  • \n

  • [Required] Deep understanding of SRE principles including SLOs/SLIs, error budgets, incident management, and postmortem culture, with a track record of driving adoption and maturity
  • \n

  • [Required] Fluency in the Golden Signals, RED, and USE methods as applied frameworks for monitoring and alerting design
  • \n

  • [Required] Hands-on experience with: Terraform, Kubernetes, Jenkins or other CI/CD tooling, Kafka, Github, and configuration management tools such as Puppet, Chef, or Ansible
  • \n

  • [Required] Deep, hands-on expertise with infrastructure-as-code (IaC) tools and practices (e.g., Terraform, CloudFormation, CDK, Pulumi), with a track record of driving IaC adoption at scale across a cloud platform organization
  • \n

  • [Required] Relevant experience with configuration and implementation of IaaS, Infrastructure as Code, AWS, Azure, etc.
  • \n

  • [Required] Expert working knowledge of infrastructure design and components, such as servers, operating systems, networks, and storage
  • \n

  • [Required] Basic understanding of good delivery practices and continual integration and improvement; Agile/Lean background for projects and project delivery
  • \n

  • [Required] Bachelor’s degree, preferably in a technical discipline (Computer Science, Mathematics, etc.), or equivalent combination of education and experience required; Master’s degree and relevant experience also considered
  • \n

  • [Required] 10+ years’ experience in IT systems installation, operations, administration, and maintenance of cloud systems / virtualized servers, including 5+ years in a technical leadership role
  • \n

\n


\n

    \n
  • [Required] AWS Solutions Architect Associate Certification or higher strongly desired
  • \n

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...