Director Platform Engineering - SRE / Observability
Job Description
Director, Platform Engineering
\n
\n
SALARY: $175k - $240k plus 30% bonus
\n
\n
LOCATION: CHICAGO, IL
\n
HYBRID 3 DAYS ONSITE
\n
\n
You will manage a team of SREs and managers. 7-10 people and managers. Keys SRE observability infrastructure kubernetes kafka aws terraform. j
\n
\n
The Director will lead and manage this team to drive reliability, observability, and cloud platform engineering excellence across a large, complex cloud-based computing environment. The ideal candidate is a hands-on, data-driven technical leader, who can personally raise the bar on SRE and observability practices, reduce waste, optimize cloud efficiency, and improve performance, in close partnership with a dedicated SRE/monitoring team and a centralized architecture function.
\n
\n
Qualifications:
\n
\n
- \n
- [Required] 5+ years of demonstrated experience leading engineering teams, with an emphasis on developing key talent and cultivating positive, high-performing cultures
- [Required] 10+ years of progressive, hands-on experience in software engineering with an understanding of large-scale computing solutions (primarily AWS), including software design and development, database architectures, IP networking, security, cloud operations, and performance tuning
- [Required] Demonstrated, hands-on expertise building and maturing SRE and observability practices at scale, able to personally raise the technical bar for a dedicated SRE/monitoring team, not just consume their output
- [Required] Demonstrated track record owning the definition and governance of SLOs and SLAs for mission-critical systems, and personally driving resilience engineering practices (chaos engineering, load/performance testing) to validate reliability ahead of incidents
- [Required] Demonstrated track record driving reliability through a data-driven culture, waste/toil reduction, cloud efficiency measures, and performance optimization, treating metrics as the primary lens for every decision
- [Required] Experience defining, instrumenting, and acting on software delivery performance metrics (DORA metrics, cycle time, deployment frequency, lead time, etc.) to drive engineering improvement initiatives
- [Required] Strong consultative, communication, team player, and analytical skills, with the ability to regularly interact between various teams distributed across the US
- [Required] Strong technical team leadership and technical project management skills
- [Required] Relevant experience leading highly technical team members through adopting new technologies while maintaining highly available, mission-critical systems, with a proven track record of success
- [Required] Ability to clearly communicate verbally and in writing to business and technology leaders, architects, developers, and team members
- [Required] Must be able to collaborate effectively with a group of high-performing, technical individuals
- [Required] Experience acting as a product owner, defining roadmap, requirements, and priorities, for a platform capability such as observability, ideally in partnership with a separate team that owns the underlying tooling and operations
- [Required] Experience with architecting, implementing, and maintaining highly available mission-critical environments for 24x7 availability
- [Required] Demonstrated history of working within deadlines and ability to work well under pressure
- [Required] Experience managing work tasks using Agile methodology/scrum desired
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
Technical Skills:
\n
- \n
- [Required] Deep expertise in OpenTelemetry, including instrumentation standards, auto-instrumentation, semantic conventions, and the OTel Collector, as the foundation for a vendor-neutral, paved-road instrumentation strategy
- [Required] Hands-on experience with metrics engines and time-series databases, including Prometheus and PromQL, plus scale-out options such as Mimir, Thanos, VictoriaMetrics, or Amazon Managed Prometheus, including cardinality management
- [Required] Hands-on experience with tracing and logging backends such as Tempo/Jaeger, Loki/Elastic/Splunk (Splunk is common in financial services), and AWS X-Ray
- [Required] Experience with Kubernetes and Kafka observability specifically, including EKS metrics, kube-state-metrics, consumer lag, and broker health
- [Required] Experience with alerting and incident tooling such as PagerDuty or Opsgenie, including ServiceNow integration, alert routing, and noise reduction
- [Required] Hands-on experience with SLO-as-code frameworks such as OpenSLO, Sloth, or Nobl9, defining and governing SLOs in the repo rather than only in a vendor UI
- [Required] Experience delivering golden paths and templates (Terraform modules, Helm charts, pipeline templates) that ship with logging, metrics, tracing, dashboards, alerts, and SLO defaults out of the box
- [Required] Experience with resilience validation practices, including chaos engineering (e.g., AWS FIS, Gremlin) and load/performance testing
- [Required] Deep, hands-on mastery of observability tooling and practices (metrics, distributed tracing, centralized logging, dashboards, alerting), e.g., Datadog, Prometheus/Grafana, CloudWatch, or equivalent, sufficient to elevate, not just consume, a dedicated SRE team’s capability
- [Required] Deep understanding of SRE principles including SLOs/SLIs, error budgets, incident management, and postmortem culture, with a track record of driving adoption and maturity
- [Required] Fluency in the Golden Signals, RED, and USE methods as applied frameworks for monitoring and alerting design
- [Required] Hands-on experience with: Terraform, Kubernetes, Jenkins or other CI/CD tooling, Kafka, Github, and configuration management tools such as Puppet, Chef, or Ansible
- [Required] Deep, hands-on expertise with infrastructure-as-code (IaC) tools and practices (e.g., Terraform, CloudFormation, CDK, Pulumi), with a track record of driving IaC adoption at scale across a cloud platform organization
- [Required] Relevant experience with configuration and implementation of IaaS, Infrastructure as Code, AWS, Azure, etc.
- [Required] Expert working knowledge of infrastructure design and components, such as servers, operating systems, networks, and storage
- [Required] Basic understanding of good delivery practices and continual integration and improvement; Agile/Lean background for projects and project delivery
- [Required] Bachelor’s degree, preferably in a technical discipline (Computer Science, Mathematics, etc.), or equivalent combination of education and experience required; Master’s degree and relevant experience also considered
- [Required] 10+ years’ experience in IT systems installation, operations, administration, and maintenance of cloud systems / virtualized servers, including 5+ years in a technical leadership role
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
- \n
- [Required] AWS Solutions Architect Associate Certification or higher strongly desired
\n
