Search

Platform Engineer

PublishedPublished: 6/14/2022
Technology

Job Description

Need a Kubernetes SME / Kubernetes Architect / Kubernetes Engineer with

\n

    \n
  • Deep Kubernetes experience (Troubleshooting, configuring for scale), doing this at scale involves working with larger data sets (aka Big Data) and they use Spark on Kubernetes to run their processes.
  • \n

  • Spark: Spark would be 50% of the work.
  • \n

  • SQL querying will assess on this as well.
  • \n

\n


\n

Senior Kubernetes Engineer Overview:

\n

    \n
  • We are seeking an expert level Kubernetes Engineer to configure, deploy, and operate mission-critical Amazon EKS infrastructure supporting petabyte-scale data processing workloads.
  • \n

  • This role requires deep technical expertise in Kubernetes internals, large-scale cluster management, and Apache Spark optimization to support thousands of concurrent jobs processing large volumes of Bigdata.
  • \n

\n


\n

Core Responsibilities:

\n

    \n
  • Design, deploy, and maintain production-grade Amazon EKS clusters architected for petabyte-scale data processing with high availability and fault tolerance.
  • \n

  • Operate in air-gapped private VPC environments without internet access, managing secure package repositories and container registries.
  • \n

  • Debug and resolve complex distributed systems issues across EKS, Karpenter, and Spark, including scheduling bottlenecks, node scaling delays, and cascading failures under heavy load.
  • \n

  • Implement Karpenter consolidation and disruption policies balancing cost optimization with job resiliency.
  • \n

  • Manage spot and on-demand instance strategies with robust node interruption handling.
  • \n

  • Define and enforce ResourceQuotas, LimitRanges, and PriorityClasses to ensure fair resource distribution and prevent resource starvation.
  • \n

  • Configure persistent volume claims and provision container-native storage using Amazon EBS CSI driver for block storage and Amazon EFS CSI driver for shared file access.
  • \n

  • Optimize storage configurations for resiliency, high-throughput data processing workloads.
  • \n

  • Implement logging, alerting, and anomaly detection to identify provisioning failures, executor loss, and throughput degradation before broader system impact.
  • \n

  • Design fault-tolerant architectures with retry strategies, checkpointing, and graceful degradation patterns minimizing re-computation on failure.
  • \n

  • Develop and manage configs in a private VPC without internet access.
  • \n

\n


\n

Preferred Qualifications:

\n

    \n
  • Kubernetes Certifications (CKA, CKAD, or CKS) and AWS Certifications.
  • \n

  • Active contributions to open-source Kubernetes, Karpenter or Spark projects.
  • \n

  • Familiarity with FinOps practices and cost optimization at scale.
  • \n

  • Background in data engineering or analytics platforms.
  • \n

\n


\n

#LI-CGTS

\n

#TS-3142

Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...
Loading...