Job Description
- \n
- Compute: HPC environments, AMD/NVIDIA GPUs, Linux Kernel/OS, NUMA awareness ,CPU pinning, huge pages, and performance optimization
- Storage: NVMe storage and distributed storage systems, High performance Block, file, and Object Storage Experience
- Software-Defined Networking (SDN): Design, deployment, and operational support of SDN infrastructure
\n
\n
\n
\n
For L1 role (Major skill required)
\n
- \n
- GPU infrastructure (is not mandatory)
- Networking
- Block storage
- File storage
- Production support and reliability engineering
\n
\n
\n
\n
\n
\n
\n
Hi ,
\n
\n
Embedded Platform/Infrastructure Engineer (Level 2 Engineer)
\n
Sunnyvale, CA or San Jose, CA (5 days Onsite, Final round F2F)
\n
Contract & FTE (Both)
\n
\n
Skill required – GPU (Must), Exposure SRE, Storage & Network, GPU Infrastructure & SDN.
\n
\n
\n
As an Embedded PE, we will function as a member of a foundation engineering team,
\n
owning reliability, performance, scalability, and operational excellence for critical
\n
infrastructure supporting fastest-growing AI cloud platforms.
\n
\n
This role requires deep technical expertise in a specific infrastructure domain and the
\n
ability to operate across the full service lifecycle, including design, implementation,
\n
optimization, incident response, and roadmap planning.
\n
\n
Required Qualifications
\n
10 to 15+ years of infrastructure engineering, platform engineering, SRE, or systems engineering
\n
experience.
\n
Deep expertise in a specific infrastructure domain.
\n
Strong Linux systems knowledge.
\n
Demonstrated experience operating production-critical infrastructure.
\n
Hands-on incident response and troubleshooting experience.
\n
Experience building automation and operational tooling.
\n
Ability to clearly articulate technical contributions to complex projects.
\n
\n
Responsibilities
\n
Embed within a foundation engineering team and operate as a domain expert.
\n
Participate in on-call rotations and production incident management.
\n
Design, implement, and improve infrastructure reliability and scalability.
\n
Collaborate with architects and engineers on long-term technical roadmaps.
\n
Analyze complex infrastructure performance bottlenecks and system failures.
\n
Build automation to improve operations, deployment, and observability.
\n
Define and track service reliability objectives (SLIs/SLOs).
\n
Contribute to platform standards and engineering best practices.
\n
Partner with data center and operations teams during service-impacting events.
\n
Drive continuous improvements in service stability, efficiency, and performance.
\n
\n
Direct: 469-421-5604 , Ext- 218 • nitesh.j@tekgence.com
\n
Linkedin: linkedin.com/in/nitesh-ch-a378b5222
\n
6655 Deseo Dr, Suite 104,Irving, TX , 75039 • www.tekgence.com
