Job Description
Optomi, in partnership with a large telecommunications company, is seeking a Principal Software Engineer to join their team in the Bedminster, NJ Area. The ideal candidate will have experience building, deploy, and scale production AI inference infrastructure. You’ll work across Kubernetes, vLLM, Python, GPUs, and cloud/on-prem environments to make model serving reliable, scalable, and highly automated.
\n
\n
What The Right Professional Will Enjoy
\n
- \n
- Building and operating LLM inference infrastructure at scale.
- Deploying and managing models using vLLM and Kubernetes.
- Developing Python automation for model deployment, scaling, updates, rollbacks, and health monitoring.
- Managing GPU capacity, performance, reliability, and production workloads.
- Building tools that allow engineering teams to deploy and access models without managing underlying infrastructure.
- Supporting both R&D and production inference environments.
- Troubleshooting and optimizing model-serving performance and reliability.
- Remote flexibility
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
Apply Today If Your Background Includes
\n
- \n
- Strong software engineering and Python skills.
- Deep hands-on Kubernetes experience.
- Production experience with LLM inference/model serving, ideally vLLM.
- Experience with GPU infrastructure (Not required), containers, and CI/CD.
- Understanding of scaling, automation, observability, and reliability for distributed systems.
- Experience with cloud and/or on-prem infrastructure preferred.
\n
\n
\n
\n
\n
\n
\n
