L2 Production Support Engineer/SRE with Banking/Finance, Cloud, Kubernetes, Python experience
Job Description
Role: Production Support Engineer /Site Reliability Engineer SRE
\n
Location: New York - New York (Onsite/ Hybrid)
\n
Duration: FTE
\n
\n
About Smart IT Frame:
\n
At Smart IT Frame, we connect top talent with leading organizations across the USA. With over a decade of staffing excellence, we specialize in IT, healthcare, and professional roles, empowering both clients and candidates to grow together.
\n
\n
Job Description:
\n
\n
- \n
- Production Support Engineer Site Reliability Engineer SRE
- Experience 5 Years
- Domain Banking Financial Services Capital Markets FinTech Digital Assets
\n
\n
\n
\n
Job Summary
\n
We are seeking an experienced Production Support Engineer / Site Reliability Engineer (SRE) to support mission-critical applications and cloud-native platforms. The ideal candidate will have strong expertise in Kubernetes, application support, observability, automation, and cloud infrastructure, with a focus on ensuring high availability, reliability, scalability, and operational excellence across enterprise environments.
\n
Key Responsibilities
\n
- \n
- Provide L2/L3 production support and incident management for business-critical applications.
- Perform troubleshooting, root cause analysis (RCA), and implement preventive measures to reduce recurring incidents.
- Support and optimize cloud environments across Microsoft Azure and AWS.
- Manage and troubleshoot Kubernetes (AKS), Docker, virtual machines, and distributed systems.
- Build and maintain observability and monitoring solutions using Site24x7, Splunk, Prometheus, Grafana, Azure Monitor, and CloudWatch.
- Develop automation scripts using Python, Bash, or PowerShell to improve operational efficiency.
- Support CI/CD pipelines, deployment automation, and release management processes.
- Troubleshoot application, API, database, and infrastructure-related issues.
- Ensure compliance through IAM management, certificate renewals, vulnerability remediation, and security best practices.
- Contribute to AI-driven operational automation, intelligent monitoring, and AIOps initiatives.
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
\n
Required Skills
\n
Cloud & Infrastructure
\n
- \n
- Azure
- AWS
- AKS (Azure Kubernetes Service)
- Azure Monitor
- CloudWatch
- Terraform
\n
\n
\n
\n
\n
\n
\n
Containers & Orchestration
\n
- \n
- Kubernetes
- Docker
- Helm
\n
\n
\n
\n
Monitoring & Observability
\n
- \n
- Site24x7 (Mandatory)
- Splunk
- Prometheus
- Grafana
- New Relic
\n
\n
\n
\n
\n
\n
DevOps & CI/CD
\n
- \n
- Jenkins
- Azure DevOps
- GitHub Actions
- GitLab CI/CD
\n
\n
\n
\n
\n
Programming & Scripting
\n
- \n
- Python (Mandatory)
- Bash
- PowerShell
\n
\n
\n
\n
Databases
\n
- \n
- PostgreSQL
- SQL Server
- MySQL
\n
\n
\n
\n
Service Management
\n
- \n
- ServiceNow
- Jira
- ITIL Processes
\n
\n
\n
\n
Operating Systems
\n
- \n
- Linux (RHEL, Ubuntu)
\n
\n
Preferred Skills
\n
- \n
- Experience in Banking, Financial Services, Capital Markets, FinTech, or Digital Assets domains.
- Exposure to MLOps, MLflow, AIOps, and AI/LLM-based operational automation.
- Hands-on experience with Infrastructure as Code (IaC) using Terraform.
- Experience working in regulated financial services environments.
\n
\n
\n
\n
\n
Key Competencies
\n
- \n
- Strong troubleshooting and analytical skills.
- Incident management and RCA expertise.
- Reliability engineering mindset.
- Automation-first approach.
- Excellent stakeholder management and communication skills.
- Cross-functional collaboration and customer-focused attitude.
\n
\n
\n
\n
\n
\n
\n
Skills
\n
Mandatory Skills : Python, SITE24X7
\n
\n
📩 Apply today or share profiles at deighton@smartitframe.com
