Search

325,607 Jobs

No logo available
Wyetech
locationAnne Arundel County, MD, USA
PublishedPublished: 7/26/2026
No logo available
Check Point Software Technologies
locationDallas, TX, USA
PublishedPublished: 7/26/2026
No logo available
WAGSTAFF INC
locationSpokane, WA, USA
PublishedPublished: 7/26/2026
No logo available
System One
locationMidvale, UT, USA
PublishedPublished: 7/26/2026
No logo available
System One
locationDallas, TX, USA
PublishedPublished: 7/26/2026
No logo available
Edifice Protection Group
locationCleveland, OH, USA
PublishedPublished: 7/26/2026
No logo available
Subway
locationPalm Springs North, FL 33015, USA
PublishedPublished: 7/26/2026
No logo available
Tumi Hospitality
locationAustin, TX, USA
PublishedPublished: 7/26/2026
No logo available
City of Philadelphia
locationPhiladelphia, PA, USA
PublishedPublished: 7/26/2026
No logo available
La Scuola International School
locationEast Palo Alto, CA, USA
PublishedPublished: 7/26/2026

Site Reliability Engineer

Resolve Tech Solutions
locationIrving, TX, USA
PublishedPublished: 6/14/2022
Technology
Full Time

Job Description


\n

About the Company

\n


\n

Role: SRE RunOps Engineer

\n


\n

Location: Irving, TX

\n


\n

Onsite job

\n


\n

About the Role

\n


\n

Production Support & Incident Management

\n


\n

Serve as a primary responder for production incidents, ensuring rapid triage, mitigation, and resolution.

\n


\n

Responsibilities

\n


\n

    \n
  • Lead root cause analysis (RCA) and drive long‑term corrective actions.
  • \n

  • Maintain and improve incident response processes, runbooks, and escalation paths.
  • \n

  • Collaborate with engineering, QA, and product teams to prevent recurrence of issues.
  • \n

\n


\n

AWS Infrastructure Operations

\n


\n

    \n
  • Support and optimize AWS services such as EC2, ECS/EKS, Lambda, S3, CloudWatch, IAM, RDS, and VPC networking.
  • \n

  • Monitor system health, performance, and capacity across cloud environments.
  • \n

  • Implement infrastructure best practices around reliability, scalability, and cost efficiency.
  • \n

  • Assist with deployments, environment configuration, and CI/CD pipelines.
  • \n

\n


\n

Database & Storage Support

\n


\n

    \n
  • Manage and troubleshoot MongoDB clusters, including performance tuning, replication, backups, and failover.
  • \n

  • Diagnose query performance issues and collaborate with developers on schema optimization.
  • \n

  • Ensure data integrity, availability, and recovery readiness.
  • \n

\n


\n

Monitoring, Observability & Alerting

\n


\n

    \n
  • Use New Relic, CloudWatch, and other observability tools to monitor application and infrastructure performance.
  • \n

  • Build dashboards, alerts, and telemetry that provide actionable insights.
  • \n

  • Continuously refine monitoring thresholds to reduce noise and improve signal quality.
  • \n

  • Experience with on-call rotations and 24/7 production environments.
  • \n

  • Work cross-functionally with the various teams in the organization and help establish SLOs and achieve those SLOs.
  • \n

\n


\n

Qualifications

\n


\n

    \n
  • 5+ years of experience in SRE, DevOps, Production Support, or similar operational roles.
  • \n

\n


\n

Required Skills

\n


\n

    \n
  • Strong hands-on experience with AWS services and cloud-native architectures.
  • \n

  • Proficiency with MongoDB administration and troubleshooting.
  • \n

  • Experience with New Relic or similar APM/observability platforms.
  • \n

  • Experience using additional tools like Postman, Intune, and Firebase, Service Now, Cloudwatch.
  • \n

  • Strong understanding of Linux systems, networking, and distributed systems.
  • \n

  • Solid scripting skills (Python, Bash, or similar).
  • \n

  • 5+ years Monitoring and Alarming in all environments and familiar with tools like Mongo Charts, New Relic, Cloudwatch, Service Now.
  • \n

  • Proven experience managing high-severity incidents and driving RCA processes.
  • \n

  • Familiarity with CI/CD tools (Jenkins, GitHub Actions, GitLab CI, etc.).
  • \n

  • Experience with container orchestration (ECS, EKS, Kubernetes).
  • \n

  • Knowledge of message queues (Kafka, SQS, RabbitMQ).
  • \n

  • Exposure to microservices architectures.
  • \n

  • Certifications such as AWS Solutions Architect, AWS SysOps, or MongoDB DBA.
  • \n

  • Working experience with IoT devices, and Microsoft Intune.
  • \n

\n


\n

Pay range and compensation package

\n


\n

[Pay range or salary or compensation]

\n


\n

Equal Opportunity Statement

\n


\n

[Include a statement on commitment to diversity and inclusivity.]

\n