← All roles

Site Reliability Engineer

US client via staffing partner · Plano, TX

Location
Plano, TX · Onsite
Salary band
Band not stated
Type
Full-time
Level
Senior
Work authorization
Not stated

Stack

Site Reliability Engineering · MuleSoft · TIBCO · Distributed Systems · Dynatrace · Splunk · Prometheus · Grafana · Jenkins · Git · Ansible · Shell · Python · PowerShell · Linux/Unix · Windows · Microservices · CI/CD · Kubernetes · Terraform

About the role

Role: Site Reliability Engineer Location: Plano, TX – Onsite Full Time only Description We are seeking an experienced Site Reliability Engineer SRE Lead to drive platform reliability observability and operational excellence across the API Services ecosystem This role combines - Production engineering and reliability leadership - Platform security and vulnerability remediation - Ownership of large-scale distributed runtime environments Key responsibilities include - Leading reliability engineering for high-scale API platforms 40K runtimes - Driving EOL remediation and platform stabilization efforts - Implementing SRE best practices - SLIs SLOs error budgets - Incident management and postmortem culture - Enhancing observability monitoring and proactive fault detection - Building resilient platforms capable of handling AI-driven usage patterns and threat models - Supporting global production environments with on-call and escalation coverage Required Skills - Strong experience in Site Reliability Engineering Production Engineering - Hands on expertise with - MuleSoft TIBCO or similar middleware platforms - Large-scale distributed systems and runtime management - Deep understanding of - System reliability scalability and high availability design - Incident management root cause analysis and problem management - Experience with - Observability tools: eg Dynatrace Splunk Prometheus Grafana - CI/CD pipelines Jenkins Git Ansible - Strong scripting automation skills - Shell Python PowerShell - Experience managing Linux/Unix and Windows production environments - Knowledge of - Microservices API platforms and cloud-based architectures - Understanding of - Platform security vulnerability remediation and risk mitigation in production systems - Excellent troubleshooting skills in high-pressure real-time environments Desired Skills - Experience implementing SRE frameworks SLIs SLOs error budgets - Familiarity with - Kubernetes containerized platforms - Infrastructure as Code Terraform Ansible - Exposure to - AI-driven operational monitoring or security tooling - Large-scale platform modernization or migration programs - Middleware certifications MuleSoft or equivalent - Experience in regulated environments eg financial services Mandatory Skills: Ansible, Automation & Scripting, Dynatrace, Jenkins, Prometheus

Recruiter contact

Unlock this recruiter's name, email and phone for a one-time $100. Payment is per role — you only pay for the introductions you actually want.

Want the full job-search package instead? Talk to us →

The recruiter's details are not public. Helen makes the introduction, checks your resume against the requirements first, and follows up on the reply.

Ask your consultant for an introduction.