Site Reliability Engineer
US client via staffing partner · Sunnyvale, CA
- Location
- Sunnyvale, CA · Onsite
- Salary band
- Band not stated
- Type
- Contract
- Level
- Senior
- Work authorization
- Not stated
Stack
Distributed Systems · Performance Engineering · Java · Kubernetes · Linux · PostgreSQL · Azure · Python · Bash · PowerShell · ActiveMQ · Load Testing · CI/CD Pipelines · Monitoring Tools · Large Language Models (LLMs) · Capacity Planning · Observability · Networking · Database Optimization · Automation
About the role
Position: Site Reliability Engineer - 10+ Year exp required Location: Sunnyvale CA (Onsite) Long term contract Job Summary We are seeking an experienced engineer who can analyze, diagnose, and optimize performance and reliability of large-scale distributed systems. This role requires deep technical understanding across the entire application stack, the ability to read and reason about code, and the capability to provide data-backed answers to both engineering teams and business stakeholders. This role goes beyond traditional operations or DevOps. The successful candidate will think like a software engineer, act like a systems engineer, and operate with a production-first mindset. Key Responsibilities Performance & Reliability Engineering - Analyze and resolve performance issues such as high latency, slow login, throughput degradation, and system instability. - Perform deep, end-to-end investigations across the full stack including: - Load balancers and traffic routing - Web server and application runtime configurations - Middleware and messaging systems - Database performance (queries, indexing, pooling) - Kubernetes clusters (pods, resources, scaling behavior) - Linux OS tuning (CPU, memory, I/O, ulimits, networking) - Identify root causes and propose clear, actionable engineering solutions. Distributed Systems Design - Design, review, and influence high-performance, highly-available distributed architectures. - Evaluate trade-offs related to scalability, latency, fault tolerance, and cost. - Partner with development teams early to prevent reliability and performance issues before production. Capacity Planning & Scalability - Assess system readiness for growth scenarios such as: - "We plan to onboard 10,000 users in 6 months — can the system support it?" - Perform capacity and scale analysis for: - Application tiers - Databases - Messaging systems - Kubernetes compute and storage - Provide evidence-based recommendations supported by metrics, benchmarks, and production data. Engineering Collaboration - Work closely with software engineering teams to: - Review performance-critical code paths - Propose improvements at code, configuration, or infrastructure level - Improve system observability (metrics, logs, traces) - Communicate complex technical findings clearly to both engineers and business stakeholders. Required Technical Skills - Strong understanding of distributed systems and performance engineering - Ability to read, analyze, and troubleshoot Java code - Hands-on experience with: - Kubernetes (resource management, scaling, container behavior) - Linux internals and tuning - PostgreSQL (queries, indexing, performance optimization) - Proven experience building or operating high-availability, high-throughput systems - Strong analytical and problem-solving skills with a data-driven approach Nice to Have - Experience with Azure cloud services - Messaging systems such as ActiveMQ - Load testing and benchmarking experience - Background in roles such as SRE, Performance Engineering, Platform Engineering Required Skills & Qualifications Technical Skills - Hands-on experience with cloud platforms (Azure) - Strong scripting skills (e.g., Python, Bash, PowerShell, or similar) - Experience with deployment pipelines, automation, and monitoring tools - Solid understanding of cloud infrastructure, networking, and application operations LLM & AI Experience - Practical experience working with Large Language Models (LLMs) - Familiarity with applying LLMs to engineering or operational workflows is required Professional Attributes - Strong desire to learn and deeply understand complex systems - Self-starter with the ability to take ownership and drive initiatives independently - Demonstrates leadership, accountability, and problem-solving mindset - Strong collaboration and communication skills
Recruiter contact
Unlock this recruiter's name, email and phone for a one-time $100. Payment is per role — you only pay for the introductions you actually want.
Want the full job-search package instead? Talk to us →
The recruiter's details are not public. Helen makes the introduction, checks your resume against the requirements first, and follows up on the reply.