Machine Learning Evaluation Engineer
Major consumer tech company · Cupertino, CA
- Location
- Cupertino, CA · Hybrid
- Salary band
- Band not stated
- Type
- Contract
- Level
- Senior
- Work authorization
- No sponsorship — now or in future
Stack
Python · Machine Learning · LLM · Generative AI · Conversational AI · NLP · RAG · Agentic AI · ML Evaluation · Data Pipelines · Statistical Analysis · CI/CD · Prompt Engineering · Embeddings · Human-in-the-Loop Evaluation · Recommendation Systems · Multimodal AI · Responsible AI · Distributed Data Processing
About the role
Job Title: Machine Learning Evaluation Engineer Location: Cupertino, CA (Hybrid) Duration: Contract Job Description: As a ML Evaluation Engineer, you will design and build scalable evaluation systems for LLM, Generative AI, Conversational AI, and Agentic AI product. Minimum Qualifications: 5+ years of exp in ML Engineering, ML Evaluation, Software Engineering, Data Science, Quality Engineering, or a related technical field. Strong programming skills in Python and exp developing production-quality software, ML systems, data pipelines, or evaluation infrastructure. Exp developing or evaluating LLMs, Generative AI, Conversational AI, NLP, recommendation systems, or other ML-driven products. Exp designing automated ML evaluation frameworks, metrics, benchmarks, datasets, or experimentation methodologies. Understanding of modern LLM application architectures, including prompting, embeddings, retrieval-augmented generation (RAG), tool use, and agentic workflows. Exp with model-based evaluation techniques and an understanding of the strengths and limitations of approaches such as LLM-as-a-Judge. Exp with Human-in-the-Loop evaluation, annotation, or data-quality workflows. Strong understanding of statistical analysis, experimentation, sampling, and measurement methodologies. Exp performing model error analysis, failure analysis, and root-cause investigation. Ability to work effectively across Machine Learning, Engineering, Product, Quality, and Data teams. Excellent written and verbal communication skills, with the ability to translate complex technical findings into clear, actionable recommendations. Preferred Qualifications: Exp building evaluation infrastructure for production-scale LLM or Generative AI applications. Exp evaluating RAG, conversational systems, AI agents, personalization, recommendations, or multimodal AI. Exp building golden datasets, regression suites, automated quality gates, and continuous evaluation pipelines. Exp integrating ML evaluation into CI/CD and production release processes. Exp with prompt evaluation, model comparison, experiment tracking, and AI observability. Exp evaluating multilingual AI experiences across languages, locales, and markets. Familiarity with responsible AI evaluation, including robustness, safety, bias, and adversarial testing. Exp developing internal ML platforms, developer tooling, or self-service evaluation capabilities used across multiple teams. Exp working with large-scale datasets and distributed ML or data-processing infrastructure.
Recruiter contact
Unlock this recruiter's name, email and phone for a one-time $100. Payment is per role — you only pay for the introductions you actually want.
Want the full job-search package instead? Talk to us →
The recruiter's details are not public. Helen makes the introduction, checks your resume against the requirements first, and follows up on the reply.