| Location | Santa Clara, CA |
Role: Senior LLMOps / MLOps Engineer(23382-1)
Location: Santa Clara, CA
Onsite Requirement - Yes
Number of days onsite - 5 days
Must Have Skills
Skill 1 - Strong proficiency in Python and software engineering best practices
Skill 2 - 14+ years of experience in MLOps, LLMOps, AI/ML Platform Engineering
Skill 3 - Strong expertise in LLM Inferencing and Model Hosting using vLLM, SGLang, TGI, Triton, Ray Serve, Azure ML, or Databricks Model Serving
Good To have Skills -
Skill 1 - Exposure to AI Observability, Governance, and Responsible AI practices
Mandatory if Applicable
Domain Experience (If any) - Senior LLMOps / MLOps Engineer
Summary
We are looking for a highly skilled Senior LLMOps / MLOps Engineer with strong expertise in LLM inferencing, model hosting, and serving Large Language Models (LLMs) at scale. The ideal candidate should be a hands-on engineer with proven experience deploying and optimizing open-source LLMs, building high-performance inference platforms using technologies such as vLLM, SGLang, TGI, Triton, and Ray Serve, and driving GPU utilization, latency, throughput, and cost optimization. This is a highly technical role requiring active involvement in designing, building, troubleshooting, and optimizing production AI systems. Experience in MLOps platforms and scalable AI infrastructure is essential.
Must-Have Skills
Good-to-Have Skills