| Location | Santa Clara, CA |
Job Title: GenAI Engineer LLM Infrastructure & Inference ServicesWork Location: Santa Clara, CA 95051Vendor Rate: $90/hrContract duration: 6monthsTarget Start Date: 17 Aug 2026Job Details:GenAI Engineer LLM Infrastructure & Inference Services (Concise Summary) Deploy, host, and manage Large Language Models (LLMs) on GPU infrastructure for production environments. Build scalable and high-performance inference services using vLLM, TensorRT-LLM, Triton Inference Server, and Ray Serve. Optimize model serving for latency, throughput, GPU utilization, and cost efficiency. Develop AI platform services and APIs using Python, FastAPI, microservices, and Kubernetes. Implement RAG pipelines, vector databases, and agentic AI frameworks such as LangChain and LangGraph. Manage GPU infrastructure, containerization, and cloud deployments across AWS, Azure, or GCP. Establish MLOps/LLMOps practices including CI/CD, model deployment, monitoring, observability, and governance. Perform performance tuning, benchmarking, capacity planning, and production support for enterprise GenAI platforms. Collaborate with architects, data scientists, and product teams to deliver scalable, secure, and reliable AI solutions.
Project Code: Foundation track FP code