Job title: AWS DevOps Engineer AI/ML & MLOpsWork Location: Tampa, FL
Job Description:
We are looking for an experienced AWS DevOps Engineer with strong expertise in Terraform, CI/CD automation, and AI/ML platform deployment. The ideal candidate will be responsible for building, automating, and managing scalable cloud infrastructure on AWS while enabling AI/ML workloads through robust DevOps practices. This role requires hands-on experience in Infrastructure as Code (IaC), containerization, cloud-native technologies, MLOps, and automation.
Key Responsibilities
Cloud Infrastructure & Automation- Design, deploy, and manage highly available and secure AWS cloud environments.
- Develop and maintain Infrastructure as Code (IaC) using Terraform.
- Automate cloud provisioning, configuration management, and environment setup.
- Implement cloud governance, security, compliance, and cost optimization strategies.
DevOps & CI/CD- Design and manage CI/CD pipelines using GitHub Actions, Jenkins, GitLab CI/CD, or AWS CodePipeline.
- Automate application deployments across development, testing, and production environments.
- Implement GitOps and DevSecOps best practices.
- Manage source control repositories and branching strategies.
AI/ML & MLOps- Deploy, automate, and manage AI/ML solutions on AWS.
- Support ML lifecycle management, including model training, validation, deployment, and monitoring.
- Work with Amazon SageMaker for model development and deployment.
- Implement MLOps pipelines for continuous model integration and delivery.
- Collaborate with Data Scientists and AI Engineers to operationalize machine learning models.
Containerization & Orchestration- Build and manage containerized workloads using Docker.
- Deploy and manage Kubernetes clusters using Amazon EKS.
- Implement Helm charts and Kubernetes best practices for scalable deployments.
- Monitoring & Security
- Configure monitoring, logging, and alerting using CloudWatch, Prometheus, Grafana, and ELK Stack.
- Implement IAM policies, security controls, secrets management, and vulnerability scanning.
- Monitor infrastructure health and optimize system performance.
Nice-to-Have- Generative AI deployment experience using Amazon Bedrock, OpenAI, Anthropic, or Hugging Face models.
- Experience with LLM deployment, vector databases, and RAG architectures.
- Knowledge of LangChain, AI Agents, and AI workflow automation.
- Exposure to Data Engineering tools such as Glue, Athena, EMR, or Redshift.
- Experience implementing AI governance and model security frameworks.