Deploy, monitor, and maintain AI and ML models in production environments, ensuring performance and reliability Infrastructure Management: Provision, configure, and maintain cloud and on-premises infrastructure, including GPU servers and high-performance computing resources CI/CD Pipeline Development: Build and manage continuous integration and continuous deployment pipelines for AI applications Automation and Scripting: Automate repetitive tasks, infrastructure provisioning, and model deployment using scripting languages like Python Collaboration: Work closely with cross-functional teams, including engineers, data scientists, and product managers, to design and implement AI solutions Monitoring and Security: Implement monitoring, logging, and security best practices to ensure AI systems operate safely and efficiently Documentation and Training: Create technical documentation and provide training to end-users or team members on AI system usage and maintenance. Requirements: Programming: Proficiency in Python is essential; familiarity with other languages like Bash or Java is beneficial Cloud Platforms: Experience with cloud services such as AWS, Azure, or Google Cloud for AI deployment AI/ML Knowledge: Understanding of machine learning models, data pipelines, and AI frameworks (e.g., TensorFlow, PyTorch) is highly desirable DevOps Tools: Experience with CI/CD tools (Jenkins, GitLab CI), containerization (Docker, Kubernetes), and infrastructure-as-code (Terraform, Ansible) is important Problem-Solving: Ability to troubleshoot complex system issues and optimize AI workflows Communication: Strong collaboration and communication skills to work with technical and non-technical stakeholders.