Senior AI/ML Platform Engineer
We're seeking a Senior AI/ML Platform Engineer to build and scale enterprise AI/ML platforms that support production-grade machine learning and generative AI workloads. This is a hands-on engineering role focused on creating secure, reliable, observable, and scalable AI/ML infrastructure across cloud and data platforms.
Rather than developing one-off AI solutions, you'll build the foundational systems, automation, and governance that enable AI/ML teams to deploy and operate models at scale.
Key Responsibilities
- Design, build, and support enterprise AI/ML platforms across Databricks, AWS, MLflow, model registries, model serving, feature stores, and related technologies.
- Develop reusable patterns for model development, deployment, monitoring, security, and production support.
- Implement CI/CD pipelines, infrastructure automation, secrets management, access controls, and deployment frameworks.
- Support batch, streaming, real-time, and API-based model deployment architectures.
- Define engineering standards for experimentation, model promotion, observability, governance, and operational support.
- Partner with data, security, infrastructure, and architecture teams to deliver production-ready AI/ML capabilities.
- Build reference architectures, templates, and platform enablement resources to support enterprise adoption.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Systems, or related field (or equivalent experience).
- Experience building, operating, or supporting production AI/ML platforms within cloud environments.
- Hands-on experience with at least one AI/ML platform such as Databricks, AWS SageMaker, MLflow, Azure ML, or Vertex AI.
- Strong background with CI/CD, Infrastructure as Code, environment management, secrets management, access controls, and deployment automation.
- Experience supporting model development and deployment beyond experimentation and notebook-based workflows.
- Solid understanding of cloud-native architecture, APIs, containers, compute, storage, and observability.
- Experience creating reusable engineering frameworks, templates, and platform standards.
- Proven ability to support AI/ML workloads in governed, production environments.
- Strong troubleshooting skills across platform, deployment, performance, and integration challenges.
Preferred Qualifications
- Deep Databricks experience including Unity Catalog, MLflow, Model Serving, Jobs/Workflows, Clusters, Permissions, and Cost Optimization.
- AWS expertise including IAM, S3, Lambda, ECS/EKS, API Gateway, SageMaker, Bedrock, Networking, and Security.
- Experience in highly regulated or mission-critical environments such as energy, industrial, healthcare, finance, or manufacturing.
- Background in platform cost management and workload optimization.
- Experience creating enablement materials for engineers and data science teams.
Ideal Background
Candidates with experience in MLOps, ModelOps, AI Platform Engineering, Machine Learning Infrastructure Engineering, or AI/ML Enablement Platforms will be particularly successful in this role.