| Location | Frisco, TX |
This role serves as a vendor-provided Site Reliability Engineer responsible for improving and protecting the reliability, scalability, and performance of a platform currently hosted on Microsoft Azure that will be migrated to Tencent Kubernetes Engine (TKE) in the future. It manages availability, latency, performance, security, and capacity while enabling efficient, automated software delivery across both the current Azure environment and the upcoming TKE platform. The role differentiates by combining deep Azure cloud infrastructure expertise with Kubernetes-based container orchestration, positioning the team for a smooth cloud-to-TKE migration. Success is measured by improved system uptime, faster incident resolution, and a reliable, well-supported migration path. The work directly impacts IT service quality, operational resilience, and customer experience through the transition and beyond.
Responsibility | Approx. % of Time |
Monitor, troubleshoot, and resolve incidents affecting availability, latency, and performance of current Azure-hosted workloads | 20% |
Provision, configure, and manage Azure infrastructure (VMs, networking, storage, IAM) to support production and non-production environments | 20% |
Support planning and execution of the platform's migration from Azure to Tencent Kubernetes Engine (TKE), including workload containerization and cutover activities | 20% |
Design, build, and maintain CI/CD pipelines that support automated deployment and testing today on Azure and going forward on TKE | 15% |
Build and maintain observability tooling - dashboards, alerts, logging, and health checks - to proactively identify and address system risks across both environments | 10% |
Drive automation and infrastructure-as-code practices to reduce manual toil and improve deployment consistency | 10% |
Collaborate with internal engineering teams and stakeholders to support incident response, capacity planning, and migration readiness | 5% |
Education: Bachelor's degree in Computer Science, Engineering, or related field, or equivalent practical experience.
Experience: 4+ years in Site Reliability Engineering, DevOps, or Cloud Infrastructure roles, including hands-on production experience with Microsoft Azure (compute, networking, storage, IAM). Experience with Tencent Kubernetes Engine (TKE) or comparable Kubernetes platforms strongly preferred, as the platform will migrate to TKE.
Technical Skills: Proficiency with CI/CD tooling (e.g., Azure DevOps, Jenkins, GitLab CI), containerization and orchestration (Kubernetes, Docker, TKE), infrastructure-as-code (Terraform, ARM/Bicep), scripting/automation (Python, Bash, PowerShell), and monitoring/observability platforms (e.g., Grafana, Prometheus, Azure Monitor).
Other: Strong troubleshooting and incident-response skills; experience supporting cloud platform migrations a plus; ability to work effectively as an embedded vendor resource within a client engineering team; on-call availability as required.
Azure certifications (e.g., AZ-104, AZ-400, AZ-500) and/or Kubernetes certifications (CKA/CKAD). Prior experience migrating workloads from a cloud VM-based platform to Kubernetes/TKE, including containerization of legacy services.