Kubernetes Engineer Lead

Apolis

Chicago, IL

JOB DETAILS
SALARY
$70–$75 Per Hour
SKILLS
Amazon Web Services (AWS), Artificial Intelligence (AI), Automation, Capacity and Performance Management, Cloud Computing, Configuration Management, Continuous Deployment/Delivery, Continuous Integration, DevOps, Emerging Technology, Environmental Management, Fleet Management, Hybrid Cloud, Identify Issues, Incident Management, Linux Administration, Machine Tool, Onboarding, Performance Tuning/Optimization, Production Management, Production Systems, Reliability Engineering, Risk Management, Root Cause Analysis, Security Architecture, Software Engineering, Systems Administration/Management
LOCATION
Chicago, IL
POSTED
Today
Role: Kubernetes Engineer Lead
Rate: ***/hour all inclusive
Location: East coast - Remote - Preferred Miami, FL due to client visits may need to go onsite eventualy.

JD
Position Summary
We are seeking a highly skilled Senior Kubernetes Platform Engineer to lead the design, automation, operation, and modernization of enterprise Kubernetes platforms across hybrid cloud environments. This role is responsible for platform reliability, automation, observability, security, GitOps adoption, and cloud-native engineering practices that enable scalable and resilient application delivery.
Key Responsibilities
Platform Engineering & Operations
Own the lifecycle management of EKS Anywhere Kubernetes platforms, including provisioning, upgrades, security, scalability, and operational excellence.
Design and build internal platform services, tooling, and automation frameworks.
Develop custom Kubernetes operators and controllers to simplify platform administration and application onboarding.
Drive platform standardization, governance, and engineering best practices.
Cloud Native Automation & GitOps
Lead GitOps-based deployment and configuration management using ArgoCD and Helm.
Build self-service capabilities and reusable deployment patterns for engineering teams.
Automate platform operations through Infrastructure as Code (IaC) and cloud-native tooling.
Support continuous integration and continuous delivery (CI/CD) practices.
Observability & Reliability Engineering
Design and operate monitoring, logging, and distributed tracing solutions.
Implement OpenTelemetry-based observability frameworks integrated with enterprise monitoring platforms.
Lead incident management, root cause analysis (RCA), and reliability improvement initiatives.
Drive capacity planning, performance optimization, and operational maturity programs.
Security & Secrets Management
Implement secure platform architectures and Kubernetes security controls.
Design and manage secrets management solutions using HashiCorp Vault and Kubernetes-native integrations.
Automate credential management, rotation, and compliance-related controls.
Collaborate with security teams to maintain platform governance and risk mitigation.
Cloud Infrastructure & Automation
Develop and manage cloud infrastructure using Terraform and AWS services.
Build event-driven automation solutions using serverless and cloud-native technologies.
Support multi-account cloud environments and infrastructure modernization initiatives.
Champion automation-first approaches to reduce operational overhead.
AI-Enabled Platform Innovation
Develop and integrate AI-assisted operational tooling for diagnostics, troubleshooting, and productivity enhancement.
Evaluate emerging technologies and automation opportunities within platform engineering.
Promote innovation through intelligent operations and engineering efficiency improvements.

Required Qualifications
Experience
Strong experience in Platform Engineering, Cloud Engineering, DevOps, or Site Reliability Engineering (SRE).
Hands-on experience managing production Kubernetes environments.
Proven background in infrastructure automation and cloud-native technologies.
Experience supporting mission-critical enterprise platforms.
Technical Skills Summary
Kubernetes, EKS, EKS Anywhere
AWS Cloud Services
Terraform
ArgoCD, GitOps
Helm
OpenTelemetry
Datadog, Splunk, Prometheus
HashiCorp Vault
Go, Python, Bash
Linux Administration
CI/CD and Infrastructure as Code.
Detailed Technical Skill requirements
1. EKS Anywhere Expertise
Management Cluster
Bootstrap Cluster
Workload Cluster
Cluster API (CAPI)
eksctl-anywhere CLI
Cluster creation/deletion/upgrade - complete cluster life cycle
Worker scaling
Node replacement
Cluster troubleshooting
Cluster configuration YAMLs
Generated cluster-state YAMLs
kube-vip
HAProxy
LoadBalancer Services
StorageClass
PVC
PV
Trust Manager
Sealed Secrets
Vault Integration
HashiCorp Vault
Image Security
Policy enforcement
2. Rafay Platform Administration (EKS setup)
Rafay Organization structure
Projects
Clusters
Cluster Blueprints
Namespaces
Environment Manager
Addon management
Multi-Cluster Management
Fleet management
Cluster grouping
Cluster templates
Blueprint lifecycle
Versioning
Inheritance
Policy enforcement
3. Argo CD Administration
ArgoCD Architecture
App of Apps Pattern
ApplicationSets
Projects
Repositories
Clusters
Sync Policies
Cluster Generator
List Generator
Matrix Generator
Template Management
Fleet Deployment
Vault Integration
Cluster Integration
4. Harness Platform Administration
Projects
Organizations
Connectors
Secrets
Environments
Services
Infrastructure Definitions
Templates
Repositories
Clusters
Sync Policies
Cluster Generator
List Generator
Matrix Generator
Template Management
Fleet Deployment
Vault Integration
Cluster Integration
5. Helm Charts
Helm Architecture
Charts
Releases
Repositories
Dependencies
Templates
Values.yaml
Overrides
Environment-specific customization
Secret references
Resource tuning

About the Company

A

Apolis

Since 1996, RJT has provided successful SAP, Oracle, and IT consulting solutions and staffing services to clients around the world. The new Apolis brings you the same personalized service fortified with a greater array of IT solutions, global expertise, and cost-management strategies.

We are a global IT consultancy that seamlessly integrates experts and leading-edge solutions into your organization so you can focus on what really matters.

COMPANY SIZE
500 to 999 employees
INDUSTRY
Computer/IT Services
EMPLOYEE BENEFITS
Paid Sick Days, Employee Referral Program, Employee Events, Retirement / Pension Plans
WEBSITE
https://www.apolisrises.com/