Senior Kubernetes Platform Engineer Lead

Apolis

New York, NY

JOB DETAILS
SALARY
$65–$65 Per Hour
SKILLS
Automation, Budgeting, Cloud Architecture, Cloud Computing, CompTIA Security+, Continuous Deployment/Delivery, Continuous Improvement, Continuous Integration, Customer Support/Service, DevOps, Documentation, Environmental Compliance, Federal Information Processing Standards (FIPS), GitHub, Hybrid Cloud, Incident Response, Jenkins, Linux Operating System, Metrics, On Call, Process Improvement, Python Programming/Scripting Language, Regulatory Compliance, Reliability Engineering, Scripting (Scripting Languages), Software Engineering, Systems Reliability, Technical Leadership, U.S. National Institute of Standards and Technology (NIST), United States Citizen, United States Department of Defense (DoD)
LOCATION
New York, NY
POSTED
3 days ago
Job Description: US Citizen or GC Holder Required - FedRAMP Requirement
REMOTE - prefer PST hours
5 days per week/ 8 hours per day

Senior Kubernetes Platform Engineer Lead

Top Skills Required :

1. Kubernetes (GKE / AKS / EKS) Production grade, regulated environments
2. CI/CD & Automation (GitHub Actions, Jenkins, ArgoCD, GitOps)
3. FedRAMP High / IL5 Compliance & Observability (Prometheus, Grafana, ELK, OpenTelemetry)

About the Role
We build technology that simply works reliable, secure, and easy to use. We re looking for a Senior Site Reliability Engineer (SRE) - Technical Leader to help us design, operate, and scale a Kubernetes-based platform supporting highly regulated environments, including FedRAMP High and DoD IL5.
This role sits at the intersection of software engineering and infrastructure. You ll work closely with engineers across the stack to ensure our platform is resilient, observable, compliant, and developer-friendly without slowing teams down.

What You ll Do
" Design, build, and operate production-grade Kubernetes platforms in regulated environments
" Improve system reliability through automation, thoughtful design, and continuous iteration
" Define and drive SLOs, SLIs, and error budgets to guide reliability decisions
" Build and evolve CI/CD pipelines that are secure, scalable, and easy to use
" Implement robust observability (metrics, logs, traces) to make systems understandable and actionable
" Reduce operational toil by automating repetitive processes and improving workflows
" Partner with security and compliance teams to meet FedRAMP High and IL5 requirements without sacrificing developer velocity
" Support ATO processes, including documentation, controls implementation, and audit readiness
" Participate in on-call rotations supporting customer requests and paging alerts
" Participate in incident response, blameless postmortems, and continuous improvement efforts
" Help shape a platform that engineers enjoy using

What You Bring
" 10+ years of experience in SRE, DevOps, or infrastructure engineering
" Strong experience running Kubernetes in production (EKS, AKS, GKE, or upstream)
" Hands-on experience working in FedRAMP High and/or DoD IL5 environments
" Solid understanding of cloud infrastructure, Linux systems, and networking fundamentals
" Experience with Infrastructure as Code (Terraform preferred)
" Familiarity with CI/CD systems (GitHub Actions, GitLab CI, Jenkins, ArgoCD)
" Proficiency in scripting or programming (Python, Go)
" Experience building or operating observability platforms (Prometheus, Grafana, OpenTelemetry, ELK)
" Working knowledge of compliance frameworks (e.g., NIST 800-53, STIGs, RMF)

Nice to Have
" Experience with service mesh technologies (Istio, Linkerd)
" Familiarity with policy-as-code (OPA/Gatekeeper, Kyverno)
" Experience with GitOps workflows
" Exposure to multi-cluster or hybrid cloud architectures
" Knowledge of FIPS-compliant systems or DoD Cloud SRG
" Relevant certifications (CKA, CKS, cloud provider certs, Security+)

How We Work
" We value simplicity, transparency, and collaboration
" We believe in blameless culture and learning from incidents
" We focus on building tools and platforms that empower other engineers
" We balance reliability, security, and developer experience not one at the expense of the others

What Success Looks Like
" Our platform is reliable, scalable, and easy to operate
" Engineers can deploy confidently in high-compliance environments
" Observability provides clear, actionable insights
" Operational overhead is minimized through automation
" Compliance requirements are met seamlessly as part of the platform
Years of Experience: 14.00 Years of Experience

About the Company

A

Apolis

Since 1996, RJT has provided successful SAP, Oracle, and IT consulting solutions and staffing services to clients around the world. The new Apolis brings you the same personalized service fortified with a greater array of IT solutions, global expertise, and cost-management strategies.

We are a global IT consultancy that seamlessly integrates experts and leading-edge solutions into your organization so you can focus on what really matters.

COMPANY SIZE
500 to 999 employees
INDUSTRY
Computer/IT Services
EMPLOYEE BENEFITS
Paid Sick Days, Employee Referral Program, Employee Events, Retirement / Pension Plans
WEBSITE
https://www.apolisrises.com/