What You'll Do: Pick up well-scoped platform work across the systems above and ship it end to end Build and maintain GitLab CI/CD pipelines and GitOps delivery through ArgoCD Write and review infrastructure-as-code with Terraform for AWS resources Add and tune observability - Prometheus metrics, Grafana dashboards, Alertmanager rules Help fulfill and then automate recurring requests (ingress, DNS, service accounts, IAM roles) Help investigate platform incidents and document what you learn Learn to recognize common failure modes (IP exhaustion, resource limits, reconciliation lag) and escalate them early Technologies We Use: Kubernetes / EKS (multi-cluster, multi-region), Karpenter, cert-manager, external-dns Prometheus, Alertmanager, Grafana, OpenTelemetry (and long-term storage / sharding for Prometheus) AWS networking (VPC, VPC peering, Transit Gateway, Route53, NAT Gateway, security groups, subnet/CIDR design across accounts and regions) Terraform, AWS (IAM, EKS, S3, EBS) ArgoCD, GitLab CI/CD, Nexus (artifact registry), Docker, container image build pipelines Vault, OpenCost Engineering Practices We Employ: Agile software development Infrastructure as Code (IaC) Pair programming Test-Driven Development (TDD) Continuous Delivery Qualifications What We Look For: 1-2 years of experience in software or infrastructure engineering Bachelor's degree or equivalent experience Comfortable in the terminal and reading YAML, Terraform, or similar declarative config Some exposure to cloud (AWS or equivalent) and containers - production Kubernetes experience is a plus, not a requirement Genuine interest in operating real systems - especially observability - and growing deep in them Willing to learn how systems fail and to ask questions when something looks off Effective communication; enjoys pair programming and code review Nice to Have: Hands-on exposure to Kubernetes, Terraform, or GitLab/GitHub CI/CD Any experience with Prometheus/Grafana or another metrics stack Scripting in Python, Bash, or Go What Success Looks Like: You reliably deliver well-scoped platform work end to end with decreasing oversight You're building genuine depth in at least one system we own, not just familiarity with the tools Other engineers find the workflows you touch easier and more predictable to use Additional information This is a hybrid role requiring 3 days a week in office. You'll contribute to operating them and build depth in them over time, starting with guidance from senior engineers: Observability & monitoring - Prometheus, Alertmanager, Grafana, and OpenTelemetry across production regions.