Job Title: Site Reliability Engineer (SRE)
Location: Urbandale, IA (Open to relocation)
Duration: 12 Months
Role Summary:
Lead application reliability, resiliency, security, and incident management for large-scale cloud environments. Drive automation, monitoring, performance optimization, and recovery solutions.
Key Responsibilities:
- Lead incident response, recovery, and post-mortem analysis
- Design and automate failure detection, recovery, and resiliency solutions
- Define and manage Service Level Objectives (SLOs)
- Improve application performance, security, quality, and cost efficiency
- Develop monitoring, metrics, dashboards, and operational playbooks
- Collaborate with engineering teams to ensure reliable production systems
- Act as a technical advisor for complex reliability and infrastructure challenges
Required Skills:
- Strong Site Reliability Engineering (SRE) or DevOps experience
- AWS Cloud Services, Kubernetes, Terraform, Datadog
- Incident Management, RCA, SLO/SLI implementation
- Automation and Infrastructure as Code (IaC)
- Programming experience in one or more: Java, Scala, JavaScript, .NET, Go, or Python
- Experience with application monitoring, observability, security, and performance tuning