| Location | Louisville, KY |
| Job Type | Full-time, Employee |
| Salary | $0–$60 Per Year |
| Headquarters | Louisville, KY, US |
Site Reliability Engineer with strong experience in AWS, Kubernetes and Terraform
Location :- Louisville, KY (Hybrid) (on site Tuesday and Thursday )
F2F interview Must
Job Description :-
Key Responsibilities
**Reliability & Observability**
- Define and own Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for restaurant technology systems; use error budgets to balance reliability with velocity.
- Build and maintain monitoring, alerting, and observability platforms that provide meaningful signal - not noise.
- Lead blameless post-incident reviews; drive root cause analysis and ensure permanent corrective actions are implemented.
- Proactively identify reliability risks across the stack before they become incidents.
- Establish and track reliability metrics; report on system health to engineering leadership.
**Architecture & Infrastructure**
- Design and implement scalable, secure, and highly available cloud architectures primarily on AWS, with working knowledge of Azure.
- Architect and manage containerized workloads using Kubernetes (EKS, AKS), including edge Kubernetes deployments in restaurant environments.
- Design and implement serverless and container-based solutions, including AWS Fargate and other managed services.
- Develop and maintain Infrastructure as Code (IaC) using Terraform.
- Own architecture across the full restaurant technology stack - cloud, edge, networking, and device management - not just the cloud layer.
**DevOps, Automation & Tooling**
- Build and optimize CI/CD pipelines using GitLab CI/CD and modern DevOps practices.
- Build internal tools, automations, and middleware integrations that eliminate repetitive operational work.
- Use AI-assisted development to accelerate scripting, troubleshooting, and documentation.
- Champion a culture of engineering solutions over repeated manual fixes - if something is done twice, it should be automated.
**Restaurant & Edge Technology**
- Design and support Kubernetes-based edge systems deployed in restaurant locations.
- Support mobile application deployments and troubleshoot deployment issues across restaurant endpoints.
- Manage and optimize Mobile Device Management (MDM) platforms covering the restaurant device fleet.
- Configure and troubleshoot enterprise networking - primarily switches and restaurant-facing network infrastructure.
- Lead and participate in incident response for restaurant technology systems, including on-call coverage and post-incident review.
- Reduce mean time to detection (MTTD) and mean time to resolution (MTTR) through better tooling, runbooks, and automation.
**Security & Governance**
- Establish and enforce cloud governance, security policies, and architectural standards.
- Implement cloud security best practices: IAM strategy, network segmentation, encryption, and secrets management.
- Conduct security architecture reviews; identify vulnerabilities, misconfigurations, and compliance gaps.
- Integrate security into CI/CD pipelines (DevSecOps - Development, Security, and Operations), including automated scanning, policy validation, and vulnerability management.
**Leadership & Collaboration**
- Collaborate with engineering, DevOps, and security teams to ensure secure-by-design solutions across cloud and restaurant tech.
- Provide technical leadership and mentorship to engineering teams.
- Create and maintain documentation, runbooks, and architectural decision records.
- Continuously evaluate emerging technologies and recommend improvements.
Required Qualifications
- 6+ years in IT infrastructure, with 3+ years focused on site reliability engineering, cloud architecture, or platform engineering.
- Hands-on experience with AWS (VPC, EC2, ECS, EKS, Fargate, Lambda, IAM, RDS, S3).
- Working experience with Microsoft Azure.
- Strong expertise in Kubernetes and container orchestration, including edge or distributed deployments.
- Experience with GitLab CI/CD and CI/CD pipeline design.
- Solid experience with Terraform for infrastructure provisioning.
- Experience with enterprise networking - switch configuration, VLANs, network troubleshooting.
- Familiarity with Mobile Device Management (MDM) platforms.
- Experience with automation and scripting (Python, Bash, Go, or equivalent).
- Proven ability to build internal tooling and API integrations, not just configure managed services.
- Experience defining and operating against SLOs, SLIs, and error budgets.
- Comfortable working in Linux command-line environments; Windows familiarity a plus where restaurant endpoints require it.
Preferred Qualifications
- Experience designing serverless architectures (AWS Lambda, Fargate, API Gateway, EventBridge).
- Experience with DevSecOps tooling (SAST - Static Application Security Testing, DAST - Dynamic Application Security Testing, container scanning, IaC scanning).
- Familiarity with security frameworks (CIS, NIST, ISO 27001, SOC 2).
- AWS and/or Azure certifications.
- Experience with monitoring and observability tools (CloudWatch, Prometheus, Grafana, or SIEM solutions).
- Background in restaurant, retail, or distributed edge technology environments.
- Experience using AI-assisted development tools for scripting, troubleshooting, and documentation.
Key Competencies
- Generalist mindset - comfortable moving between cloud, edge, networking, and device management in the same week.
- Reliability-first thinking - treats toil reduction, error budgets, and post-incident learning as core engineering disciplines, not afterthoughts.
- Bias toward permanent fixes and automation over repeated manual intervention.
- Ability to balance strategic architecture with hands-on execution.
- Strong communication and cross-functional collaboration skills.
- Curious, proactive, and detail-oriented approach to systems design and operations.