Job Title: DevOps / Platform Reliability Engineer
Location: Tampa, FL / Atlanta, GA / Dallas, TX / Richardson, TX 75082
Contract Duration: 6 Months
Work Mode: Hybrid Client
Experience: 6 8 Years
Job Summary
We are looking for an experienced DevOps / Platform Reliability Engineer to build, automate, and maintain highly available and scalable cloud-native platforms.
The ideal candidate should have strong hands-on experience with Docker, Amazon EKS, Grafana, Git/Version Control, Performance Tuning, and System Design. The role will focus on platform reliability, monitoring, observability, automation, scalability, and operational excellence.
Must-Have Skills
- Docker
- Amazon EKS
- Grafana
- Git / Version Control
- Performance Tuning
- System Design
Key Responsibilities
- Design, build, deploy, and maintain highly available cloud-native platforms.
- Manage and support Docker and Kubernetes-based environments, including Amazon EKS.
- Develop and maintain monitoring and observability solutions using Grafana.
- Implement and manage CI/CD pipelines and deployment automation.
- Use Git and other version control systems for source code and infrastructure management.
- Perform system and application performance tuning to improve reliability and scalability.
- Design resilient and scalable platform architectures.
- Automate infrastructure, deployments, and operational processes.
- Monitor system health, troubleshoot issues, and support production environments.
- Implement SRE practices focused on reliability, availability, scalability, and operational efficiency.
- Collaborate with development, infrastructure, and application teams to improve platform performance and stability.
Preferred Skills
- Kubernetes
- AWS Cloud Services
- CI/CD
- Infrastructure Automation
- Monitoring & Observability
- SRE Practices
- Cloud-Native Architecture
- Troubleshooting and Incident Management