We are looking for an SRE / Cloud Operations Engineer with 1-2 years of hands-on experience supporting cloud-based applications. The person will help maintain reliable, secure, and scalable platforms across AWS, GCP, and Azure, with a strong focus on Kubernetes-based workloads.
The role involves monitoring services, troubleshooting production issues, supporting deployments, automating operational tasks, and collaborating with development and platform teams to improve availability and performance.
Team Size
The candidate will join the SRE team, partnering closely with application developers, security teams, and cloud infrastructure teams. The team owns platform availability, observability, incident response, deployment reliability, and operational automation.
Candidate Experience
- 1-2 years of experience in SRE, DevOps, Cloud Operations, Systems Engineering, or Infrastructure Support.
- Hands-on experience working with Kubernetes, including deploying and troubleshooting pods, deployments, services, namespaces, ConfigMaps, and Secrets.
- Exposure to supporting production applications and handling incidents or service requests.
- Experience working in at least one public cloud: AWS, GCP, or Azure.
- Hands-on exposure to Linux administration, scripting, monitoring, containers, and Docker.
- Basic understanding of CI/CD pipelines and software deployment practices.
Skills (Non-negotiables & good to have)
- Hands-on experience with at least one cloud platform: AWS, GCP, or Azure.
- Basic Kubernetes knowledge: pods, deployments, services, namespaces, ConfigMaps, Secrets, logs, and troubleshooting.
- Docker/container fundamentals.
- Linux troubleshooting and administration skills.
- Scripting knowledge in Bash or Python.
- Understanding of networking basics: DNS, HTTP/HTTPS, TCP/IP, load balancers, security groups/firewalls.
- Familiarity with monitoring and alerting concepts, such as metrics, logs, dashboards, and incident response.
- Good communication and problem-solving skills.