Should have experience in Incident Management and aware of ITIL framework
Scope of Work
Provide support for client reported user issues.
Lead Incident responses, perform RCA and provide on-call support
Automation and tooling using scripting (shell, python etc) .
Track recurring issue, contribute to Problem Management.
Embed Reliability practices & support production releases.
Escalate unresolved technical problems to Engineering team.
Maintain records of support interactions and resolution runbooks
Educate & train L1 team on common issue and help in Shift left initiative
Site Reliability Engineer- W2 Role*
Technical proficiency: Strong Proficiency in Java, Strong understanding of Database concepts (Oracle, SQL, Dynamo DB etc.) Industry standard SRE Tools like Prometheus, Grafana, Data Dog Etc
Good to have skills: Cloud Concepts / AWS, Terraform, MongoDB, Spring Boot Framework, Kafka, RESTful API, Camunda