Mandatory Skills: 1. Java Spring boot 2. Apache Kafka 3. Dev Ops 4. CI/CD automation
Years of experience required: 8-10
Job Description: Automation & Efficiency · Automate the top 5 high-volume support and request types · Build self-service and agent-driven solutions to reduce manual work · Harden operational workflows for consistency, auditability, and resilience · Implement auto-retry and backoff for recurring failure patterns
Reliability Engineering · Define and manage Service Level Objectives (SLOs) for critical services and batch processes · Apply error budget concepts to guide reliability and release decisions · Improve batch reliability through standardized recovery patterns and monitoring
Observability & Metrics · Build reliability dashboards tracking incidents, repeat issues, failure rates, and automation coverage · Improve operational reporting and visibility across incidents, problems, and changes
Runbooks & Self-Service · Develop and expand runbooks for key production scenarios · Convert runbooks into automated remediation workflows · Enable self-service for repeat operational requests · Drive conversion of repeat incidents into permanent fixes and known problems
Self-Healing & Intelligent Operations · Implement self-healing capabilities to minimize manual intervention · Optimize alerting systems (e.g., Moogsoft) to reduce noise and improve signal quality · Leverage automation and AI to resolve recurring issues with minimal human involvement