Sr. SRE engineer

Spark Tek Inc

  • Owings MIlls, GA
  • 2 days ago

    Highlights

    Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS. •     Triage incidents by identifying root causes, distinguishing infrastructure issues from application defects, and restoring service within defined SLAs.

    Numbers & Facts

    LocationOwings MIlls, GA

    Description

    Role: Sr. SRE engineer  
    Location: Atlanta, GA       

    Qualifications
    •     Strong experience supporting production systems hosted on AWS, including EC2, VPC, ALB/NLB, RDS, Lambda, and EKS.
    •     Hands-on experience with incident management and 24/7 production support models.
    •     Proficiency with monitoring and observability tools such as CloudWatch, Dynatrace, and Quantum Metric.
    •     Experience building and maintaining monitoring dashboards.
    •     Strong troubleshooting skills across infrastructure, networking, and application layers.
    •     Working knowledge of CI/CD pipelines and AWS deployment processes.
    •     Experience working with databases and Unix/Linux environments.

    Key Responsibilities
    Incident Management and Production Support
    •     Provide Level 1 and Level 2 support for production incidents across AWS-hosted applications and infrastructure.
    •     Triage incidents by identifying root causes, distinguishing infrastructure issues from application defects, and restoring service within defined SLAs.
    •     Escalate code-level defects to development teams with clear diagnostics, supporting logs, and impact assessments.
    •     Participate in on-call rotations, major incident bridges, and post-incident reviews.
    •     Investigate application defects, configuration issues, and infrastructure anomalies reported through monitoring tools or user incidents.

    Monitoring and Operational Health
    •     Perform regular health checks across applications, infrastructure, and AWS services.
    •     Monitor system health using CloudWatch, Dynatrace, Quantum Metric, and ThousandEyes.
    •     Respond proactively to alerts related to resource utilization, latency, errors, and availability.
    •     Maintain and improve monitoring and observability dashboards."               

    Similar Jobs

    See more jobs