Alibaba Cloud-Site Reliability Engineer-Sunnyvale

Alibaba Group Holding Ltd

  • Sunnyvale, CA
  • 11 days ago
  • $104,400–$171,000 Per Year

Highlights

The mission of the Cloud Intelligence Group SRE (Site Reliability Engineering) Team is to ensure the stability of production environments, enterprise-grade cloud data reliability, and service continuity for the Cloud Intelligence Group. Our greatest challenge lies in guaranteeing uninterrupted business operations for cloud-based customers and achieving availability that exceeds 99.99%.

Numbers & Facts

LocationSunnyvale, CA
Salary$104,400–$171,000 Per Year

Description

The mission of the Cloud Intelligence Group SRE (Site Reliability Engineering) Team is to ensure the stability of production environments, enterprise-grade cloud data reliability, and service continuity for the Cloud Intelligence Group. Our greatest challenge lies in guaranteeing uninterrupted business operations for cloud-based customers and achieving availability that exceeds 99.99%.

Objectives of the Cloud Intelligence Group SRE Team

Our goal is to establish a systematic stability assurance framework that integrates technology and management, including but not limited to:

  1. Developing stability standards and metrics

    • Covering robust architecture, R&D quality, release management, production environment operations, and more.
    • Embedding stability into Alibaba Cloud's technical R&D system.
    • <
    • Driving major stability governance campaigns

      • Initiatives such as full-stack disaster recovery, phased change rollout, the 1-5-10 emergency response mechanism (1-minute alerting, 5-minute triage, 10-minute recovery), and financial-loss prevention.
      • Rapidly and continuously mitigating stability risks.
      • <
      • Building a stability-focused technical platform

        • Platform capabilities for unattended change management, red/blue team drills, emergency collaboration, risk and vulnerability inspection, and monitoring/alerting.
        • Simplifying stability engineering through automation and tooling.
        • <
        • Executing production incident management

          • Emergency response, cross-team coordination, root cause analysis, rapid recovery, and post-incident reviews to drive systemic improvements.
          • <
          • Ensuring stability for large-scale customer events

            • Technical and operational support for critical activities such as Olympics and customer business peak periods.
            • <
            • On-call responsibilities

              • Responding to customer issues within Service Level Agreement (SLA) timeframes, resolving problems proactively, and enhancing customer experience.
              • <

                Responsibilities The objective of the Cloud Intelligence Group's SRE team is to establish a systematic stability assurance framework that integrates technology and management, including but not limited to:

                1. Daily operations and maintenance of applications, databases, and middleware, as well as troubleshooting and answering customer inquiries;
                2. Collaborating with R&D to develop critical support plans based on customer business requirements during peak periods, including preparation during the standby period, on-duty support during critical periods, and post-standby review;Minimum qualification:

                  • A degree in Computer Science or related field
                  • 3 years of experience as a Site Reliability Engineer (SRE) or above.
                  • Proficiency with Linux environments or cloud infrastructure
                  • Exceptional system diagnostic and problem-solving skills
                  • Strong teamwork spirit and ability to work well under pressure
                  • <

                    Preferred qualification:

                    • In-depth understanding of Kubernetes or monitoring systems
                    • Expertise in programming languages such as Golang, Python, or Java
                    • Experience in designing and implementing distributed systems

                    The pay range for this position at commencement of employment is expected to be between $104,400 and $171,000/year. However, base pay offered may vary depending on multiple individualized factors, including market location, job-related knowledge, skills, and experience.

                    If hired, employee will be in an "at-will position" and the Company reserves the right to modify base salary (as well as any other discretionary payment or compensation program) at any time, including for reasons related to individual performance, Company or individual department/team performance, and market factors.

Similar Jobs

See more jobs