Cloud Site Reliability Engineer - DCS Cloud

Beijing ByteDance Technology Co Ltd

Seattle, WA

JOB DETAILS
SKILLS
Amazon Web Services (AWS), Analysis Skills, Automation, Blog, C++ Programming Language, CUDA (Compute Unified Device Architecture), Capacity Management, Change Management, Cloud Computing, Communication Skills, Computer Networks, Computer Science, Corporate Compliance, Cost Control, Customer Support/Service, DevOps, Distributed Control Systems (DCS), Docker, Ecosystems, GCP (Good Clinical Practices), GPU (Graphics Processing Unit), Identify Issues, Incident Response, Information/Data Security (InfoSec), K Virtual Machine (KVM), Large-Scale Systems, Linux Operating System, Machine Tool, Microsoft Windows Azure, Network Operations Center, Network Performance/Analysis, On Call, Open Source, Operating Systems, Patents, Private Cloud, Process Improvement, Product Lifecycle, Production Systems, Programming Languages, Public Cloud, Python Programming/Scripting Language, Regulatory Compliance, Reliability Engineering, Root Cause Analysis, Software Engineering, Stress Testing, System Operations, Team Player, Technical Operations, Testing, Topology, Vehicle Fleets, Virtualization
LOCATION
Seattle, WA
POSTED
30+ days ago

Our Infrastructure Engineering team supports the companys fast growth by building and operating hyper-scale datacenters, managing the life cycle of server fleet, providing cloud solutions, and developing various infrastructure services and making sure they are scalable and are reliable.

Responsibilities - What Youll Do

  • Design, build, scale, and operate ByteDance's global infrastructure, including large-scale systems spanning public and private clouds.
  • Develop tools, automation frameworks, visualizations, and monitoring systems to streamline operations and drive optimization of global infrastructure.
  • Create, manage, and standardize cloud AMIs/images for use across multiple environments, ensuring strict alignment with the companys global compliance standards.
  • Thrive in a fast-paced environment, engaging in technical operations and on-call rotations to address incidents related to cloud, OS, network, performance, and reliability.
  • Drive improvements across the entire infrastructure lifecycle, from ideation and design through development, deployment, user support, and continuous refinement.

Minimum Qualifications

  • Bachelor's degree or above in Computer Science, Software Engineering, Information Security, or a related field.
  • 2+ years of experience in Linux operations, SRE, or DevOps; experience operating large-scale production environments is a strong plus.
  • Proficient in at least one programming language such as Go, Python, or C++, with solid engineering capabilities in platform development, system tooling, and automation.
  • Strong computer science fundamentals, with deep understanding of Linux OS principles, computer networks, storage systems, GPU systems, and databases, along with systematic troubleshooting and root-cause analysis skills.
  • Familiar with core reliability practices, including monitoring and alerting, capacity management, change management, canary/gray releases, incident response, and postmortem processes.
  • Strong communication and collaboration skills, with the ability to proactively identify problems, drive cross-team execution, and demonstrate strong ownership and results-oriented mindset.

Preferred Qualifications

  • Hands-on experience operating public cloud platforms, or deep familiarity with major cloud providers such as OCI, AWS, Azure, GCP, etc, including understanding of their underlying mechanisms.
  • Experience with large-scale cloud host delivery, image/AMI systems, resource scheduling, network adaptation, and virtualization technologies such as KVM/QEMU.
  • Familiar with containers and cloud-native ecosystems, including Docker, Kubernetes, and containerd, with a solid understanding of isolation mechanisms like cgroups and namespaces.
  • Experience maintaining GPU clusters, including drivers, CUDA, MIG, topology awareness, troubleshooting, stress testing, and GPU delivery pipelines.
  • Proven experience in reliability-focused initiatives such as failure drill systems, capacity governance, change governance, observability platforms, and resource cost optimization.
  • Open-source contributions, technical blogs, patents, or technical sharing experience are highly preferred.

About the Company

B

Beijing ByteDance Technology Co Ltd