The IT storage team manages petabytes of on-premises, clustered POSIX storage for AI modeling and is developing the next-generation storage solutions. Responsibilities: Onboard new customers onto our caching solution, helping them update their applications and verifying successful integration.
Numbers & Facts
Location
Foster City, CA
Salary
$75–$85.19 Per Hour
Description
JOB TITLE: Site Reliability Engineer - Data & Caching Systems LOCATION: Foster City, CA (5 days onsite) DURATION: 6 months RATE RANGE: $75 - $85/hr.
Job Description:
AI systems require efficient data handling at scale. We are seeking an individual passionate about optimizing software, hardware, and data transfer to support our AI initiatives across the country. The IT storage team manages petabytes of on-premises, clustered POSIX storage for AI modeling and is developing the next-generation storage solutions. This includes building a geo-distributed file system/data lake to support autonomous robotaxis operations nationally and globally. Our initial focus is on a high-performance caching system significantly outperforming AWS S3.
Responsibilities:
Onboard new customers onto our caching solution, helping them update their applications and verifying successful integration.
Help scale and operate our fleet of caching servers throughout the world by leveraging automation. (Terraform, Spacelift, SaltStack, Ansible, etc.)
Help improve observability and monitoring of the caching system. (OTEL, Grafana, Zabbix, etc.)
Improve CI/CD workflows and help with the software release cycle. (GitHub, Buildkite, Bamboo, Bazel, etc.)
Improve user-facing documentation and operational runbooks for the caching system.
Report bugs to software developers. (Jira, Confluence, etc.)
Optionally help fix bugs in the caching system. (Rust)
Qualifications:
2+ years of experience operating distributed services.
Excellent written and verbal communication skills.
Able to organize and work on several long-running processes at one time.
Ability to troubleshoot complex problems between the system and customer s application.
Expertise in automation and Infrastructure as Code (IoC). (Terraform, Spacelift, Ansible, SaltStack, etc.)
Expertise in monitoring, alerting, and observability platforms. (OTEL, Grafana, Zabbix, OpsGenie, Incident.io, etc.)
Motivated to learn new technologies and think differently.
Bonus Qualifications:
Experience with Python, C++, Rust, or similar programming languages.
Experience with CI/CD workflows. (Buildkite, Bamboo, GitHub, etc.)
BENEFITS SUMMARY: Individual compensation is determined by skills, qualifications, experience, and location. Compensation details listed in this posting reflect the base hourly rate or annual salary only, unless otherwise stated. In addition to base compensation, full-time roles are eligible for Medical, Dental, Vision, Commuter and 401K benefits with company matching. IND123