Lead Software Engineer, Middleware Reliability Engineering Visa Technology and Operations LLCLead Software Engineer, Middleware Reliability EngineeringFoster City, CA$192,300–$307,600 / yearFamiliarity with AI/ML frameworks and chatbot integrations for operational automation Middleware Expertise • Understanding of middleware technologies (Message Queues, Service Bus, API • Gateways) • Experience with application servers (Tomcat, JBoss, WebSphere) • Knowledge of integration platforms and their deployment patterns • Proficiency in troubleshooting and performance optimization Education • Bachelor's or Master's degree in Computer Science or related field, or equivalent experience • We value hands-on experience and continuous learning over specific degrees. • Solid understanding of Linux/Unix systems, networking protocols, certificate management, secret management, system design, cloud platforms (AWS, • Azure, GCP), and containerization (Kubernetes, Docker • Proficiency with monitoring tools (Prometheus, Grafana, Datadog, etc.), logging systems (ELK stack, Splunk), and tracing tools (Jaeger, Zipkin).
Supplier Quality Engineer Raas Info Solutions Pvt LtdSupplier Quality EngineerFremont, CA$50–$54.99 / hourContractorThe role involves evaluating thermal/mechanical designs, supplier manufacturing capability, reliability test reports, and process quality to ensure products meet performance, manufacturability, serviceability, and compliance requirements. Job Summary: Responsible for supplier quality management of data center rack systems and advanced cooling solutions.
DevOps Platform Development Engineer (Bilingual Mandarin) ComriseDevOps Platform Development Engineer (Bilingual Mandarin)Palo Alto, CA$150,000–$240,000 / yearFull timeCollaborate closely with platform engineering teams in China as well as overseas infrastructure, networking, security, and business teams to manage implementation schedules, dependencies, and risks, while providing feedback to improve platform capabilities for international deployment scenarios. Participate in the day-to-day operations and incident response of overseas production environments, quickly troubleshoot deployment, infrastructure, and platform integration issues, and drive root cause analysis and long-term corrective actions.
NewSupplier Quality Engineer JobotSupplier Quality EngineerSunnyvale, CA$50–$85 / hourThe ideal candidate will have experience supporting suppliers in electro-mechanical and electronics manufacturing environments, including printed circuit board assemblies (PCBAs), manufactured components, and complex hardware assemblies. Information collected and processed as part of your Jobot candidate profile, and any job applications, resumes, or other information you choose to submit is subject to Jobot's Privacy Policy, as well as the Jobot California Worker Privacy Notice and Jobot Notice Regarding Automated Employment Decision Tools which are available at jobot.com/legal.
NewSenior Sales Engineer JobotSenior Sales EngineerPleasanton, CA$300,000–$330,000 / yearThe founding team includes senior leaders and technical pioneers from industry giants like AWS, Cisco, VMware, and Gigamon - holding dozens of patents and having built critical systems at some of the most respected tech companies in the industry. Information collected and processed as part of your Jobot candidate profile, and any job applications, resumes, or other information you choose to submit is subject to Jobot's Privacy Policy, as well as the Jobot California Worker Privacy Notice and Jobot Notice Regarding Automated Employment Decision Tools which are available at jobot.com/legal.
Staff Data Engineer Visa Technology and Operations LLCStaff Data EngineerFoster City, CA$146,200–$233,700 / yearVisa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid. At Visa, you'll have the opportunity to create impact at scale — tackling meaningful challenges, growing your skills and seeing your contributions impact lives around the world.
NewData Engineer - Sr. Consultant level Visa Technology and Operations LLCData Engineer - Sr. Consultant levelFoster City, CA$169,100–$270,800 / yearVisa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid. Masters, MBA, JD, MD) or 2 years of work experience with a PhD Preferred Qualifications • 9 or more years of relevant work experience with a Bachelor Degree or 7 or more relevant years of experience with an Advanced Degree (e.g.
Lead Software Engineer Visa Technology and Operations LLCLead Software EngineerFoster City, CA$192,300–$307,600 / yearVisa is a world leader in payments technology, facilitating transactions between consumers, merchants, financial institutions and government entities across more than 200 countries and territories, dedicated to uplifting everyone, everywhere by being the best way to pay and be paid. Essential Functions Provide deep technical leadership within Shared Services Product Development, with a strong understanding of how shared platforms enable downstream product innovation.
Network DevOps Engineer Epitec StaffingNetwork DevOps EngineerPalo Alto, CASummary Our client is seeking a hands-on Network DevOps Engineer to support network automation, cloud infrastructure deployment, and multi-cloud connectivity initiatives across a large-scale enterprise environment. The ideal candidate enjoys solving complex infrastructure challenges, building automated solutions, and supporting highly available environments across multiple cloud platforms.
Database admin/Database Engineer Pinnacle Technical ResourcesDatabase admin/Database EngineerSunnyvale, California$75–$80 / hourContractorThe specific compensation for this position will be determined by several factors, including the scope, complexity, and location of the role, as well as the cost of labor in the market; the skills, education, training, credentials, and experience of the candidate; and other conditions of employment. Successfully placed or hired candidates would only be asked for banking details after accepting an offer from us during our official onboarding processes as part of payroll setup.
Mechanical Engineer III Pinnacle Technical ResourcesMechanical Engineer IIICupertino,, California$60–$65 / hourContractorThe specific compensation for this position will be determined by a number of factors, including the scope, complexity and location of the role as well as the cost of labor in the market; the skills, education, training, credentials and experience of the candidate; and other conditions of employment. Provide design and tradeoff guidance to Product Design teams through early-stage simulations, particularly for fast charging and high-power functions.
FullStack Data / ADK Engineer CYNET SYSTEMSFullStack Data / ADK EngineerSanta Clara, CA$50–$55 / hourTemporaryContractorPart timeAs a nationally and locally certified Minority Business Enterprise (MBE), Cynet Systems is committed to helping organizations build high-performing teams while empowering professionals to grow rewarding careers. We deliver agile, scalable talent solutions across IT, engineering, life sciences, clinical, and professional staffing, powered by a high-performing recruitment engine operating across North America and Asia.
Test Engineer III Pinnacle Technical ResourcesTest Engineer IIICupertino,, California$65–$70 / hourContractorThe specific compensation for this position will be determined by a number of factors, including the scope, complexity and location of the role as well as the cost of labor in the market; the skills, education, training, credentials and experience of the candidate; and other conditions of employment. Successfully placed or hired candidates would only be asked for banking details after accepting an offer from us during our official onboarding processes as part of payroll setup.
CI/CD and Automation Infrastructure Engineer Tanisha SystemsCI/CD and Automation Infrastructure EngineerCupertino, CAFull timeRole Purpose: Build and maintain scalable CI, automation and dashboard infrastructure that enables reliable test execution across media, platform and device validation workflows. • Build and maintain XCTest/XCUITest frameworks, Jenkins pipelines and supporting test infrastructure.
Data Engineer II Pinnacle Technical ResourcesData Engineer IISunnyvale,, CaliforniaContractorThe specific compensation for this position will be determined by several factors, including the scope, complexity, and location of the role, as well as the cost of labor in the market; the skills, education, training, credentials, and experience of the candidate; and other conditions of employment. The ideal candidate will be responsible for designing, building, and maintaining scalable data pipelines and systems to support our organization's data needs.
NewSenior AWS DevOps Engineer JobotSenior AWS DevOps EngineerMilpitas, CA$155,000–$165,000 / yearInformation collected and processed as part of your Jobot candidate profile, and any job applications, resumes, or other information you choose to submit is subject to Jobot's Privacy Policy, as well as the Jobot California Worker Privacy Notice and Jobot Notice Regarding Automated Employment Decision Tools which are available at jobot.com/legal. Our platform integrates with employers and payroll systems to provide real-time wage access and financial wellness tools that reduce reliance on payday loans, overdraft fees, and other costly alternatives.
NewSenior Forward Deployed Engineer (DevOps/SRE) JobotSenior Forward Deployed Engineer (DevOps/SRE)Pleasanton, CA$300,000–$350,000 / yearThe founding team includes senior leaders and technical pioneers from industry giants like AWS, Cisco, VMware, and Gigamon - holding dozens of patents and having built critical systems at some of the most respected tech companies in the industry. 6+ years of experience in Site Reliability Engineering, DevOps, Platform Engineering, or similar infrastructure-focused roles, including technical leadership or end-to-end customer delivery.
Quality & Reliability Engineer, Trainium Manufacturing, Quality & Reliability Amazon.com IncQuality & Reliability Engineer, Trainium Manufacturing, Quality & ReliabilityCupertino, CAThe Trainium Manufacturing, Quality and Reliability (MQR) Team is part of AWS Annapurna Labs focused on Machine Learning products that designs cutting AI platforms for the world's largest Cloud Services provider. Lead identifying and validating product/component risks and work with design teams to mitigate them and define the test methodology and test coverage to assure product reliability.
Cell Reliability Engineer, Cell Quality & Reliability Tesla IncCell Reliability Engineer, Cell Quality & ReliabilityFremont, CA$109,600–$164,400 / yearAnalyze Field Reliability/Quality data for cell related failures and failure modes to predict expected failure rates, affected populations, verify effectiveness of the corrective actions at Tesla and at suppliers. Perform risk assessment with process engineers for proper documentation of control plan, pfmea/dfmea, ppap, IQC/OQC metrics along with supporting key production metrics to achieve yield, OEE.
Hardware Reliability Engineer - Mac System Reliability Apple IncHardware Reliability Engineer - Mac System ReliabilityCupertino, CABachelor's Degree in technical field ( mechanical engineering, materials engineering, electrical engineering, physics or related fields) Collaborative attitude and ability to cross-functionally work effectively Good written and verbal English communication skills. A proactive, can-do attitude, with a passion for working alongside an amazing team and innovative productsIdentifying high-risk failure modes early in the design process and working closely with design engineering teams to mitigate risks.
Reliability Engineer, Mechanical Systems, NA Vantage Data Centers Management Co LLCReliability Engineer, Mechanical Systems, NASanta Clara, CAFor each of the major systems Electrical, Mechanical, and Controls, the Reliability Engineering team is responsible for ensuring success in the commissioning stages of new construction, evaluating and improving the reliability and performance of existing critical infrastructure, sustaining equipment operational availability through maintenance program design, providing ongoing technical support to the Site Operations Teams, as well as providing systems reliability and maintainability feedback to the Design Engineering teams for future design considerations. Developing and operating across North America, EMEA and Asia Pacific, Vantage has evolved data center design in innovative ways to deliver dramatic gains in reliability, efficiency and sustainability in flexible environments that can scale as quickly as the market demands.
Sr. Site Reliability Engineer Starlink Space Exploration Technologies CorpSr. Site Reliability Engineer StarlinkPalo Alto, CA$165,000–$280,000 / yearBASIC QUALIFICATIONS: Bachelor''s degree in computer science, engineering, math, or scientific discipline and 5 years of software development experience; OR 7+ years of professional experience building software with site reliability or DevOps in lieu of a degree. ITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State.
Reliability Engineer Yantran LLCReliability EngineerCupertino, CAThe engineer will be responsible for leading and executing reliability tests on new devices and supplier technologies, development of new test procedures to quantify the reliability of a design, and failure analysis resulting from these tests. Key Qualifications: Experience performing reliability testing including Mechanical Stress Tests, Shock/Drop/Vibration Testing, Environmental Testing.
Senior Site Reliability Engineer- Palo Alto, the US KodySenior Site Reliability Engineer- Palo Alto, the USPalo Alto, CAYou will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America. Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements.
Senior Site Reliability Engineer- Sunnyvale, CA, the US KodySenior Site Reliability Engineer- Sunnyvale, CA, the USSunnyvale, CAYou will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America. Data-driven mindset with the ability to leverage metrics, telemetry, trends, and service-level indicators to prioritize reliability investments and operational improvements.
Hardware Reliability Engineer Google LLCHardware Reliability EngineerMountain View, CADrive the failure analysis process for all failures discovered during reliability testing and lead the failure analysis process with internal and external cross-functional teams for all failures discovered during reliability testing. Proficiency in statistical data analysis including DOE, significance testing, sample sizes and confidence level using statistical tools such as JMP, Python.
Sr. Solar Field Reliability Engineer Tesla IncSr. Solar Field Reliability EngineerPalo Alto, CA$100,000–$228,000 / yearYou will be the technical authority on solar field performance: parsing massive datasets to uncover subtle failure signatures, diagnosing complex failure modes that span electrical, thermal, and mechanical domains, and driving corrective actions that improve reliability for hundreds of thousands of customers. Produce Field Quality Reports: Author monthly/quarterly reliability reports for leadership, summarizing active investigations, failure rate trends, countermeasure efficacy, and cost projections.
Nuclear Systems Reliability Engineer Lawrence Livermore National LaboratoryNuclear Systems Reliability EngineerLivermore, CA$146,340–$222,564 / yearCollaborate with team members to support and contribute to the analysis and testing of programmatic and nuclear facility equipment, including equipment vibration assessment, thermography, equipment operational data trending, deriving post-maintenance testing plans, writing testing procedures, and performing safety calculations. Contribute to and participate in the development of a predictive reliability program to ensure operational readiness and minimize downtime for critical systems, including power distribution, control systems, scientific tools & diagnostics, containment ventilation, cryogenic liquids, compressed gases, gloveboxes, and fire suppression.
Sr. Site Reliability Engineer Illumio IncSr. Site Reliability EngineerSunnyvale, CAYour Impact: We are looking for an experienced Senior Site Reliability Engineer (SRE) with a strong background in AWS & Azure cloud platforms to play a key role in ensuring the reliability, scalability, and performance of our cloud-based systems and applications. Powered by the Illumio AI Security Graph, our breach containment platform identifies and contains threats across hybrid multi-cloud environments stopping the spread of attacks before they become disasters.
Senior Site Reliability Engineer- Palo Alto, The US KodyPay LtdSenior Site Reliability Engineer- Palo Alto, The USPalo Alto, CAYou will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America. Cross-Border Collaboration: Act as a key technical bridge between our US operations and international engineering hubs, leveraging bilingual communication to streamline complex technical alignment.
Senior Site Reliability Engineer- Sunnyvale, CA, The US KodyPay LtdSenior Site Reliability Engineer- Sunnyvale, CA, The USSunnyvale, CAYou will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America. Diagnose, triage, mitigate, and coordinate resolution of production incidents across payment services, Kubernetes platforms, databases, messaging systems, and cloud infrastructure.
Site Reliability Engineer Akkodis Group AG.Site Reliability EngineerSunnyvale, CA$64–$68 / hourSite Reliability Engineer Job Responsibilities include: Design, deploy, and maintain highly available and scalable cloud infrastructure on AWS using services such as EC2, EKS, S3, RDS, Lambda, IAM, and VPC. The ideal candidate with experience maintaining highly available, scalable cloud infrastructure and driving operational excellence through automation, monitoring, and incident management.
Reliability Engineer, Mechanical Systems, NA Vantage Data CentersReliability Engineer, Mechanical Systems, NASanta Clara, CaliforniaFor each of the major systems Electrical, Mechanical, and Controls, the Reliability Engineering team is responsible for ensuring success in the commissioning stages of new construction, evaluating and improving the reliability and performance of existing critical infrastructure, sustaining equipment operational availability through maintenance program design, providing ongoing technical support to the Site Operations Teams, as well as providing systems reliability and maintainability feedback to the Design Engineering teams for future design considerations. Developing and operating across North America, EMEA and Asia Pacific, Vantage has evolved data center design in innovative ways to deliver dramatic gains in reliability, efficiency and sustainability in flexible environments that can scale as quickly as the market demands.
Wireless Module Reliability Engineer Apple IncWireless Module Reliability EngineerSunnyvale, CAAct as the primary technical liaison between internal cross-functional teams and reliability teams at vendors, ensuring alignment on reliability campaigns, determining stress test conditions and approving test hardware. As a Reliability Engineer for the Wireless Design group, you will be responsible for driving all reliability requirements across multiple Wireless vendors.
Site Reliability Engineer (SRE) / Software Engineer (SWE) PhaxisSite Reliability Engineer (SRE) / Software Engineer (SWE)Mountain View, CA$60In this hybrid role, you will be responsible for maintaining and enhancing our production applications, monitoring system alerts and logs, troubleshooting and resolving code issues, and ensuring the reliability and scalability of our infrastructure. The responsibility for this role is to deploy cloud infrastructure, build a robust, secure, and maintainable service through Terraform, and build an efficient developer experience.
NewSenior/Staff Power Electronics Hardware Reliability Engineer Zipline International IncSenior/Staff Power Electronics Hardware Reliability EngineerSouth San Francisco, CA$140,000–$230,000 / yearDeep experience developing and validating safety-critical high-voltage and high-power electrical systems such as grid-connected converters, industrial power supplies, EV charging systems, inverters, high band gap semiconductors, renewable-energy systems, aerospace power systems, or similar equipment. Our customers include the world's largest and most prominent healthcare systems, governments, retailers, restaurants and global businesses who rely on us to save lives, reduce emissions, increase economic opportunity, and provide delivery from point A to point B as fast as possible.
Senior Site Reliability Engineer OutSystemsSenior Site Reliability EngineerMenlo Park, CaliforniaAs an SRE at OutSystems here are your key responsibilities and duties: Lead and onboard services and teams to the reliability tenets; Establish and maintain Service Level Objectives (SLOs) and Service Level Agreements (SLAs); Design and implement scalable, reliable, and secure infrastructure, while ensuring cloud-native best practices; Collaborate with software development teams to ensure systems are resilient (observable, fault-tolerant, recoverable, scalable) and performant; Implement monitoring, alerting, logging, and tracing solutions to detect and respond to incidents; Lead incident response efforts, ensuring quick resolution and minimal downtime, and conduct RCA/post-mortems; Automate every operational task, with a special focus on fast incident detection & recovery; Programming in Python supported by Gen AI tooling to accelerate development of mission critical automation and tools. (CKA, CKAD, CKS certifications are valued); Experience with automation and Infrastructure as Code (IaC) tools, such as AWS CloudFormation, Terraform, Puppet, Chef, Spacelift, etc; Experience with Python, Go, Bash/Shell scripting, or other automation tools/languages; Familiarity with AWS services like EC2, RDS, ELB, CloudFront, Lambda, etc; Proficiency in monitoring and troubleshooting complex distributed systems; Experience with Grafana, ELK stack, Prometheus, or others; Strong understanding of designing resilient and fault-tolerant systems; Expertise in debugging complex distributed systems.
NewSenior Lead Site Reliability Engineer JPMorgan Chase Bank, N.A.Senior Lead Site Reliability EngineerPalo Alto, CAFull timeStrong experience building production-grade RESTful APIs and designing message queue architectures (Kafka, RabbitMQ, SQS) for event-driven systems; and expertise in graph databases (Neo4j, TigerGraph), vector databases (Pinecone, Weaviate, Chroma), and integrating multiple data stores for AI-powered systems. Hands-on experience building AI Agents and autonomous systems with proficiency in AI frameworks (LangChain, LangGraph, AutoGen, CrewAI) and leveraging AI development tools (GitHub Copilot, Claude, etc.) to accelerate development and innovation and Expertise in designing and implementing logging pipelines (Fluentd, Logstash, Vector) and systems for metrics collection, analysis, and distributed tracing.
Senior CockroachDB Database Engineer / Site Reliability Engineer (SRE) Resource Logistics, Inc.Senior CockroachDB Database Engineer / Site Reliability Engineer (SRE)Sunnyvale, CAWe are seeking a highly skilled CockroachDB Database Engineer with strong Site Reliability Engineering (SRE) experience to design, implement, manage, and optimize large-scale distributed database platforms. The role requires close collaboration with development, infrastructure, and platform engineering teams to ensure highly available, resilient, and scalable database services.
Senior Site Reliability Engineer- Palo Alto, The US KodySenior Site Reliability Engineer- Palo Alto, The USPalo Alto, CASenior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America.
Senior Site Reliability Engineer- Sunnyvale, CA, The US KodySenior Site Reliability Engineer- Sunnyvale, CA, The USSunnyvale, CASenior Site Reliability Engineer (Payments Infrastructure) Kody is seeking a Senior Site Reliability Engineer to ensure the reliability, availability, scalability, and operational excellence of our global payment platform. You will own production observability, incident response, service-level management, and cloud infrastructure reliability across mission-critical payment processing systems operating in Europe, Asia, and North America.
Reliability Engineer Hardware Lightmatter IncReliability Engineer HardwareMountain View, CA$142,000–$200,000 / yearYou will enable our products to meet or exceed reliability requirements throughout their lifecycle by developing robust test protocols, analyzing failures, and working with cross-functional teams to drive improvements in design and manufacturing processes. Perform detailed root cause analysis of device failures, and collaborate with cross-functional teams to resolve reliability issues, refine product designs, and improve process robustness.
Sr. Design for Reliability Engineer, Vehicle Systems Rivian Automotive IncSr. Design for Reliability Engineer, Vehicle SystemsPalo Alto, CA$146,900–$183,600 / yearRivian may use your Candidate Personal Data for the purposes of (i) tracking interactions with our recruiting system; (ii) carrying out, analyzing and improving our application and recruitment process, including assessing you and your application and conducting employment, background and reference checks; (iii) establishing an employment relationship or entering into an employment contract with you; (iv) complying with our legal, regulatory and corporate governance obligations; (v) recordkeeping; (vi) ensuring network and information security and preventing fraud; and (vii) as otherwise required or permitted by applicable law. Support the development of voice of customer vehicle testing to help understand the hardware and software integration robustness of Rivian products Set and communicate Reliability requirements and targets for Rivian products and work with other reliability engineers to cascade those targets to sub-system and component level requirements.
NewSite Reliability Engineer HonorVet TechnologiesSite Reliability EngineerSanta Clara, CARemote$55–$58 / hourHonorVet Technologies is a Service Disable Veteran-Owned IT staffing firm, ISO 9001 and ISO 27001 certified, working with federal agencies, state governments, and Fortune 500 enterprise clients across the US. You'll work cross-functionally to create alignment and deliver results alongside builders who have helped to shape the success of companies such as Google, Okta, AWS, Snowflake.
NewSenior Site Reliability Engineer : 26-02282 Akraya, Inc.Senior Site Reliability Engineer : 26-02282San Jose, CARemote$55–$60 / hourMost recently, we were recognized Stevie Employer of the Year 2025, SIA Best Staffing Firm to work for 2025, Inc 5000 Best Workspaces in US (2025 & 2024) and Glassdoor's Best Places to Work (2023 & 2022)! This role seeks a seasoned Site Reliability Engineer to develop and manage critical platform services for secure, high-compliance government environments.
NewStaff Site Reliability Engineer, AI Foundations, F1 Query Google LLCStaff Site Reliability Engineer, AI Foundations, F1 QuerySan Jose, CAF1 Query powers over 400 production systems across Ads, Finance, Play, Cloud, YouTube, Google DeepMind and Enterprise AI, and helps users address various ad-hoc and low-latency data processing, serving, investigative, and batch use cases. Engage in software engineering on services written in Java, C++, and Go (including instrumenting client software development kita (SDKs), server-side performance changes, capacity planning, experiments, monitoring, and more), and performance enhancement of existing systems.
Lead Infrastructure and Reliability Engineer (Systems & Scale) Luma AI IncLead Infrastructure and Reliability Engineer (Systems & Scale)Palo Alto, CARequired: • Deep expertise in Linux and distributed systems • Experience operating GPU / accelerator clusters in real production environments • Strong fluency in Kubernetes and modern open-source infrastructure • Comfortable debugging across hardware kernel runtime orchestration • You understand how systems behave under contention and at scale • You write code and build automation • You think in bottlenecks, failure modes, and tradeoffs • Engineers trust your judgment, especially when things break. • Scaling Training & Inference Define how infrastructure and workloads evolve as cluster size and concurrency grow Design scheduling, placement, and resource management approaches for increasingly complex jobs Work directly with research to build the systems required for new model capabilities Ensure inference platforms scale rapidly without sacrificing reliability or latency Anticipate where today's abstractions will fail and redesign ahead of them.
Site Reliability Engineer, Cloud SQL SRE Google LLCSite Reliability Engineer, Cloud SQL SRESunnyvale, CAWe"re looking for engineers who bring fresh ideas from all areas, including information retrieval, distributed computing, large-scale system design, networking and data storage, security, artificial intelligence, natural language processing, UI design and mobile; the list goes on and is growing every day. Google Cloud offers a portfolio of managed database services, catering to a wide range of customer needs from enterprise applications to modern, cloud-native solutions, including those powering Gen AI/ML.
Site Reliability Engineer, Customer Systems, IS&T Apple IncSite Reliability Engineer, Customer Systems, IS&TSunnyvale, CAIn this highly visible position, you will:Innovate, architect, build, and document highly available, scalable, reliable, secure Infrastructure Troubleshoot application specific, network, system & performance issues Build and maintain CI/CD infrastructure to enable fast delivery cycles for software engineering teams Envision and build automation tools to deliver infrastructure services reliably and in a repeatable fashion Collaborate with other site reliability engineers, software engineers, quality engineers, to gather, define, and analyze non-functional/technical requirements3+ years of experience with deploying/managing Kubernetes using Helm Experience with Shell Scripting, Python, Ansible Experience in monitoring using Splunk, Grafana, Prometheus, Alertmanager Deep understanding of networking protocols: DNS, TCP, HTTP/HTTPS Experience in setting up and managing CI/CD pipelines Bachelors or Masters in Computer Science or equivalent experience Experience in deploying, monitoring and supporting java applications Experience with ArgoCD and GitOps model Experience in defining, monitoring and achieving key operational metrics like MTTR and SLO Experience with GenAI tools in workflow automation for infrastructure management Ability to learn new technologies in a short time Excellent problem solving, critical thinking, and interpersonal skills Good communication skills to collaborate with distributed teams. In this role you will design, build and deliver highly scalable, reliable, secure cloud infrastructure which powers the applications and services used by Apple's customers every day.
Sr. Reliability Engineer Kaygen Inc.Sr. Reliability EngineerMilpitas, CA$45–$55 / hourWe specialize in providing high-volume contingent staffing, direct hire staffing and project-based solutions to companies worldwide ranging from startups to Fortune 500 and Managed Service Providers (MSP) across a wide variety of industries. Minimum of 5 years of reliability engineering experience or experience performing tests to collect experimental data and performing statistical analyses to interpret results.