Senior Database Reliability Engineer Crunchyroll, IncSenior Database Reliability EngineerSan Francisco, CARemote$203,400–$254,200 / yearCollaborative Growth & Development: Work alongside a seasoned team of database engineering specialists (leveraging existing senior architectural depth on the team) to systematically scale platform features while executing a continuous learning roadmap to expand personal depth in native AWS database services (RDS/Aurora) and complex SQL tuning. Database SRE & Site Operations: Manage large-scale data infrastructures, execute cluster management, capacity planning, data governance, compliance reviews, and handle complex data store migrations (such as MariaDB to Aurora/DynamoDB) and major version upgrades safely during non-US low traffic hours.
Reliability Test Engineer Mainspring Energy IncReliability Test EngineerMenlo Park, CA$108,000–$123,600 / yearCommercial, industrial, and utility leaders are choosing Mainspring over traditional options like engines, turbines, and fuel cells to quickly and reliably deliver local power for EV charging, commercial facilities, data centers, and grid-scale operations. Backed by top-tier investors including Khosla Ventures, Bill Gates, American Electric Power, Lightrock, and General Catalyst Mainspring designs, manufactures and delivers its products to customers across the U.S. today, and were quickly scaling for international expansion.
Principal Engineer, Cloud Site Reliability Engineering NVIDIAPrincipal Engineer, Cloud Site Reliability EngineeringUs, CaliforniaThis group works with various other groups within NVIDIA such as Graphics Processors, Mobile Processors, Deep Learning, Artificial Intelligence and Autonomous Vehicles to cater to their infrastructure needs. Are you passionate about distributed infrastructure and looking for sophisticated, critical issues, ready to build the next generation of cloud services, design creative solutions, mine through data to uncover real problems and fix them?.
Software Engineer, Infrastructure & Reliability CrewAISoftware Engineer, Infrastructure & ReliabilitySan Francisco, CABuild the tooling and automation that lets field engineers and customers run self-hosted installs themselves- Helm charts, environment config, release artifacts, pre-flight checks, and install runbooks - so engineering does fewer hands-on installs over time. Manage production observability and telemetry infrastructure: logs, metrics, traces, dashboards, Sentry/OpenTelemetry plumbing, actionable alerts, and telemetry export to customers' own monitoring systems.
System Speed and Reliability Co-Design Engineer NVIDIA CorpSystem Speed and Reliability Co-Design EngineerSanta Clara, CADesign and implement automation tools for system speed modeling; apply AI and LLM-assisted workflows (e.g., automated log analysis, pattern detection, scripting acceleration) to compress characterization and debug cycles. Demonstrated use of AI or LLM-based tools (e.g., Claude, Copilot, ChatGPT) in an engineering workflow-scripting acceleration, log triage, data analysis-with clear judgment about output validation and where automation introduces risk.
System Speed And Reliability Co-Design Engineer NvidiaSystem Speed And Reliability Co-Design EngineerSanta Clara, CADesign and implement automation tools for system speed modeling; apply AI and LLM-assisted workflows (e.g., automated log analysis, pattern detection, scripting acceleration) to compress characterization and debug cycles. Demonstrated use of AI or LLM-based tools (e.g., Claude, Copilot, ChatGPT) in an engineering workflow-scripting acceleration, log triage, data analysis-with clear judgment about output validation and where automation introduces risk.
Sr. Software Engineer (Flight Reliability) SpaceXSr. Software Engineer (Flight Reliability)Hawthorne, CA$160,000–$225,000 / yearThe Flight Reliability software team creates mission critical applications that are used throughout SpaceX to accelerate launch vehicle production and flight as well as systems that allow Starlink to grow into a worldwide fast, reliable Internet service. To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State.
Full Stack Software Engineer (Build Reliability) SpaceXFull Stack Software Engineer (Build Reliability)Hawthorne, CA$125,000–$145,000 / yearTo conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. You may also be eligible for long-term incentives, in the form of company stock, stock options, or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan.
Lead Software Engineer (Full Stack) - Build Reliability SpaceXLead Software Engineer (Full Stack) - Build ReliabilityHawthorne, CA$160,000–$225,000 / yearThis involves streamlining manufacturing quality control processes, simplifying new product introduction, ensuring product traceability, as well as optimizing rocket and spacecraft reusability across the Falcon, Dragon, and Starship programs. To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State.
Software Engineer (Flight Reliability) SpaceXSoftware Engineer (Flight Reliability)Hawthorne, CA$125,000–$145,000 / yearThe Flight Reliability software team creates mission critical applications that are used throughout SpaceX to accelerate launch vehicle production and flight as well as systems that allow Starlink to grow into a worldwide fast, reliable Internet service. To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State.
NewPrincipal Failure Analysis Engineer WDPrincipal Failure Analysis EngineerIrvine, CAYou will collaborate with worldwide FA teams and review failure analysis reports from the labs and collaborate with labs and subject matter experts to decide on best course of action for identifying underlying causes of failures and drive problems. Pro-actively engage with subject matter experts in PCBA, Firmware, Reliability, Data Analytics, FA Tools, Customer Technical Support and Customer Quality Managers, Factory FA and Quality teams on customer issues.
Senior Site Reliability Engineer OutSystems IncSenior Site Reliability EngineerSan Francisco, CASite Reliability Engineer Role As an SRE at OutSystems here are your key responsibilities and duties: Lead and onboard services and teams to the reliability tenets; Establish and maintain Service Level Objectives (SLOs) and Service Level Agreements (SLAs); Design and implement scalable, reliable, and secure infrastructure, while ensuring cloud-native best practices; Collaborate with software development teams to ensure systems are resilient (observable, fault-tolerant, recoverable, scalable) and performant; Implement monitoring, alerting, logging, and tracing solutions to detect and respond to incidents; Lead incident response efforts, ensuring quick resolution and minimal downtime, and conduct RCA/post-mortems; Automate every operational task, with a special focus on fast incident detection & recovery; Programming in Python supported by Gen AI tooling to accelerate development of mission critical automation and tools. Containerization technologies and orchestration platforms, mainly Kubernetes and EKS (CKA, CKAD, CKS certifications are valued); Experience with automation and Infrastructure as Code (IaC) tools, such as AWS CloudFormation, Terraform, Puppet, Chef, Spacelift, etc; Experience with Python, Go, Bash/Shell scripting, or other automation tools/languages; Familiarity with AWS services like EC2, RDS, ELB, CloudFront, Lambda, etc; Proficiency in monitoring and troubleshooting complex distributed systems; Experience with Grafana, ELK stack, Prometheus, or others; Strong understanding of designing resilient and fault-tolerant systems; Expertise in debugging complex distributed systems.
NewSite Reliability Engineer Iceye UsSite Reliability EngineerIrvine, CaliforniaSome travel may be necessary – the ability to travel by car, plane, train, bus, vessel, or metro, operate a motor vehicle and maintain a valid Driver’s License and/or effectively navigate public transportation is required. Working alongside a team of skilled DevOps engineers, you'll find yourself in a collaborative and supportive environment where your ideas are valued, and your contributions make a tangible impact.
Sr. Software Engineer (Flight Reliability) Space Exploration TechnologiesSr. Software Engineer (Flight Reliability)Hawthorne, CA$160,000–$225,000 / yearThe Flight Reliability software team creates mission critical applications that are used throughout SpaceX to accelerate launch vehicle production and flight as well as systems that allow Starlink to grow into a worldwide fast, reliable Internet service. ITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State.
NewSoftware Engineer (Flight Reliability) Space Exploration TechnologiesSoftware Engineer (Flight Reliability)Hawthorne, CA$125,000–$145,000 / yearThe Flight Reliability software team creates mission critical applications that are used throughout SpaceX to accelerate launch vehicle production and flight as well as systems that allow Starlink to grow into a worldwide fast, reliable Internet service. ITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State.
NewFull Stack Software Engineer (Build Reliability) Space Exploration TechnologiesFull Stack Software Engineer (Build Reliability)Hawthorne, CA$125,000–$145,000 / yearITAR REQUIREMENTS: To conform to U.S. Government export regulations, applicant must be a (i) U.S. citizen or national, (ii) U.S. lawful, permanent resident (aka green card holder), (iii) Refugee under 8 U.S.C. § 1157, or (iv) Asylee under 8 U.S.C. § 1158, or be eligible to obtain the required authorizations from the U.S. Department of State. You may also be eligible for long-term incentives, in the form of company stock, stock options, or long-term cash awards, as well as potential discretionary bonuses and the ability to purchase additional stock at a discount through an Employee Stock Purchase Plan.
Lead Mechanical Design Engineer – High-Speed Rotating Machinery Sapphire TechnologiesLead Mechanical Design Engineer – High-Speed Rotating MachineryCypress, CaliforniaSapphire Technologies is seeking a Lead Mechanical Design Engineer to lead the development and production launch of advanced high-speed rotating machinery incorporating active magnetic bearings, high-speed electric machines, power electronics, and precision mechanical systems. Grow into a technical authority for high-speed magnetic-bearing systems, with opportunities to lead platform architecture, engineering teams, reliability initiatives, technology development, and future product programs.
Director - Intel Foundry Pre-Silicon Design Quality And Reliability Intel Corp.Director - Intel Foundry Pre-Silicon Design Quality And ReliabilitySanta Clara, CA$271,620–$383,460 / yearAs the leader of Intel's Foundry pre-Silicon Quality and Reliability team, you will lead a high-performing team of Design and Quality and Reliability engineers responsible for the tools and flows which allow Foundry customers to design reliable, high performance products on Intel Silicon technologies. Mentor and develop a team of design and quality and reliability engineers and build disciplined collaboration structures with internal and external partners, promoting a collaborative, inclusive, and productive work environment that aligns with Intel's values.
Reliability Manager Ben ArisReliability ManagerCaliforniaThe Reliability Manager will lead and oversee the Fixed Equipment Reliability Engineers, along with the Chief Refinery Inspector, focusing on enhancing and maintaining the reliability and integrity of the refinery's fixed equipment. Provides leadership and guidance to the Chief Refinery Inspector to ensure that required inspection tasks are being completed and equipment integrity risks are known and being effectively managed.
Reliability Engineer, Facilities & Infrastructure Impulse Space Propulsion Inc.Reliability Engineer, Facilities & InfrastructureRedondo Beach, CA$100,000–$145,000 / yearPerform Root Cause Analysis (RCA), Failure Mode and Effects Analysis (FMEA), Fault Tree Analysis (FTA), and other reliability analyses to identify and mitigate repetitive or high-consequence failuresIdentify opportunities to improve equipment reliability, maintainability, and lifecycle cost through engineering modifications and maintenance strategy changes. This position requires applicants to be either U.S. Persons (i.e., U.S. citizen, U.S. national, lawful permanent U.S. resident (green card holder), an individual granted asylum in the U.S., or an individual admitted in U.S. refugee status) or persons eligible to obtain an export license from the U.S. Departments of State, Commerce, or other applicable U.S. government agencies.
Senior Design Reliability Engineer Impulse Space Propulsion Inc.Senior Design Reliability EngineerRedondo Beach, CA$140,000–$180,000 / yearTo conform to U.S. Government space technology export regulations, including the International Traffic in Arms Regulations (ITAR) you must be a U.S. citizen, lawful permanent resident of the U.S., protected individual as defined by 8 U.S.C. 1324b(a)(3), or eligible to obtain the required authorizations from the U.S. Department of State. Demonstrated experience supporting design-for-manufacturing of builds, integration, or test of primary structures and/or complex mechanical or electromechanical systems.
NewQuality Engineer (Mechanical Product) JobotQuality Engineer (Mechanical Product)Santa Barbara, CA$140,000–$175,000 / yearInformation collected and processed as part of your Jobot candidate profile, and any job applications, resumes, or other information you choose to submit is subject to Jobot's Privacy Policy, as well as the Jobot California Worker Privacy Notice and Jobot Notice Regarding Automated Employment Decision Tools which are available at jobot.com/legal. This individual will partner with engineering, operations, and supply chain teams to implement inspection strategies, manage supplier performance, and lead structured problem solving to improve product reliability and manufacturing performance.
Senior Lead Site Reliability Engineer JPMorgan Chase & CoSenior Lead Site Reliability EngineerPalo Alto, CAStrong experience building production-grade RESTful APIs and designing message queue architectures (Kafka, RabbitMQ, SQS) for event-driven systems; and expertise in graph databases (Neo4j, TigerGraph), vector databases (Pinecone, Weaviate, Chroma), and integrating multiple data stores for AI-powered systems. Hands-on experience building AI Agents and autonomous systems with proficiency in AI frameworks (LangChain, LangGraph, AutoGen, CrewAI) and leveraging AI development tools (GitHub Copilot, Claude, etc.) to accelerate development and innovation and Expertise in designing and implementing logging pipelines (Fluentd, Logstash, Vector) and systems for metrics collection, analysis, and distributed tracing.
Manager, Build Reliability K2 Space CorpManager, Build ReliabilityLos Angeles, CABacked by $450M from leading investors including Altimeter Capital, Redpoint Ventures, T. Rowe Price, Lightspeed Venture Partners, Alpine Space Ventures, and others - with an additional $500M in signed contracts across commercial and US government customers - we're mass-producing the highest-power satellite platforms ever built for missions from LEO to deep space. Investigate failures, drive corrective actions, and lead containment efforts that reduce risk and improve reliability across production; feed lessons learned back into the design for future component iterations and increase the repeatability of simple, robust flight unit builds.
Manager, Build Reliability K2 SpaceManager, Build ReliabilityLos Angeles, CaliforniaBacked by $450M from leading investors including Altimeter Capital, Redpoint Ventures, T. Rowe Price, Lightspeed Venture Partners, Alpine Space Ventures, and others – with an additional $500M in signed contracts across commercial and US government customers – we’re mass-producing the highest-power satellite platforms ever built for missions from LEO to deep space. Investigate failures, drive corrective actions, and lead containment efforts that reduce risk and improve reliability across production; feed lessons learned back into the design for future component iterations and increase the repeatability of simple, robust flight unit builds.
Reliability Test Engineer, Special Projects Mainspring EnergyReliability Test Engineer, Special ProjectsMenlo Park, California$108,000–$123,600 / yearThe company began commercial shipments of its Mainspring Linear Generators in 2020 and today has hundreds of megawatts in advanced development and field operations for leading Fortune 500 companies, data centers, and utilities. At Mainspring, we are committed to building a diverse, inclusive, flexible, and collaborative environment, so if you want to help us transition the world to clean and affordable electricity, and don’t meet all posted requirements for a particular role, we’d still love to hear from you.
Applications Engineer for High Reliability Compute Power Monolithic Power SystemsApplications Engineer for High Reliability Compute PowerSan Jose, CaliforniaThis individual works closely with customers, internal teams such as marketing, sales and field application engineers to develop new products, and supports design-in activities of MPS compute products into customer projects. MPS is seeking a self-motivated senior level engineer to drive system level architecture, product definition and application support for power management solutions for High Reliability Compute Power applications.
Senior Software Engineer, Site Reliability Engineering GoogleSenior Software Engineer, Site Reliability EngineeringSan Francisco, CANote: By applying to this position you will have an opportunity to share your preferred working location from the following: San Francisco, CA, USA; Pittsburgh, PA, USA; Durham, NC, USA; Raleigh, N.C., USA; San Bruno, CA, USA; Sunnyvale, CA, USA; New York, NY, USA . Practices such as limiting time spent on operational work, blameless postmortems and proactive identification of potential outages factor into iterative improvement that is key to both product quality and interesting and dynamic day-to-day work.
Senior Staff Software Engineer, Site Reliability Engineering GoogleSenior Staff Software Engineer, Site Reliability EngineeringSan Jose, CASRE ensures that Google's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to users' needs and a fast rate of improvement. Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems.
Staff Software Engineer, Site Reliability Engineering GoogleStaff Software Engineer, Site Reliability EngineeringSunnyvale, CASRE ensures that Google's services—both our internally critical and our externally-visible systems—have reliability, uptime appropriate to users' needs and a fast rate of improvement. Site Reliability Engineering (SRE) combines software and systems engineering to build and run large-scale, massively distributed, fault-tolerant systems.
Hardware Reliability Engineer - Watch System Reliability Apple IncHardware Reliability Engineer - Watch System ReliabilitySan Diego, CAA deep understanding of material and device properties (mechanical, electrical, optical, thermal) and the underlying physics Proficiency in statistical data analysis and clear reporting, providing design risk assessments and guiding product improvements Strong theoretical knowledge in physical failure analysis techniques, such as SEM/EDX/X-Ray and other characterization techniques Proven ability to develop new reliability test procedures, create detailed test plans, and analyze results to assess design risksBachelor's Degree in technical field (mechanical engineering, materials engineering, electrical engineering, physics or related fields etc.) with 3+ years of relevant engineering experience Familiar with Failure Analysis techniques (Optical Microscopy, X-ray/CT, Scanning Electron Microscopy/Energy Dispersive Spectroscopy, etc.), and the ability to use failure analysis methodology to derive a root cause of failure Ability to make clear and concise slides and presentations through keynote, excel, etc. Ability to travel internationally without restriction (up to 10%)M.S. or PhD in Materials Science, Mechanical Engineering, Electrical Engineering, or an equivalent field Statistical experience such as Weibull, JMP, or familiarity with accelerated test models Excellent communication skills, both written and verbal Sharp attention to detail Confidence in presenting test results and risk assessments to cross-functional and leadership teams Ability to manage multiple projects simultaneously Capability to perform tests and troubleshoot issues independently Openness to take on new tasks A proactive, can-do attitude, with a passion for working alongside an amazing team and innovative products Collaborative attitude and ability to work cross-functionally effectively.
SRE/Devops Engineer- San Jose, the US KodyPay LtdSRE/Devops Engineer- San Jose, the USSan Jose, CAWe are looking for a unique engineering mindset: someone who brings a positive, collaborative energy to the daily grind, but can instantly pivot into a hyper-focused, high-ownership responder when an incident strikes. Cross-Border Collaboration: Act as a key technical bridge between our US operations and international engineering hubs, leveraging bilingual communication to streamline complex technical alignment.
SRE/DevOps Engineer- Palo Alto, the US KodyPay LtdSRE/DevOps Engineer- Palo Alto, the USPalo Alto, CAWe are looking for a unique engineering mindset: someone who brings a positive, collaborative energy to the daily grind, but can instantly pivot into a hyper-focused, high-ownership responder when an incident strikes. Cross-Border Collaboration: Act as a key technical bridge between our US operations and international engineering hubs, leveraging bilingual communication to streamline complex technical alignment.
Fixed And Rotating Equipment Manager PBF EnergyFixed And Rotating Equipment ManagerMartinez, CA$127,218.49–$226,895.29 / yearThis position is responsible for developing and executing equipment integrity and reliability strategies, leading a team of engineering professionals, and partnering with Operations, Maintenance, Technical Services, Capital Projects, and Business Services to optimize equipment performance and support refinery business objectives. The Fixed and Rotating Equipment Manager provides strategic leadership and technical oversight to the refinery's Fixed and Rotating Equipment teams to ensure the safe, reliable, and compliant operations of refinery assets.
NewFixed and Rotating Equipment Manager PBF EnergyFixed and Rotating Equipment ManagerMartinez, California$127,218.49–$226,895.29 / yearThis position is responsible for developing and executing equipment integrity and reliability strategies, leading a team of engineering professionals, and partnering with Operations, Maintenance, Technical Services, Capital Projects, and Business Services to optimize equipment performance and support refinery business objectives. The Fixed and Rotating Equipment Manager provides strategic leadership and technical oversight to the refinery’s Fixed and Rotating Equipment teams to ensure the safe, reliable, and compliant operations of refinery assets.
Senior Software Engineer - Observability and Reliability Sigma Computing IncSenior Software Engineer - Observability and ReliabilitySan Francisco, CA$170,000–$240,000 / yearThe round was led by Princeville Capital, with new strategic investors Databricks Ventures, ServiceNow Ventures, and Workday Ventures participating alongside returning investors Altimeter Capital, Avenir Growth Capital, D1 Capital Partners, K5 Global, NewView Capital, Spark Capital, Sutter Hill Ventures, and XN. This milestone follows Sigma reaching $200M in annual recurring revenue in April 2026, with more than 100% year-over-year growth and 1.1 million new active users added in the latest fiscal year.
Hardware Reliability Engineer - Mac System Reliability Apple IncHardware Reliability Engineer - Mac System ReliabilityCupertino, CABachelor's Degree in technical field ( mechanical engineering, materials engineering, electrical engineering, physics or related fields) Collaborative attitude and ability to cross-functionally work effectively Good written and verbal English communication skills. A proactive, can-do attitude, with a passion for working alongside an amazing team and innovative productsIdentifying high-risk failure modes early in the design process and working closely with design engineering teams to mitigate risks.
Manager, Site Reliability Engineering Aya Healthcare IncManager, Site Reliability EngineeringCA$230,000–$255,000 / yearYou''ll shape our reliability architecture, lead complex operational initiatives, and drive the adoption of AI-native operations (AIOps) and automation to eliminate toil and advance performance - owning measurable business outcomes across uptime, customer trust, and platform efficiency, and leading with the radical ownership Aya expects of every leader. Own the reliability strategy for customer-facing products and internal platforms - defining SLOs, SLIs, and error budgets in partnership with product and engineering leadership, and operationalizing them in the release process.
Reliability Engineering Technical Leader CiscoReliability Engineering Technical LeaderSan Jose, CaliforniaThe applicable full salary ranges for this position, by specific state, are listed below: New York City Metro Area: $162,500.00 - $236,200.00 Non-Metro New York state & Washington state: $144,100.00 - $214,500.00 * For quota-based sales roles on Cisco’s sales plan, the ranges provided in this posting include base pay and sales target incentive compensation combined. The Global Process Operations (GPO) group is a core part of the Product Operations Central team, providing centralized expertise that strengthens the strategic and operational framework for Business Units (BUs) across Cisco’s product portfolio.
Senior Build Reliability Engineer (Machining) Impulse SpaceSenior Build Reliability Engineer (Machining)Redondo Beach, CA$140,000–$170,000 / yearThis position requires applicants to be either U.S. Persons (i.e., U.S. citizen, U.S. national, lawful permanent U.S. resident (green card holder), an individual granted asylum in the U.S., or an individual admitted in U.S. refugee status) or persons eligible to obtain an export license from the U.S. Departments of State, Commerce, or other applicable U.S. government agencies. You will work hands-on with technicians, manufacturing engineers, and supporting teams to prevent defects before they occur, rapidly diagnose failures, and continuously improve hardware reliability across spacecraft builds.
Site Reliability Engineer II, tvScientific Pinterest IncSite Reliability Engineer II, tvScientificSan Francisco, CARemote$114,297–$235,319 / yearWe are seeking a Site Reliability Engineer to help operate, scale, and continuously improve a cloud-native platform built on AWS, Kubernetes/EKS, and ArgoCD-driven GitOps workflows. Our platform is built by industry leaders with a long history in programmatic advertising, digital media, and ad verification who have now purpose-built a CTV performance platform advertisers can trust to grow their business.
Senior Release Train Engineer Cox AutomotiveSenior Release Train EngineerFair Oaks, GA$101,500–$169,100Facilitation and relentless improvement of periodic ART reviews with organizational leaders and stakeholders to provide clarity on the expected value delivery and to gain stakeholder feedback on changes to direction the release train may need to take. Assists the ART to relentlessly improve through the facilitation of ART Retrospectives quarterly, or more often as needed, to improve backlog management, ART reviews, sprint reviews, or ART coordination and higher-level function.
Senior Software Engineer, Machine Learning Infrastructure (Tinder LLC, West Hollywood, California) Match Group IncSenior Software Engineer, Machine Learning Infrastructure (Tinder LLC, West Hollywood, California)West Hollywood, CA$190,000–$246,000 / year2 years of professional experience working with modern cloud platforms (including AWS, Azure, or GCP) and utilizing infrastructure-as-code practices, containerization tools (Docker on managed orchestration platforms including Amazon EKS or Amazon ECS), and monitoring systems based on Prometheus metrics and Grafana dashboards, including experience operating services backed by a timeseries metrics store including Grafana Mimir. 1 years of professional experience designing and implementing CI/CD automation pipelines and GitOps practices for ML infrastructure, using tools including Terraform, Terragrunt, Helm, and internal GitOps systems (including Scaffold) together with continuous integration systems (including Jenkins or Buildkite) to manage deployment strategies including canary releases, bluegreen deployments, and zerodowntime migrations of backend services.
Quality & Reliability Engineering AppleQuality & Reliability EngineeringCupertino, CAExperience partnering with off-shore manufacturers, high volume production and traveling to manufacturing sites + Strong statistical and data analytics skills - statistical process control and sampling methods + Data driven and results oriented approach to problem solving, failure analysis techniques + Direct experience program managing functional engineering teams (ME, EE, TE, DFX or other) + Exceptional ability to build relationships; clear, consistent communication; data-driven and results- oriented; enthusiastic and motivated + Obsessively passionate and inquisitive, who seeks to pursue everyday problems in innovative ways + Laser-focused on the smallest details and able to use data forensics to solve complex manufacturing assembly quality issues Our Organization is currently hiring for the following roles: -FATP Product Quality Engineer - iPhone, Mac, Audio, VisionPro -Mechanical Engineering Product Quality Manager - iPhone, VisionPro -Sustaining Product Quality Engineer - Mac -Core Technologies Quality Engineer -Core Technologies Supplier Quality Engineer -Operations Reliability Engineer - iPhone **Minimum Qualifications** + 3+ years of Quality experience in one or more of the following disciplines: New Product Introduction / Design Engineering / Manufacturing Operations or Reliability Engineering + BS in Mechanical, Electrical, Optical, Materials, or Industrial Engineering or equivalent.
Site Reliability Engineer, Diagnostics Tesla IncSite Reliability Engineer, DiagnosticsPalo Alto, CA$140,000–$312,000 / yearWe're the small, expert and growing team creating the next-generation diagnostics software and services to support the growing fleets of Tesla products, and we're looking for seasoned SREs with domain expertise in one or more of: containers, public clouds and cloud-native apps. Along with competitive pay, as a full-time Tesla employee, you are eligible for the following benefits at day 1 of hire: Medical plans > plan options with $0 payroll deduction.
Senior Reliability Design Engineer (Teradyne, San Jose, CA) TeradyneSenior Reliability Design Engineer (Teradyne, San Jose, CA)San Jose, CA$155,500–$248,900 / yearIn this role, you will partner with instrumentation, automation, and system design teams to apply Teradyne's Design for Reliability (DfR) methodology, evaluate new technologies, reduce reliability risks, and improve product robustness throughout the development lifecycle. Teradyne's Memory Test Division is seeking a Senior Reliability Design Engineer to help develop next-generation semiconductor test systems by driving reliability into products from concept through qualification and release.
Sr. Reliability Engineer Sunrise Systems IncSr. Reliability EngineerMilpitas, CA$50–$53 / hourMinimum of 5 years of reliability engineering experience or experience performing tests to collect experimental data and performing statistical analyses to interpret results. Manage reliability projects by implementing tests, analyzing data, assessing reliability risks, and reporting accurate information to stakeholders.
Senior Engineer, AI Site Reliability Fox CorpSenior Engineer, AI Site ReliabilityLos Angeles, CA$114,000–$165,000 / yearThe senior engineer will serve as an SME for solving thundering herd problems and leading efforts to leverage AI to diagnose and remediate production incidents, automate tooling, and partnering with backend api teams to run deployments, load testing, and capacity planning for daily traffic and large live events. Fox is hiring a Senior Engineer, AI Site Reliability to help build and operate infrastructure and platforms to support APIs around our live direct to consumer APIs for major live events such as the Super Bowl, World Cup, and World Series.
NewSoftware Development Engineer – Performance & Reliability AVEVASoftware Development Engineer – Performance & ReliabilityLake Forest, California92,300.00 - $153,900.00 T his pay range represents the minimum and maximum compensation that the position offers, and final compensation can vary within the range depending on work location, job experience, skills, and relevant educational attainment and/or training. AVEVA is creating software trusted by over 90% of leading industrial companies.
Senior Software Engineer, Agents - Hiring Sprint AirbyteSenior Software Engineer, Agents - Hiring SprintSan Francisco, CAOur runtime resolves entities across hundreds of business systems, assembles the right context, chooses the appropriate connectors and skills, validates evidence, enforces permissions, executes actions safely, and returns responses users can trust. We believe the future isnt simply giving LLMs access to APIs, its building the infrastructure that allows agents to reason over enterprise context, retrieve evidence, invoke actions safely, and explain every decision they make.