NewSoftware Engineer, Storage Infrastructure Apple IncSoftware Engineer, Storage InfrastructureCupertino, CAIn this role, you will: Shape the products features and architecture as it scales orders of magnitude Build systems that simply work, empowering billions of meaningful moments for people around the world Design storage device-optimized low-level storage solutions from first principles Develop and maintain large-scale distributed systems with mission-critical reliability Optimize high-performance IO stacks for maximum availability and durability Drive operational excellence across storage platform infrastructure Collaborate cross-functionally to deliver resilient platforms that enable innovation across Apple Demonstrate passion for large-scale distributed systems and creating robust storage solutions3 years of professional software development experience Strong analytical and problem-solving skills, with meticulous attention to detail. Experience in building storage systems 2+ years of coding in one or more of these programming languages: Rust, C++, Java or C# Experience with scripting languages (Bash, Python, Perl) Bachelor's Degree in Computer Science, an engineering-related field, or equivalent related experience.
NewSite Reliability Engineer, Inference Infrastructure CohereSite Reliability Engineer, Inference InfrastructureSan Francisco, CAFull‑Time Employees At Cohere Enjoy These PerksAn open and inclusive culture and work environmentWork closely with a team on the cutting edge of AI researchWeekly lunch stipend, in‑office lunches & snacksFull health and dental benefits, including a separate budget to take care of your mental health100% Parental Leave top‑up for up to 6 monthsPersonal enrichment benefits towards arts and culture, fitness and well‑being, quality time, and workspace improvementRemote‑flexible, offices in Toronto, New York, San Francisco, London and Paris, as well as a co‑working stipend️ 6 weeks of vacation (30 working days!)#J-18808-Ljbffr. We're training and deploying frontier models for developers and enterprises who are building AI systems to power magical experiences like content generation, semantic search, RAG, and agents.
NewStaff Software Engineer (Cloud Infrastructure) Crusoe EnergyStaff Software Engineer (Cloud Infrastructure)San Francisco, CA$215,000–$260,000 / yearWe're looking for problem-solving, opportunity-finding teammates with a sense of urgency, who believe in the scale of our ambition and thrive on a path not fully paved - people who want to grow their careers alongside a team of experts across energy, manufacturing, data center construction, and cloud services. The ideal candidate will be hands-on with GPU rack-level troubleshooting and work closely with data center operations, engineering, and vendors to support cutting-edge infrastructure featuring the latest NVIDIA and AMD GPUs.
Staff+ Software Engineer, Safeguards ML Infrastructure AnthropicStaff+ Software Engineer, Safeguards ML InfrastructureSan Francisco, CA$320,000–$485,000 / yearThis research continues many of the directions our team worked on prior to Anthropic, including: GPT-3, Circuit-Based Interpretability, Multimodal Neurons, Scaling Laws, AI & Compute, Concrete Problems in AI Safety, and Learning from Human Preferences. For sales roles, the range provided is the role's On Target Earnings ("OTE") range, meaning that the range includes both the sales commissions/sales bonuses target and annual base salary for the role.
Senior Software Engineer, Infrastructure Parafin IncSenior Software Engineer, InfrastructureSan Francisco, CA$230,000–$265,000 / yearWe're a tight-knit team of innovators hailing from Stripe, Square, Plaid, Coinbase, Robinhood, CERN, and more - all united by a passion for building tools that help small businesses succeed. We partner with companies like DoorDash, Amazon, Worldpay, and Mindbody to offer fast and flexible funding, spend management, and savings tools to their small business users via a simple integration.
Staff Software Engineer, Data Infrastructure, AI Compute Platform CZ BiohubStaff Software Engineer, Data Infrastructure, AI Compute PlatformRedwood City, CA$214,000–$295,000 / yearSuccess requires excellence across five interconnected pillars: training frontier AI models specifically for biology; building engineering systems that maximize research velocity and efficiency; executing a sophisticated data strategy that fuels AI development; operating a world-class AI compute platform; and creating impactful products that transform AI capabilities into accessible scientific tools. Design and implement flexible, scalable, and performant systems to address our stakeholders' needs, leveraging technologies like Argo Workflows, Slurm, Ray, AWS Parallel Cluster for mass-scale job processing and orchestration; Vast Data, Delta Lake, Databricks and Apache Iceberg for data management and access; and cloud and on-prem HPC resources.
DevOps Engineer - Infrastructure KodiakDevOps Engineer - InfrastructureMountain View, CA$180,000–$200,000 / yearShould the position require, and Kodiak determines that a candidate's residence, U.S. person status, and/or citizenship status necessitate an export license, bar the candidate from the position, or otherwise fall under national security-related restrictions, Kodiak will consider the candidate for alternative positions unaffected by such restrictions, under terms and conditions set forth at Kodiak's sole discretion, or, as an alternative, opt not to proceed with the candidate's application. California Pay Range $180,000—$200,000 USD At Kodiak, we strive to build a diverse community working towards our common company goals in a safe and collaborative environment where harassment of any kind is strictly prohibited.
Devops Engineer - Infrastructure KodiakDevops Engineer - InfrastructureMountain View, CA$180,000–$200,000 / yearShould the position require, and Kodiak determines that a candidate's residence, U.S. person status, and/or citizenship status necessitate an export license, bar the candidate from the position, or otherwise fall under national security-related restrictions, Kodiak will consider the candidate for alternative positions unaffected by such restrictions, under terms and conditions set forth at Kodiak's sole discretion, or, as an alternative, opt not to proceed with the candidate's application. Kodiak is looking for an individual who is customer focused, passionate about the end user experience, enjoys challenging themselves constantly to improve, and empowering the end user community by providing IT knowledge and tools.
Software Engineer, Infrastructure Chime USA, Inc.Software Engineer, InfrastructureSan Francisco, CARemoteOur in-office work policy is designed to keep you connected - with four days a week in the office and Fridays from home for those near one of our offices, plus team and company-wide events depending on location. Whether it''s starting a savings account, purchasing a first car or home, launching a business, or pursuing higher education, we''re proud to have helped millions unlock their financial potential.
Infrastructure Security Monitoring Engineer Meta Platforms IncInfrastructure Security Monitoring EngineerMenlo Park, CADegree must be completed prior to joining Meta 3+ years of development experience in at least one programming language (Python, Go, etc.) with the ability to apply that to security tool development, automation, and overall programmatic solutions that will be used to defend infrastructure 1+ years of experience in offensive/defensive security or systems engineering Knowledge of network protocols (TCP/IP, computer networking, routing and switching) and Unix based systems Experience researching, building, and implementing defensive security systems that are used against internal and external attack vectors Experience designing and building out application, system and network security monitoring to aid in detection or forensic investigations Experience developing baselines and investigating anomalies in order to identify suspicious behavior Experience adhering to and implementing responsible, ethical AI practices (e.g., risk assessment, bias mitigation, quality and accuracy reviews) Understanding of MITRE ATT&CK Framework and associated threat actor techniques Experience developing automation and utilizing frameworks to scale detection, mitigation or response tools Background in intrusion detection, security investigations, and incident response Experience threat hunting, i.e. using threat intel to proactively and iteratively investigate potential risks and finding suspicious behavior Demonstrated ongoing AI skill development (e.g., prompt/context engineering, agent orchestration) and staying current with emerging AI technologies Demonstrated ability to integrate AI tools to optimize/redesign workflows and drive measurable impact (e.g., efficiency gains, quality improvements) Experience mentoring and promoting industry security practicesMeta builds technologies that help people connect, find communities, and grow businesses. Iterate security posture to better protect against attacks and detect new vectors Lead efforts to mitigate and investigate security incidents Utilize frameworks to develop and scale detection, mitigation and response automation tooling Evaluate and test new vendor and home-grown initiatives for security issues Mentor and evangelize security practices through cross functional work with engineering teams throughout Meta Keep Meta safe through active operation and defense of critical infrastructureCurrently has, or is in the process of obtaining a Bachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience.
NewQuality Systems Engineer THINK SurgicalQuality Systems EngineerFremont, CAThe Quality System Engineer is responsible for supporting the implementation, maintenance, and continuous improvement of the Quality Management System (QMS), with a primary focus on CAPA, Complaint Handling, Nonconformance Management, audit processes, and quality system infrastructure. SUPERVISORY RESPONSIBILITIESN/AQUALIFICATIONSRequiredBachelor's degree in engineering, science, or related discipline, or an equivalent combination of education and experience.3–6 years of experience in Quality Assurance in the medical device field.
Senior Lead Software Engineer - Windows Server Infrastructure JPMorgan Chase & CoSenior Lead Software Engineer - Windows Server InfrastructureSan Francisco, CAAs a Senior Lead Software Engineer - Windows Server Engineering at JPMorgan Chase within the Corporate Sector Compute Infrastructure Platform (CIP) organization, you exhibit both depth and breadth of knowledge regarding software, applications, and technical processes across multiple technical disciplines. Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs/outputs, and adherence to resiliency and security expectations; experience coaching senior engineers/leads on compliant usage patterns and controls.
Senior Infrastructure/Platform Engineer Slash FinancialSenior Infrastructure/Platform EngineerSan Francisco, CaliforniaYou’ll define our infrastructure roadmap, own critical decisions around performance, observability, security, and deployment, and play a key role in enabling Slash to serve billions of dollars in business spend reliably and securely. We combine the reliability of traditional banking (high yields, competitive rewards, and comprehensive security) with industry-specific features that make businesses more efficient, more competitive, and more profitable.
Lab Engineer, Device Services Infrastructure Apple IncLab Engineer, Device Services InfrastructureSanta Clara, CAIdentifying opportunities to improve operational efficiency through automation, telemetry, and data-driven insights Supporting AI-assisted and automated workflows that improve device health, recovery outcomes, deployment verification, and overall lab reliability Leveraging operational metrics and analytics to identify bottlenecks and drive scalable operational improvementsMaintaining a fleet of thousands of iOS/macOS devices and peripherals across multiple labs Ensuring devices and systems are up and available for testing Provisioning and deployment of released/unreleased hardware Debugging and troubleshooting device hardware/software failures Developing scripts and automation for managing machine configuration and software deployment Providing support to engineering teams in debugging systems Identifying opportunities to improve operational efficiency through automation, telemetry, and data-driven insights Supporting AI-assisted and automated workflows that improve device health, recovery outcomes, deployment verification, and overall lab reliability Leveraging operational metrics and analytics to identify bottlenecks and drive scalable operational improvements2-3+ years of experience supporting software services in a Unix-based environment Exposure to automated provisioning processes Basic scripting experience in Shell or Python Familiarity with configuration management tools such as Puppet, Ansible, or similar Working knowledge of Git or other source control systems Experience contributing to shared repositories and collaborative development workflows Strong communication skills and ability to collaborate effectively with engineering teams Willingness to learn and apply structured problem-solving approaches Strong attention to detail and a service-oriented mindsetExperience working in fast-paced engineering environments Exposure to working with service-based systems and RESTful services Basic understanding of LAN networking concepts (DNS, DHCP, routing) Experience working in structured lab environments with high-density hardware deployments Experience supporting macOS and/or iOS devices in a lab environment Familiarity with device provisioning, deployment, and lifecycle management Experience troubleshooting hardware, software, and infrastructure issues. The ideal candidate will contribute to the adoption of AI-assisted operational processes and automated workflows that reduce manual effort, improve decision-making, and support the organizations transition toward more intelligent and scalable lab operations.
Software Engineer, C++ Middleware And Runtime Infrastructure PlusAI Inc.Software Engineer, C++ Middleware And Runtime InfrastructureSanta Clara, CA$120,000–$200,000 / yearWe may use artificial intelligence (AI) tools to support parts of the hiring process, such as reviewing applications, analyzing resumes, or assessing responses and identifying potential inconsistencies or verification signals in application materials based on available information. In this role, you will contribute to existing frameworks, libraries, and tools while also designing and implementing new components across various mission-critical domains.
Staff Software Engineer, Model Infrastructure HarveyStaff Software Engineer, Model InfrastructureSan Francisco, CaliforniaIntegrate new model providers and maintain provider APIs and SDKs, enabling Harvey to rapidly adopt emerging frontier models. You'll partner closely with AI Research, Product Engineering, Infrastructure, and external model providers to build a platform that is highly reliable, scalable, observable, and efficient.
Software Engineer, Infrastructure & Reliability CrewAISoftware Engineer, Infrastructure & ReliabilitySan Francisco, CABuild the tooling and automation that lets field engineers and customers run self-hosted installs themselves- Helm charts, environment config, release artifacts, pre-flight checks, and install runbooks - so engineering does fewer hands-on installs over time. Manage production observability and telemetry infrastructure: logs, metrics, traces, dashboards, Sentry/OpenTelemetry plumbing, actionable alerts, and telemetry export to customers' own monitoring systems.
Senior Software Engineer, Cloud Infrastructure DecagonSenior Software Engineer, Cloud InfrastructureSan Francisco, California$200,000–$400,000 / yearOur technology enables industry-defining enterprises like Avis Budget Group, Block’s Cash App and Square, Chime, Oura Health, and Hunter Douglas to deploy AI agents that power personalized, deeply satisfying interactions across voice, chat, email, SMS, and every other channel. Take end-to-end ownership of deployment architecture in customer-owned cloud environments (VPC configuration, permissioning, networking, provisioning) and the full lifecycle that follows: setup, upgrades, scaling, and incident support.
IT Infrastructure Compliance Engineer NvidiaIT Infrastructure Compliance EngineerSanta Clara, CAActing as a trusted technical advisor, you will identify systemic risks in supply chain automation, intellectual property (IP) protection, B2B data exchanges, and shop-floor control systems, ensuring our global partner network operates securely, compliantly, and at peak efficiency. We are seeking a highly motivated, technically sharp Senior IT Infrastructure Compliance Engineer to lead and execute comprehensive risk assessments across our Global Manufacturing Operations and External Partner ecosystem (encompassing Tier-1 Foundries, ODMs, OSATs, and critical logistics providers).
Backend Engineer - Infrastructure HeyGenBackend Engineer - InfrastructureSan Francisco, CA$180,000–$240,000 / yearHeyGen considers factors such as scope and responsibilities of the position, candidate's work experience, education/training, key skills, and internal equity, as well as location, market and business considerations when extending an offer. ML Infrastructure: Construct the ML training infrastructure to enhance our AI researchers' productivity, and inference systems to optimize the scalability and performance of our video GenAI.
Controls Engineer, Test Infrastructure BelcanControls Engineer, Test InfrastructureBerkeley, CAJob Duties:* ideal candidate is comfortable working across controls, electrical design, automation, and safety systems, and enjoys collaborating closely with test engineers, technicians, and product teams to translate test requirements into robust hardware and software solutions. To be considered for this role you will have a Bachelor"s degree in Electrical Engineering, Controls Engineering, Mechatronics, or a related field and 3+ years of experience developing PLC-based control systems for automated equipment or test infrastructure.
AI Systems Engineer, Codex Agents OpenAIAI Systems Engineer, Codex AgentsSan Francisco, CAFor unincorporated Los Angeles County workers: we reasonably believe that criminal history may have a direct, adverse and negative relationship with the following job duties, potentially resulting in the withdrawal of a conditional offer of employment: protect computer hardware entrusted to you from theft, loss or damage; return all computer hardware in your possession (including the data contained therein) upon termination of employment or end of assignment; and maintain the confidentiality of proprietary, confidential, and non-public information. You'll work with research, infrastructure, and product to design agent harness capabilities, run experiments and ablations across the model + system prompt + harness stack, build frameworks for assessing production agent performance, and turn messy failures into durable improvements.
Global Production Systems Engineer Meta Platforms IncGlobal Production Systems EngineerMenlo Park, CADeliver maximum server fleet uptime and utilization rates, by leveraging data to understand hardware failure conditions and root cause Write and review code, develop documentation, and debug the hardest problems, live, on some of the largest and most complex systems in the world Own and develop diagnostic tooling requirements to run the fleet Own and drive the escalation process for Data Center Operations to identify, root cause, and solve complex tooling and hardware issues affecting the fleet Execute operational validation and verification activities for the new product integration Through consistent collaboration with cross-functional tooling teams, helps determine the root cause and provides input into their development process, with an operations-centric view of how open issues are affecting the fleet Build cross-functional relationships and have the ability to influence policies and procedures to improve global data center operations Mentor team members to evaluate and identify better ways to resolve issues and define updates to tools and processes Travel up to 25% to support global data center operationsBachelor's degree in Computer Science, Computer Engineering, relevant technical field, or equivalent practical experience 6+ years of experience in production systems engineering, infrastructure engineering, or systems software development for large-scale hardware environments 6+ years of experience with hardware lifecycle management, fleet automation, or data center operations systems spanning compute, storage, or networking infrastructure Experience developing systems software or tooling in Python, PHP, C, or C++ for Linux-based production environments at scale Experience in configuration and maintenance of applications such as web servers, load balancers, relational databases, storage systems and messaging systems Experience communicating technical designs and infrastructure decisions through written documentation and cross-functional stakeholder alignment across engineering and operations teams Experience designing or operating configuration management and infrastructure-as-code systems for large heterogeneous hardware fleets Experience supporting global, multi-site data center infrastructure deployments including hardware qualification and regional rollout coordination Familiarity with distributed systems monitoring, alerting, and automated remediation pipelines at hyperscaleMeta builds technologies that help people connect, find communities, and grow businesses. This role is deeply cross-functional and considers the technical needs of frontline users to identify and automate diagnostic tooling, which enables quality and efficient delivery of production servers.
Software Systems Engineer Meta Platforms IncSoftware Systems EngineerFremont, CAYou will drive reliability, efficiency, and performance improvements across the infrastructure stack, partnering closely with hardware engineering, data center operations, and platform teams to ensure Meta's production systems operate at scale. People who choose to build their careers by building with us at Meta help shape a future that will take us beyond what digital connection makes possible todaybeyond the constraints of screens, the limits of distance, and even the rules of physics.
Infrastructure Ops Engineer BaseTenInfrastructure Ops EngineerSan Francisco, CAHigh-Stakes Maintenance Orchestration: Coordinating critical maintenance cycles both externally (with vendors) and internally (with Baseten SREs) to evacuate workloads from unhealthy nodes and integrate replacement hardware with zero customer disruption. You will sit at the intersection of technical customer success and infrastructure engineering, partnering closely with our SRE and FDE teams to execute the complex hardware lifecycles that power our fleet.
Staff Software Development Engineer - Enterprise AI Infrastructure - #4898 GrailStaff Software Development Engineer - Enterprise AI Infrastructure - #4898Menlo Park, CA$169,000–$224,000 / yearThe Staff Engineer partners closely with cross-functional stakeholders across Software Engineering, Data Science, Security, Regulatory Affairs, and Product to build a unified control plane that securely connects large language models with enterprise tools and company knowledge. We have built a multi-disciplinary organization of scientists, engineers, and physicians and we are using the power of next-generation sequencing (NGS), population-scale clinical studies, and state-of-the-art computer science and data science to overcome one of medicine's greatest challenges.
NewPrincipal Machine Learning Infrastructure Engineer, Ads & Discovery RobloxPrincipal Machine Learning Infrastructure Engineer, Ads & DiscoverySan Mateo, CA$295,250–$345,040 / yearYou are comfortable working across modern ML systems technologies such as FSDP, vLLM, SGLang, CUDA, distributed training frameworks, inference engines, and GPU kernels, while remaining tool-agnostic and focused on achieving step-function improvements in model quality, throughput, latency, reliability, and cost. Every day, tens of millions of people come to Roblox to explore, create, play, learn, and connect with friends in 3D immersive digital experiences– all created by our global community of developers and creators.
ML Systems Engineer, Large-Scale Model Training & RL Infrastructure NebiusML Systems Engineer, Large-Scale Model Training & RL InfrastructurePalo Alto, CaliforniaIntegrate and extend frameworks such as Megatron-LM, DeepSpeed, PyTorch FSDP/DTensor, Ray, verl, slime, AReaL, OpenRLHF, or equivalent internal systems. We are building a full-stack AI cloud platform that supports developers and enterprises from data and model training through to production deployment, without the cost and complexity of building large in-house AI/ML infrastructure.
Staff Infrastructure Engineer Hamilton AIStaff Infrastructure EngineerSan Francisco, CaliforniaYou'll hopefully have experience with some or all of: CI/CD and build systems, containerized runtimes, observability and monitoring, API design, automation frameworks, cloud infrastructure (AWS/GCP), Node.js/TypeScript, and AI-assisted development workflows. We're hiring a Staff Infrastructure Engineer infra and internal platforms that let Hamilton's engineering team ship fast without breaking things that matter.
NewML Infrastructure Engineer Echo NeurotechnologiesML Infrastructure EngineerSan Francisco, CaliforniaThe person who fills this role will design, build, and scale infrastructure to power massive-scale data, modeling, and analysis platforms, playing a critical role in shaping a high-performance, production-grade ML ecosystem to support rapid experimentation with diverse datasets spanning neural signals, behavior, and more. The work will ultimately enable the development of cutting-edge models for neuroscientific discovery and neural decoding, empowering brain-computer interface technology to improve the lives of patients living with severe neurological conditions.
Senior Software Engineer, Data Infrastructure Decagon AI, IncSenior Software Engineer, Data InfrastructureSan Francisco, CA$200,000–$400,000 / yearOur technology enables industry-defining enterprises like Avis Budget Group, Block's Cash App and Square, Chime, Oura Health, and Hunter Douglas to deploy AI agents that power personalized, deeply satisfying interactions across voice, chat, email, SMS, and every other channel. Youll own critical data pipelines and storage layers end‑to‑end, improve reliability and performance, and create paved paths that let every Decagon engineer work confidently with data at scale.
AI Infrastructure Engineer Together AIAI Infrastructure EngineerSan Francisco, CaliforniaYou specialize in systems (operating systems, storage subsystems, networking), while implementing best practices for availability, reliability and scalability, with varied interests in algorithms and distributed systems. We believe open and transparent AI systems will drive innovation and create the best outcomes for society, and together we are on a mission to significantly lower the cost of modern AI systems by co-designing software, hardware, algorithms, and models.
Staff Software Engineer - Infrastructure Skydio IncStaff Software Engineer - InfrastructureSan Mateo, CA$230,000–$275,000 / yearThe Skydio team combines deep expertise in artificial intelligence, best-in-class hardware and software product development, operational excellence, and customer obsession to empower a broader, more diverse audience of drone users, from utility inspectors to first responders, soldiers in battlefield scenarios, and beyond. About the team: Skydio's Cloud infrastructure team is here to ensure that the Skydio Cloud platform is always available to our customers when they need it: whether that's performing a routine bridge inspection or it's aiding in rescue operations during a natural disaster.
Senior Software Engineer - Infrastructure Skydio IncSenior Software Engineer - InfrastructureSan Mateo, CA$170,000–$258,000 / yearThe Skydio team combines deep expertise in artificial intelligence, best-in-class hardware and software product development, operational excellence, and customer obsession to empower a broader, more diverse audience of drone users, from utility inspectors to first responders, soldiers in battlefield scenarios and beyond. About the team: Skydio's Cloud infrastructure team is here to ensure that the Skydio Cloud platform is always available to our customers when they need it most: whether that's performing a routine bridge inspection or it's aiding in rescue operations during a natural disaster.
Software Engineer, Cloud Infrastructure DatologyAISoftware Engineer, Cloud InfrastructureRedwood City, CaliforniaTraining on curated data can dramatically reduce training time and cost ( 7-40x faster training depending on the use case), dramatically increase model performance as if you had trained on >10x more raw data without increasing the cost of training, and allow smaller models with fewer than half the parameters to outperform larger models despite using far less compute at inference time, substantially reducing the cost of deployment. We raised a total of $57.5M in two rounds, a Seed and Series A. Our investors include Felicis Ventures, Radical Ventures, Amplify Partners, Microsoft, Amazon, and AI visionaries like Geoff Hinton, Yann LeCun, Jeff Dean, and many others who deeply understand the importance and difficulty of identifying and optimizing the best possible training data for models.
Software Engineer: Infrastructure ThatchSoftware Engineer: InfrastructureSan Francisco, CaliforniaWe’re a fully distributed early stage company using technology to change the way America does healthcare. We hire engineers across multiple levels and care most about the scope you’ve owned, the impact you’ve had, and how you make decisions.
AI & HPC Infrastructure Engineer Accenture PlcAI & HPC Infrastructure EngineerMountain View, CAArchitect and deploy with NVIDIA platform tools including Base Command Manager (BCM), NGC, NCCL, NVLink, and CUDA along with LLM inference engines (TensorRT-LLM), production serving frameworks (vLLM, SGLang), inference orchestration (Triton Inference Server, NVIDIA Dynamo, llm-d), and GPU benchmarking and validation tools (MLPerf, NCCL tests, fio, iperf) to deploy, tune, profile, and validate AI cluster performance across compute and networking layers including multi-node training and inference workloads. Deploy, configure, and manage XPU-based clusters (GPU, DPU, LPU, CPU) across bare-metal and containerized environments using workload schedulers (Slurm, Run:ai), Kubernetes orchestration, and container platforms to deliver scalable AI infrastructure services including Bare-Metal-aaS, GPUaaS, AIaaS, Token-aaS, model serving, and agentic AI frameworks.
Software Engineer, ML Infrastructure Anysphere IncSoftware Engineer, ML InfrastructureCAAbout the role The ML Infrastructure team builds large-scale compute, storage, and software infrastructure to support Cursors work building the worlds best agentic coding model. This role works closely with ML researchers and engineers to enable their work through improvements to our training framework, systems reliability/performance, and developer experience.
HPC Systems Engineer KLA CorporationHPC Systems EngineerMilpitas, CA$136,300–$231,700 / yearThe AI & Modeling Center of Excellence was setup with the mission of advancing KLA's traditional strengths in physics and data and providing implementation solutions for multiple KLA Inspection and Metrology products targeted at the semiconductor manufacturing industry. As a part of this group, you will be part of a world class team of physicists, HPC system designers, machine learning and application engineers who build cutting edge solutions for modeling complex imaging techniques and semiconductor processes.
Senior Network Systems Engineer Robotic Research OpCo LLCSenior Network Systems EngineerCA$140,000–$185,000 / yearForterra is expanding its mission through Vektor, a next-generation, edge-deployed, software-defined communications and smart data-brokering layer built for disrupted, degraded, intermittent, and low bandwidth (DDIL) environments. Forterra delivers weapons, sensors, and battlefield effects through integrated autonomous networks reaching operational areas faster, safer, and without placing human lives at risk.
Principal Systems Engineer - Cloud Services Stanford Health CarePrincipal Systems Engineer - Cloud ServicesPALO ALTO, CA$79.21–$104.97 / hourAs a Principal System Engineer - Cloud Services, You will oversee Stanford Health Care's cloud infrastructure, managing deployment, optimization, and support of cloud-native applications and services to ensure high performance and security. This role is vital for supporting initiatives related to modern application platforms, generative AI, machine learning, data engineering pipelines, Databricks environments, and DevOps frameworks.
Pv/Bess Systems Engineer JLLPv/Bess Systems EngineerMountain View, CAThis role interacts with sustainability leads, fire life safety teams, and third-party solar/battery installers (e.g., Tesla) to manage the full project lifecycle from design intake to performance telemetry onboarding. Whether you've got deep experience in commercial real estate, skilled trades or technology, or you're looking to apply your relevant experience to a new industry, join our team as we help shape a brighter way forward.
ML Research Engineer (Model Training) MetamorphicML Research Engineer (Model Training)Palo Alto, CaliforniaYou will design the systems that coordinate complex ML workflows across heterogeneous infrastructure, support rapid iteration by researchers, and ensure that models move cleanly from experimentation into robust, low-latency, cost-efficient deployment. Metamorphic is developing new approaches to intelligence by combining machine learning with large-scale experimental neuroscience, informed by the principles that make the brain efficient, flexible, and robust.
Hardware Systems Engineer Mill Industries IncHardware Systems EngineerSan Bruno, CATens of thousands of Mill's residential food recyclers are already helping households divert millions of pounds of food scraps every year, paving the way for our upcoming launch of Mill Commercial-the industry's first end-to-end solution for managing, understanding, and preventing food waste in commercial environments (e.g. We are a small team and you will work cross-functionally with Electrical/Mechanical Engineers, Algorithms and Product Teams to deliver world class industrial products designed to keep food waste out of landfills.
Facilities Systems Engineer Stanford Health CareFacilities Systems EngineerMenlo Park, CA$52.69–$69.82 / hourFIS is made up of multiple departments aligned in focus: Environmental Health & Safety, Facility Field Services, Protective Services, Security Services, Office of Emergency Management, Facilities Engineering- Systems, Operational Technology, Facilities Engineering- Infrastructure, Facilities Services Response Center, and Facilities Administration and Operations. Responsible for an Asset Management plan to monitor assets including buildings and equipment: applies value analysis to repair/replace/redesign and make/buy decisions for Capex and Opex decisions; adheres to the LifeCycle Asset Management process.
Distributed Systems Engineer - Platform Inngest IncDistributed Systems Engineer - PlatformSan Francisco, CAconcurrency over time, or function debounce)Plan and implement improvements on throughput, and latency at hundreds of thousands to millions of requests per secondContribute to systems architecture and infrastructure changes as we growCollaborate with team members to expose internal data across metrics stores, APIs, and customer dashboards we host in our cloud UIWork with backend engineers to design APIs that can be used across the Inngest cloud dashboard, dev server and CLIsDogfood the Inngest product and develop ideas for improvements, features, or integrationsCommunicate with our users through Github, email and DiscordWrite technical specs for features and documentation for our usersIdeal candidateYouve been working on distributed systems for several yearsYouve used Go or similar statically typed languages professionally for two years or moreYouve architected, or been involved in designing, systems that handle scaleYou understand engineering trade-offs and can make correct judgement calls on approaches availableYou understand how to observe, monitor, and maintain the systems you designYou appreciate simplicity, even if its harder to design and buildBonus pointsWork with compliance (SOC2, ISO27001, HIPAA, etc.) and know how to make it serve security approaches and not the other way aroundExperienced or have a solid understanding of networkingUnderstand and have experience managing and maintaining systems, eg. What we build withBackend: Go, Postgres, Redis, Clickhouse, PubSub/Kafka, k8sAPIs: gRPC internally, GraphQL and REST APIs for UIHosted on AWS, GCP and Bare MetalGithub, Linear, Slack, Notion, Figma.
Staff Control Systems Engineer Super Micro Computer, Inc.Staff Control Systems EngineerSan Jose, CA$170,000–$200,000 / yearThis role combines hands-on technical execution (device connectivity, protocol configuration, database integration) with system-level architecture and design, bridging Rack Solutions, Facilities, Production, and IT teams to create a unified, scalable, and data-driven controls environment. Job Summary: The Controls Systems Engineer is responsible for designing, implementing, and maintaining an integrated multi-site controls and automation solution spanning Supermicro's global facilities for rack integration, burn-in, and cooling infrastructure.
Systems Engineer II, AI Solutions & IT Operations Woven Planet Holdings CoSystems Engineer II, AI Solutions & IT OperationsPalo Alto, CA$112,000–$184,000 / yearIn this unique role, you will act as a full-stack owner for our Moveworks ServiceNow Platform-consulting with internal customers to build cutting-edge Agentic AI solutions-while also rolling up your sleeves to serve as an escalation point for complex L2/L3 IT system operations. Full-stack Ownership: Partner deeply with internal customers throughout the entire delivery lifecycle of AI Agents on the Moveworks ServiceNow Platform: Discussions and Planning, Solution Design/Architecture, Building, Tuning, and Prod Launch.
Software Engineer, Data Infrastructure CohereSoftware Engineer, Data InfrastructureSan Francisco, CAWe're training and deploying frontier models for developers and enterprises who are building AI systems to power magical experiences like content generation, semantic search, RAG, and agents. As a Software Engineer, Data Infrastructure, you will: Work directly on petabyte-scale storage infrastructure, and the networking and performance challenges that come with it.
Software Engineer - Maps Infrastructure Applied Intuition IncSoftware Engineer - Maps InfrastructureSunnyvale, CA$125,000–$222,000 / yearApplied Intuition is headquartered in Sunnyvale, California, with offices in Washington, D.C. San Diego; Ft. Walton Beach, Florida; Ann Arbor, Michigan; London; Stuttgart; Munich; Stockholm; Bangalore; Seoul; and Tokyo. Our product suite uses HD maps to solve our customers needs in a multitude of applications: calculating localized information, querying data across global-sized maps, as well as visualizing and inspecting maps.