Member of Technical Staff, ML Inference Engineering

Sanas

  • Palo Alto, California
  • 30+ days ago

    Highlights

    Sanas is bringing real-time speech and language models on-premise — deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them.

    Numbers & Facts

    LocationPalo Alto, California

    Description

    About the Role

    Sanas is bringing real-time speech and language models on-premise — deployed at scale directly inside sovereign data centers, not served from behind a hosted cloud endpoint. It's one of the most demanding environments in the industry: strict latency budgets, massive concurrency, and infrastructure that needs to be private and reliable.

    We're looking for a deeply hands-on, senior engineer to help lead that build. This is someone who shapes core infrastructure and architecture decisions rather than just executing against a specification, and who naturally raises the level of the engineers working alongside them.

    What You'll Do

    Performance Optimization

    • Optimize system and GPU performance for high-throughput AI workloads across multi-node training and inference
    • Analyze and improve latency, throughput, memory usage, and compute efficiency
    • Profile system performance to detect and resolve GPU- and kernel-level bottlenecks
    • Implement low-level optimizations using CUDA, Triton, and other performance tooling
    • Improve support for mixed precision, quantization, and model graph optimization
    • Build and maintain performance benchmarking and monitoring infrastructure
    • Scale inference and training systems across multi-GPU, multi-node environments

    Inference Systems & Reliability

    • Own and evolve our inference engine, enabling reliability and performance at scale
    • Develop and optimize runtime inference services for large-scale AI applications
    • Implement robust, fault-tolerant systems for data ingestion and processing

    Requirements

    Must-have:

    • 5+ years of experience writing high-quality, high-performance code
    • Familiarity with NVIDIA GPU architecture and CUDA
    • Fluency in the LLM serving stack, from kernels and quantization up to schedulers and autoscaling
    • A research-leaning or systems background in LLM, Speech-to-Text, Text-to-Speech, or Speech-to-Speech inference, with work you can point to
    • A record of shipping research or systems that other people build on, whether in a lab or in industry

    Nice-to-have:

    • Experience serving low-precision (FP4/FP8) models, multiple LoRA adapters within one model instance (Multi-LoRA), or models distributed across several GPU nodes
    • Experience developing large-scale, high-load production systems
    • Experience maintaining or contributing to open-source ML projects
    • Experience managing machine learning workloads on Kubernetes clusters
    • Experience with InfiniBand or RoCE networking
    • Experience with bare-metal provisioning and lifecycle management
    • Experience operating large-scale AI training or inference clusters
    • Experience with hardware health monitoring and predictive failure detection
    • Experience with distributed storage systems

    Similar Jobs

    See more jobs