AI Systems & Infrastructure Engineer

Triune Infomatics

  • San Jose, CA
  • 8 days ago

    Highlights

    You will bridge the gap between GPU kernels and production serving, focusing on distributed inference orchestration, hardware-software co-optimization, and the elimination of system-level bottlenecks to ensure maximum throughput and minimum latency. Profiling & Benchmarking: Systematically identify AI bottlenecks (NVLink, PCIe, HBM bandwidth) and establish rigorous TPS/TTFT benchmarking suites.

    Numbers & Facts

    LocationSan Jose, CA

    Description

    Role: AI Systems & Infrastructure Engineer
    Location: San Jose, CA (Hybrid)
    Duration: 3+ months
     
     
    Summary: We are seeking an expert AI Systems & Infrastructure Engineer to design, build, and optimize the end-to-end environment for large-scale AI models. You will bridge the gap between GPU kernels and production serving, focusing on distributed inference orchestration, hardware-software co-optimization, and the elimination of system-level bottlenecks to ensure maximum throughput and minimum latency.
     
    Key Responsibilities:
    • End-to-End Infrastructure: Architect the full lifecycle from GPU resource allocation to high-performance serving layers.
    • Distributed Orchestration: Implement Tensor and Pipeline Parallelism to deploy massive models across multi-GPU/multi-node clusters.
    • System Optimization: Maximize hardware utilization via advanced AI memory management (KV cache, PagedAttention, quantization).
    • Profiling & Benchmarking: Systematically identify AI bottlenecks (NVLink, PCIe, HBM bandwidth) and establish rigorous TPS/TTFT benchmarking suites.
    • Deployment & Testing: Build AI-specific CI/CD pipelines for automated performance gating, model validation, and regression testing.
    • Co-Design: Align model architectures with target hardware constraints to optimize underlying inference engines.
     
    Technical Qualifications:
    1. Frameworks & Acceleration
    • Core: Expert PyTorch and TensorFlow; proficient in Jupyter/Colab for profiling.
    • Serving Engines: Deep expertise in vLLM, SGLang, Ollama, and Nvidia NIM.
    • Acceleration: Advanced knowledge of Nvidia Dynamo (TorchDynamo) and graph-compilation for execution path optimization.
    1. Systems & Infrastructure
    • Distributed Compute: Proficiency in multi-node communication (NCCL, MPI) and distributed inference strategies.
    • Memory & Precision: Expert in GPU memory layouts, quantization (INT8, FP8, NF4), and PagedAttention.
    • Profiling Tools: Mastery of NVIDIA Nsight, PyTorch Profiler, and Triton for kernel and memory access analysis.
    • Deployment: Experience building AI-centric CI/CD pipelines with automated performance gating and canary deployments.
    1. Hardware & Architecture
    • GPU Architecture: Deep understanding of H100/A100 internals (SMs, Tensor Cores, HBM, NVLink/InfiniBand).
    • Bottleneck Analysis: Ability to diagnose and resolve compute-bound, memory-bound, and I/O-bound workloads.
    • Validation: Experience in stress testing, load balancing, and failover validation for distributed AI nodes.
     
    Preferred Qualifications:
    • Custom CUDA kernel development or Triton optimization.
    • Kubernetes (K8s) for GPU orchestration and scheduling.
    • Experience with distributed training (DeepSpeed, Megatron-LM).
    • OS-level knowledge of memory paging and asynchronous I/O.
     
    Soft Skills:
    • Systems Thinking: Ability to map computational operations directly to physical hardware.
    • Analytical Rigor: Data-driven approach to tuning based on profiles rather than intuition.
    • Collaboration: Ability to translate researcher requirements into concrete infrastructure specs.

    Similar Jobs

    See more jobs