Research Engineer - Training Performance & ML Compilation (Torch Compile) - Seed Infra

    Highlights

    Experience with the PyTorch compilation stack, meeting at least one of the following: direct experience using, debugging, or extending Inductor or FX; proficiency in Triton kernel development; or solid experience with PyTorch computation graph work (graph optimization, graph capture, operator fusion). Conduct performance profiling and analysis of large-scale training jobs; identify and resolve bottlenecks in collaboration with research and infrastructure teams.

    Numbers & Facts

    LocationSeattle, WA

    Description

    About the Team The Seed Infrastructures team oversees the distributed training, reinforcement learning framework, high-performance inference, and heterogeneous hardware compilation technologies for AI foundation models.

    Responsibilities

    • Optimize training performance for large-scale foundation models through compiler-level techniques, including graph optimization, operator fusion, and kernel generation.
    • Develop and extend ML compilation capabilities based on the PyTorch compilation stack (e.g. FX, Dynamo, Inductor) to improve training efficiency across heterogeneous GPU platforms.
    • Design and optimize high-performance GPU kernels for training workloads.
    • Conduct performance profiling and analysis of large-scale training jobs; identify and resolve bottlenecks in collaboration with research and infrastructure teams.Minimum Qualification(s)
    • Bachelor's degree or above in Computer Science, Electrical Engineering, or a related field.
    • Strong proficiency in C/C++ and Python; solid foundations in algorithms, data structures, and systems programming.
    • Hands-on experience in training-side performance optimization for deep learning workloads.
    • Hands-on experience writing and optimizing GPU kernels (e.g., CUDA, Triton).
    • Experience with the PyTorch compilation stack, meeting at least one of the following: direct experience using, debugging, or extending Inductor or FX; proficiency in Triton kernel development; or solid experience with PyTorch computation graph work (graph optimization, graph capture, operator fusion).

    Preferred Qualification(s)

    • Experience with TorchDynamo or bytecode-level program transformation.
    • Experience with Triton compiler internals or other ML compiler backends (e.g., MLIR, LLVM).
    • Contributions to related open-source projects (e.g., PyTorch, Triton, FlashAttention).
    • Publications in relevant venues (e.g., MLSys, OSDI, ASPLOS).

    Similar Jobs

    See more jobs