Student Researcher (AI Foundation Models Infrastructure - Seed Infra) - 2026 Start (PhD)

    Highlights

    As an Infrastructure Intern, you may work on one or more of the following areas: Design and optimize large-scale distributed training systems (e.g., data/model/pipeline parallelism, memory efficiency, fault tolerance). Build tooling and automation to improve developer productivity and system reliabilityMinimum Qualifications: Currently pursuing a PhD degree in Computer Science, Electrical Engineering, or related technical fields.

    Numbers & Facts

    LocationSan Jose, CA

    Description

    About the team The Seed Infrastructures team oversees the distributed training, reinforcement learning framework, high-performance inference, and heterogeneous hardware compilation technologies for AI foundation models.

    Responsibilities

    • As an Infrastructure Intern, you may work on one or more of the following areas:
    • Design and optimize large-scale distributed training systems (e.g., data/model/pipeline parallelism, memory efficiency, fault tolerance)
    • Contribute to reinforcement learning training frameworks and large-scale post-training systems
    • Improve inference performance, latency, and throughput for foundation models
    • Develop compiler or runtime optimizations for heterogeneous hardware (GPU/accelerator)
    • Work on system-level performance analysis, profiling, and bottleneck diagnosis
    • Build tooling and automation to improve developer productivity and system reliabilityMinimum Qualifications:
    • Currently pursuing a PhD degree in Computer Science, Electrical Engineering, or related technical fields
    • Strong programming skills in Python and/or C++
    • Solid understanding of systems, distributed computing, machine learning systems, or performance optimization
    • Experience with one or more of the following:
    • Distributed training frameworks (e.g., PyTorch FSDP, Megatron-style parallelism)
    • Reinforcement learning training systems
    • GPU programming (CUDA, Triton) or compiler technologies

    Preferred Qualifications:

    • Experience working on large-scale ML systems or infrastructure projects
    • Contributions to open-source ML systems or performance tooling
    • Publications in ML systems, distributed systems, or related areas (a plus but not required)

    As a condition of employment, all successful candidates must be able to establish authorization to work in the United States. For this position, the Company does not provide sponsorship or any immigration-related benefits.

    Similar Jobs

    See more jobs