Senior Research Scientist/Software Engineer, LLM/Agent Platform (TikTok-Content Ecology AI Innovation & Platform)

TikTok Inc

  • San Jose, CA
  • 11 days ago

    Highlights

    You will design and implement novel LLM/Agent frameworks with self-improving capabilities and architect the robust, high-performance infrastructure that brings these LLMs and agents to life at a global scale. Experience with large-scale model training and inference, including distributed training, KV cache-aware serving, GPU/accelerator optimization, and high-performance networking (e.g., RDMA, NCCL).

    Numbers & Facts

    LocationSan Jose, CA

    Description

    The Content Ecology Algorithm Team drives TikTok's AI innovations in LLMs, NLP, Computer Vision (CV), multimodal learning, and recommendation algorithms. We develop cutting-edge AI capabilities that power multiple business lines.

    We are seeking exceptional, experienced Research Scientists / Architects in the areas of LLM / Agent Platform. This is a unique role for innovators who are passionate about building the "bones" (scalable infrastructure) of next-generation AI and Agent. You will design and implement novel LLM/Agent frameworks with self-improving capabilities and architect the robust, high-performance infrastructure that brings these LLMs and agents to life at a global scale. This is an opportunity to shape the future of AI at one of the world s most dynamic technology companies.

    What You'll Do

    • Architect and build standardized, configurable, and reusable pipelines for the entire lifecycle of models and agents-from data processing and training to deployment, monitoring, and governance.
    • Partner with algorithm teams to understand their needs and provide a world-class infrastructure platform that accelerates their research and development cycles.
    • Build robust observability and evaluation frameworks to ensure the reproducibility, reliability, and cost-efficiency of AI workloads at scale.
    • Design and implement core platform infrastructure, including model/agent registries, feature stores, and high-throughput retrieval/RAG systems.
    • Design, build, and optimize advanced Agentic AI systems, focusing on core components like planning, tool use, and memory. Minimum Qualifications
    • BS/BA or Master in Computer Science or related technical field or equivalent technical experience
    • 5+ years of hands-on experience in software engineering, with a focus on machine learning, distributed systems, or AI infrastructure.
    • Strong proficiency in integrating AI tools into knowledge discovery and research workflows.
    • Familiarity with building robust evaluation frameworks and ensuring experimental reproducibility.
    • Expertise in deep learning frameworks and tensor libraries like PyTorch, Tensorflow, JAX/FLAX
    • Solid understanding of machine learning fundamentals and the modern AI stack.
    • Excellent communication skills to collaborate across teams.

    Preferred Qualifications:

    • PhD in Computer Science or related technical discipline.
    • 5+ years of experience as an architect, or technical leadership position
    • Experience with the ML infrastructure ecosystem, including GPU scheduling, model serving (Triton, TensorRT-LLM), vector databases (FAISS, Milvus), and MLOps principles.
    • Experience with large-scale model training and inference, including distributed training, KV cache-aware serving, GPU/accelerator optimization, and high-performance networking (e.g., RDMA, NCCL).
    • Experience with performance optimization of large model training and inference (e.g., DeepSpeed/ZeRO, vLLM).
    • Deep knowledge of agent architectures, including planning, tool use (e.g., LangChain, LlamaIndex), and memory systems.
    • Publications in systems and/or machine learning conferences (e.g., NeurIPS, OSDI, SOSP, ASPLOS, MLSys).

    Similar Jobs

    See more jobs