Research Engineer - Midtraining

Periodic Labs

  • Menlo Park, California
  • 7 days ago

    Highlights

    As a Midtraining Research Engineer, you'll take base models and improve their scientific reasoning: curating and generating data, building evals, and running large-scale training experiments. Design and run large-scale training experiments, partnering with supercompute engineers to scale efficiently across thousands of GPUs.

    Numbers & Facts

    LocationMenlo Park, California
    Websitehttps://periodic.com

    Description

    We're an AI and physical sciences company building state-of-the-art models to accelerate breakthroughs across materials, energy, and beyond. Backed by world-class investors and growing rapidly, we operate at the pace the frontier requires. Our team brings deep expertise, genuine ownership, and a drive to push the boundaries of what's scientifically possible.

    About the Role

    We're training frontier models to develop deep scientific knowledge and reasoning for scientific discovery. As a Midtraining Research Engineer, you'll take base models and improve their scientific reasoning: curating and generating data, building evals, and running large-scale training experiments. Your work will also lay the groundwork for our pre-training efforts down the line.

    What You'll Do

    • Identify, process, and curate novel sources of scientific data for large-scale model training.

    • Generate high-quality synthetic data to fill gaps in scientific knowledge and reasoning.

    • Build evaluations that correlate with downstream scientific task performance, working closely with RL researchers, physicists, and chemists.

    • Develop and apply techniques such as self-distillation and on-policy distillation to improve model capability.

    • Design and run large-scale training experiments, partnering with supercompute engineers to scale efficiently across thousands of GPUs.

    • Build tools for yourself and the team to investigate how data choices shape model intelligence.

    You Will Thrive in This Role If You Have

    • Experience training LLMs on curated mixes of trillions of tokens.

    • Experience on a dedicated evals team supporting a large production training run.

    • Hands-on use of self-distillation, on-policy distillation, or similar methods in a real training pipeline.

    • Experience with scaling laws and compute-optimal hyperparameters.

    • Comfort working across data, evals, and training infrastructure.

    Especially Strong Candidates May Also Have

    • Experience optimizing throughput and reliability for large-scale distributed training runs.

    • A background in AI for science or training on specialized domain data (e.g., protein, materials, or other scientific datasets).

    • Experience creating evals or synthetic data for non verifiable tasks and tracking performance over live runs.

    Mechanics

    • Minimum education: Bachelor's degree or similar experience

    • Location: Menlo Park, CA (Soon: San Francisco, too)

    • Compensation: $250,000–$350,000 + equity

    • Visa sponsorship: Yes, we sponsor visas and will do everything we can to assist in this process.

    Similar Jobs

    See more jobs