High Performance Computing Engineer for Workflow Modernization (HPC Engineer 2/3)

Los Alamos National Laboratory

  • Los Alamos, NM
  • 1 day ago

    Highlights

    Due to federal restrictions contained in the current National Defense Authorization Act, citizens of the People's Republic of China-including the special administrative regions of Hong Kong and Macau-as well as citizens of the Islamic Republic of Iran, the Democratic People's Republic of Korea (North Korea), and the Russian Federation, who are not Lawful Permanent Residents ("green card" holders) are prohibited from accessing facilities that support the mission, functions, and operations of national security laboratories and nuclear weapons production facilities, which includes Los Alamos National Laboratory. LANL's High Performance Computing Division supports the Laboratory's mission by managing a world-class supercomputing center, supporting stockpile stewardship for NNSA/DOE and accelerating scientific discovery through some of the world's largest supercomputers.

    Numbers & Facts

    LocationLos Alamos, NM

    Description

    What You Will Do
    This position is open for external candidates only to apply.

    The HPC Division is seeking an engineer to modernize how our Workload Management and Consulting teams do their own work, using AI to make our tools, documentation, and support processes more effective. This is a role between infrastructure and end users - you'll build capabilities that make both our team and our users more productive.

    This position will be filled at either the HPC Engineer 2 or HPC Engineer 3 level, depending on the skills of the selected candidate.

    Core responsibilities:
    • Identify which of the group's existing processes - documentation generation and upkeep, troubleshooting guides - are good candidates for AI-assisted automation, and build that automation. This includes surfacing relevant system and knowledge base content to assist consultants investigating complex user issues.
    • Enable AI workloads on LANL's HPC systems through investigation and deployment of relevant technologies, such as:
      • Slinky, a next-generation collection of projects bridging HPC and AI workloads.
      • RunAI, a scheduler for GPU-centric AI/ML workloads, including capacity sharing and integration with existing HPC resources.
    • Develop and deploy tools on OpenShift, for both direct user consumption and internal division use, including infrastructure that gives users visibility into their workloads.
    • Provide technical support and documentation for the tools and workflows you build, and collaborate with vendors and open-source communities on relevant scheduling, orchestration, and AI tooling.

    The tools you build will serve two audiences: HPC users trying to get work done, and HPC staff trying to support them.

    HPC Engineer 2 (106,400-176,000/year)

    At this level, responsibilities include:
    • Building and iterating on AI-assisted tools for documentation and support workflows, under guidance from senior staff.
    • Assisting with evaluation, deployment, and support of Slinky and RunAI in test and production environments.
    • Developing and deploying tools on OpenShift.
    • Providing day-to-day technical support and troubleshooting for users of these systems.
    • Contributing to documentation and knowledge base content, including using the tools you help build.
    • Collaborating with the Workload Management and Consulting teams to identify workflow pain points worth solving.

    This level applies standard principles and techniques, builds familiarity with new AI tooling and scheduling/orchestration technologies, and exercises independent judgment on well-scoped problems.

    HPC Engineer 3 (128,000-215,900/year)

    In addition to the above, this level includes:
    • Technical ownership of the group's AI-assisted workflow modernization efforts - deciding what to build, how to build it, and how to evaluate whether it's actually helping.
    • Designing more complex tooling that combines AI-assisted automation with observability into user and staff workflows.
    • Technical ownership of significant pieces of the Slinky and/or RunAI evaluation and deployment effort, drawing on hands-on experience with HPC workload management to guide design and integration decisions.
    • Root-cause troubleshooting of complex issues spanning scheduling, orchestration, and the internal tools you build.
    • Representing this work in technical discussions with vendors, other HPC teams, and the broader HPC community.
    • Mentoring junior staff and students on AI tooling, scheduling systems, and container orchestration.

    This level requires deep technical expertise in applying AI tooling to real operational problems, alongside solid grounding in scheduling and container orchestration.

    About the Organization

    LANL's High Performance Computing Division supports the Laboratory's mission by managing a world-class supercomputing center, supporting stockpile stewardship for NNSA/DOE and accelerating scientific discovery through some of the world's largest supercomputers.

    Within HPC, the Environments Group (HPC-ENV) is responsible for the user experience of our supercomputers. This position sits between 2 teams within HPC-ENV:
    • Workload Management is responsible for the schedulers, resource managers, and orchestration layers that determine how and when work runs on LANL's supercomputers, including evaluation of new technologies.
    • Consulting is the single point of contact for HPC customers, providing direct technical support and coordinating institutional knowledge across the division.

    What You Need

    Minimum Job Requirements:

    • Effective written and oral communication skills.
    • Experience with UNIX/Linux systems administration and command-line tools.
    • Experience applying AI/ML tools to practical problems - document generation/summarization, classification, retrieval-augmented systems, or similar.
    • Experience with containers (Docker, Podman, Charliecloud, Singularity/Apptaine, etc..) and container orchestration (OpenShift, Rancher, etc..).
    • Scripting ability (Python, Bash, or similar) for automation and tool development.
    • Ability to communicate across both infrastructure and end-user contexts.

    Additional Requirements for HPC Engineer 3:
    • Demonstrated depth of experience designing and deploying AI-assisted tooling for real operational or support workflows.
    • Experience evaluating new technologies and producing concrete deployment recommendations.
    • Hands-on experience with HPC systems and workload management.

    Education/Experience:

    HPC Engineer 2: Bachelor's degree in Computer Science, Computer Engineering, or related field, and 3 years of relevant experience in HPC, scalable AI computing, or data center environments, or equivalent combination of education and experience.

    HPC Engineer 3: Bachelor's degree in Computer Science, Computer Engineering, or related field, and 6 years of relevant experience in HPC, scalable AI computing, or data center environments, or equivalent combination of education and experience.

    Desired Qualifications:

    Background
    • Experience building tools around large language models - retrieval-augmented generation, fine-tuning, prompt engineering, agentic workflows, or similar - applied to documentation or support use cases.
    • Experience with HPC or cluster workload managers/schedulers (Slurm, PBS, LSF, or similar).
    • GPU scheduling and resource sharing in HPC or cloud environments.
    • Building observability tooling (metrics, logging, dashboards) for distributed or scientific computing workflows.
    • Customer service or public-facing technical support experience.
    • Evaluating software/systems against performance or usability metrics and producing improvement recommendations.

    Technologies
    • LLM tooling and frameworks (LangChain, LlamaIndex, vector databases, local/hosted model serving, etc.).
    • Slurm or other HPC workload managers.
    • RunAI or comparable Kubernetes-native AI/ML scheduling platforms.
    • OpenShift or Kubernetes application development and deployment.
    • Documentation tooling (static site generators, wikis, knowledge base platforms) and how to automate or augment them.
    • Configuration management tools (Ansible, Chef, etc.).
    • CI/CD pipelines (GitLab CI, GitHub Actions, etc.).
    • Monitoring/observability stacks (Prometheus, Grafana, ELK, etc.).

    Institutional Growth
    • Skills or perspectives relevant to the role's goals not listed above, technical or otherwise.
    • Interest in making complex systems and processes usable, not just functional, for the people who depend on them.

    Work Location: This position is hybrid, located in Los Alamos, NM. Hybrid is defined as working partially onsite/partially offsite but within a 2-hour ground commute of this location. Work locations are at the discretion of management and can change with appropriate notice.

    Position commitment: Regular appointment employees are required to serve a period of continuous service in their current position in order to be eligible to apply for posted jobs throughout the Laboratory. If an employee has not served the time required, they may only apply for Laboratory jobs with the documented approval of their Division Leader. The position commitment for this position is 1 year.

    Note to Applicants:

    We encourage applicants to include a cover letter addressing the requirements and qualifications. No single candidate is expected to meet every item below - apply regardless of the number you meet.

    Due to federal restrictions contained in the current National Defense Authorization Act, citizens of the People's Republic of China-including the special administrative regions of Hong Kong and Macau-as well as citizens of the Islamic Republic of Iran, the Democratic People's Republic of Korea (North Korea), and the Russian Federation, who are not Lawful Permanent Residents ("green card" holders) are prohibited from accessing facilities that support the mission, functions, and operations of national security laboratories and nuclear weapons production facilities, which includes Los Alamos National Laboratory.

    Where You Will Work

    Located in beautiful northern New Mexico, Los Alamos National Laboratory (LANL) is a multidisciplinary research institution engaged in strategic science on behalf of national security. Our generous benefits package includes:

    § PPO or High Deductible medical insurance with the same large nationwide network

    § Dental and vision insurance

    § Free basic life and disability insurance

    § Paid childbirth and parental leave

    § Award-winning 401(k) (6% matching plus 3.5% annually)

    § Learning opportunities and tuition assistance

    § Flexible schedules and time off (PTO and holidays)

    § Onsite gyms and wellness programs

    § Extensive relocation packages (outside a 50 mile radius)
    Additional Details

    Directive 206.2 - Employment with Triad requires a favorable decision by NNSA indicating employee is suitable under NNSA Supplemental Directive 206.2. Please note that this requirement applies only to citizens of the United States. Foreign nationals are subject to a similar requirement under DOE Order 142.3A.

    Clearance: Q (Position will be cleared to this level). Selected applicants will be subject to a background investigation conducted by or on behalf of the Federal Government, and must meet eligibility requirements* for access to classified matter. This position requires a Q clearance, and obtaining such clearance requires US Citizenship except in extremely rare circumstances. Dependent upon the position, additional authorization to access classified information may be required, which may or may not be available to dual citizens. Receipt of a Q clearance and additional access authorization ultimately is a decision of the Federal Government and not of Triad.

    *Eligibility requirements: To obtain a clearance, an individual must be at least 18 years of age; U.S. citizenship is required except in very limited circumstances. See DOE Order 472.2 for additional information.

    New-Employment Drug Test: The Laboratory requires successful applicants to complete a new-employment drug test and maintains a substance abuse policy that includes random drug testing. Although New Mexico and other states have legalized the use of marijuana, use and possession of marijuana remain illegal under federal law. A positive drug test for marijuana will result in termination of employment, even if the use was pre-offer.

    Regular position: Term status Laboratory employees applying for regular-status positions are converted to regular status.

    Internal Applicants: Regular appointment employees who have served the required period of continuous service in their current position are eligible to apply for posted jobs throughout the Laboratory. If an employee has not served the required period of continuous service, they may only apply for Laboratory jobs with the documented approval of their Division Leader. Please refer to Policy Policy P701 for applicant eligibility requirements.
    Equal Opportunity: Los Alamos National Laboratory is an equal opportunity employer. All employment practices are based on qualification and merit, without regard to protected categories such as race, color, national origin, ancestry, religion, age, sex, gender identity, sexual orientation, marital status or spousal affiliation, physical or mental disability, medical conditions, pregnancy, status as a protected veteran, genetic information, or citizenship within the limits imposed by federal, state, and local laws and regulations. The Laboratory is also committed to making our workplace accessible to individuals with disabilities and will provide reasonable accommodations, upon request, for individuals to participate in the application and hiring process. To request such an accommodation, please send an email to applyhelp@lanl.gov or call (505)-664-6947.

    Similar Jobs