HPC Consultant

Diverse Lynx, LLC

  • Redmond, WA
  • 9 days ago
  • Remote
  • $120,000 Per Year

Highlights

You should have hands-on experience operating clusters at meaningful scale: node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets. Security Notice: Our official website is www.diverselynx.com We do not operate or authorize any other websites representing Diverse Lynx LLC.

Numbers & Facts

LocationRedmond, WA (
Remote
)

Description

Role: HPC Consultant

Location: Remote

Duration: Fulltime

Work Type: Remote-first. Occasional on-site visits to customer offices

Interview Type: Video

Salary-$120k with benefits.

Area Experience Kubernetes Terraform AWS Python HPC/GPU CI/CD Monitoring Troubleshooting Production Support This role is Kubernetes-heavy. You''ll operate multi-cloud platform infrastructure where misconfigurations or failed upgrades translate directly into thousands of lost GPU-hours. The clusters are large enough that Client failure modes are routine.

Responsibilities

  • Operate Kubernetes platforms (EKS, CKS, GKE) at significant scale across providers. You''re responsible for cluster lifecycle, node pool management, networking policy, and maintaining stability during rapid growth.
  • Provision HPC infrastructure through CI/CD system across AWS, CoreWeave, Eligible to workP, and OCI, with additional providers to be expanded in the near future.
  • Manage job scheduling to allocate GPU compute across training and inference workloads.
  • Define and maintain SLIs/SLOs. Build monitoring and alerting. Participate in severity escalation response and author post-incident reviews.
  • Coordinate daily with Networking, Storage, Security, and AI/ML platform teams.

Requirements

  • 4+ years in infrastructure engineering, cloud platforms, or HPC.
  • Kubernetes is the core requirement. You should have hands-on experience operating clusters at meaningful scale: node pool sizing, scheduler debugging, CNI troubleshooting, and rolling upgrades across large fleets. Candidates whose K8s experience is limited to small or local environments are unlikely to be a fit.
  • Terraform proficiency. You''ll write and review infrastructure-as-code daily.
  • Working knowledge of AWS (EC2, S3, EFS, FSx for Lustre).
  • Python for tooling and automation.

Disclaimer: Diverse Lynx LLC is an Equal Opportunity Employer. All applicants and employees are evaluated without discrimination, based solely on their qualifications, ability, competence and performance. This email and its attachments may contain confidential or proprietary information and is intended only for the recipient(s). If you received this message in error, please disregard it and notify the sender. If you no longer wish to receive our communications, you may unsubscribe at any time.

Security Notice: Our official website is www.diverselynx.com We do not operate or authorize any other websites representing Diverse Lynx LLC.

Similar Jobs