Senior HPC Networking Engineer

Mirantis Inc

  • CA
  • 13 days ago

    Highlights

    Work with an established Silicon Valley leader in the cloud infrastructure industry; Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies; Be a part of cutting-edge, open-source innovation; Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued; Professional development and training; Attend conferences and working groups; Company outings, happy hours, hackathons, and tech talks; Receive a competitive compensation package with a strong benefits plan. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment-on-premises, in the cloud, at the edge, or in sovereign data centers.

    Numbers & Facts

    LocationCA

    Description

    Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. By combining open source innovation with deep expertise in Kubernetes orchestration, Mirantis empowers platform engineering teams to deliver composable, production-ready developer platforms across any environment-on-premises, in the cloud, at the edge, or in sovereign data centers. As enterprises navigate the growing complexity of AI-driven workloads, Mirantis delivers the automation, GPU orchestration, and policy-driven control needed to manage infrastructure with confidence and agility. Committed to open standards and freedom from lock-in, Mirantis ensures that customers retain full control of their infrastructure strategy. https://www.mirantis.com/

    Location: US

    Employment Type: Full-time

    Role Overview:

    We are seeking a highly skilled Senior HPC Networking Engineer to design, deploy, manage, and troubleshoot high-performance networking environments. The ideal candidate will have deep expertise in InfiniBand technologies, strong general networking knowledge, and hands-on experience with Fortinet solutions. You will play a critical role in ensuring the performance, reliability, and scalability of HPC infrastructure.

    Key Responsibilities:

    • Design, deploy, and maintain high-performance network infrastructures for HPC environments, with a strong focus on InfiniBand fabrics.

    • Troubleshoot complex network issues across InfiniBand and Ethernet environments, ensuring minimal downtime and optimal performance.

    • Manage and optimize InfiniBand components, including switches, HCAs, subnet managers, and fabric configurations.

    • Perform performance tuning, monitoring, and capacity planning for HPC networking systems.

    • Implement and maintain network security using Fortinet solutions (FortiGate, FortiManager, FortiAnalyzer).

    • Diagnose and resolve issues related to routing, switching, latency, and throughput across hybrid network environments.

    • Collaborate with compute, storage, and platform teams to support HPC workloads and cluster operations.

    • Develop and maintain documentation for network architecture, configurations, and operational procedures.

    • Participate in on-call rotations and provide escalation support for critical incidents.

    • Lead or contribute to network upgrades, migrations, and new deployments.

    • You will actively troubleshoot and resolve daily customer incident tickets to keep massive GPU training runs moving.

    Build, operate, and scale next-generation GPU infrastructure:

    • You will be hands-on on the front lines driving daily triage, incident resolution, and SLA management for the world's most advanced NVIDIA clusters, InfiniBand/RoCE fabrics, and AI workloads.

    Build the playbook, then grow into the platform:

    • Designed for engineers energized by standing up new operations from the ground up, this role offers a direct trajectory from operationalizing bare-metal clusters to driving platform and AI capabilities as we scale.

    Required:

    • 5+ years of experience in network engineering, with a focus on HPC or data center environments.

    • Strong hands-on experience with InfiniBand technologies (e.g., Mellanox/NVIDIA).

    • Solid understanding of networking fundamentals: TCP/IP, routing protocols (BGP, OSPF), VLANs, QoS, and network design.

    • Proven experience deploying and troubleshooting Fortinet solutions (FortiGate, FortiManager, VPNs, firewall policies).

    • Experience with network performance analysis and troubleshooting tools.

    • Familiarity with Linux systems and scripting for automation (e.g., Bash, Python).

    • Strong analytical and problem-solving skills.

    Preferred:

    • Experience with large-scale HPC clusters or AI/ML infrastructure.

    • Knowledge of RDMA, MPI, and low-latency networking concepts.

    • Certifications such as FCSS/FCNSP (Fortinet), CCNP/CCIE, or equivalent.

    • Experience with automation and Infrastructure as Code tools (e.g., Ansible, Terraform).

    Soft Skills:

    • Strong communication and collaboration skills.

    • Ability to work independently and handle complex technical challenges.

    • Detail-oriented with a proactive approach to problem-solving.

    What We Offer:

    • Opportunity to work on cutting-edge HPC infrastructure.

    • Collaborative and innovative work environment.

    • Competitive salary and benefits package.

    What does Mirantis offer you?

    • Work with an established Silicon Valley leader in the cloud infrastructure industry;
    • Work with exceptionally passionate, talented and engaging colleagues, helping Fortune 500 and Global 2000 customers implement next-generation cloud technologies;
    • Be a part of cutting-edge, open-source innovation;
    • Thrive in the high-energy environment of a young company where openness, collaboration, risk-taking, and continuous growth are valued;
    • Professional development and training;
    • Attend conferences and working groups;
    • Company outings, happy hours, hackathons, and tech talks;
    • Receive a competitive compensation package with a strong benefits plan.

    We are a Leader for Container Management in G2 (#2 after AWS)!

    Similar Jobs

    See more jobs