High Performance Computing (HPC) Engineer/Childress TX- 6= months contract

Suncap Technology, Inc.

  • Childress, TX
  • 25 days ago

    Highlights

    We are seeking a highly skilled High-Performance Computing (HPC) Engineer to design, deploy, optimize, and support Dell-based HPC and AI infrastructure solutions for enterprise, healthcare, research, education, manufacturing, and government customers. The ideal candidate will possess strong technical expertise in HPC architecture, Linux administration, clustering technologies, high-speed networking, storage systems, and workload management.

    Numbers & Facts

    LocationChildress, TX

    Description

    P

    osition: High-Performance Computing (HPC) Engineer

    Location: Childress TX ( One of the main locations)

    Requirement: Up to 50% travel throughout the United States for customer meetings, solution deployments, health checks, and executive briefings.

    Duration: 6 months or longer

    One of the main locations will be Childress-TX. Open to 50% travel the right candidate. 100% onsite is preferred.

    Description:

    Level: Mid-Senior Engineer
    Reports To: Services Delivery Manager / Practice Director

    Position Overview

    We are seeking a highly skilled High-Performance Computing (HPC) Engineer to design, deploy, optimize, and support Dell-based HPC and AI infrastructure solutions for enterprise, healthcare, research, education, manufacturing, and government customers.

    The ideal candidate will possess strong technical expertise in HPC architecture, Linux administration, clustering technologies, high-speed networking, storage systems, and workload management. This is a high-visibility customer-facing role requiring exceptional communication skills and the ability to engage with technical decision-makers, architects, and executive stakeholders.

    This position involves significant customer interaction and travel to support solution design workshops, installations, migrations, performance tuning, and ongoing technical advisory services.


    Key Responsibilities

    Solution Design & Architecture

    • Design and architect Dell HPC solutions utilizing:
      • Dell PowerEdge Servers
      • Dell PowerScale (Isilon)
      • Dell PowerStore
      • Dell ObjectScale/ECS
      • Dell Integrated Rack Solutions
    • Develop scalable compute, storage, and networking architectures for HPC and AI environments.
    • Perform capacity planning, workload analysis, and performance assessments.

    Deployment & Implementation

    • Install, configure, and integrate HPC clusters at customer locations.
    • Deploy and manage Linux-based clustered environments.
    • Configure high-speed interconnects including:
      • InfiniBand
      • Ethernet 100/200/400GbE
    • Deploy and support shared storage architectures.

    Performance Optimization

    • Conduct benchmarking and performance analysis.
    • Optimize compute, memory, storage, and network utilization.
    • Identify and resolve bottlenecks affecting HPC workloads.
    • Support customer application tuning initiatives.

    Customer Engagement

    • Lead technical discussions with customer engineering and IT teams.
    • Deliver architecture reviews and best-practice recommendations.
    • Serve as a trusted advisor throughout the project lifecycle.
    • Provide executive-level technical presentations when required.

    Operations & Support

    • Troubleshoot complex hardware, software, networking, and storage issues.
    • Participate in escalated support engagements.
    • Create technical documentation, runbooks, and implementation guides.
    • Assist with disaster recovery and business continuity planning.

    Collaboration

    • Work closely with Dell engineering, product specialists, account teams, and partner organizations.
    • Support proof-of-concept engagements and technology demonstrations.
    • Mentor junior engineers and share technical best practices.

    Required Technical Skills

    Operating Systems

    • Red Hat Enterprise Linux (RHEL)
    • Rocky Linux
    • AlmaLinux
    • SUSE Linux Enterprise Server (SLES)

    HPC Cluster Technologies

    • Slurm Workload Manager
    • OpenHPC
    • Bright Cluster Manager (preferred)
    • Warewulf
    • xCAT

    Storage Technologies

    • Dell PowerScale / Isilon
    • NFS
    • Lustre
    • BeeGFS
    • GPFS / IBM Spectrum Scale
    • Parallel file systems

    Networking

    • InfiniBand
    • RoCE
    • High-speed Ethernet (100/200/400GbE)
    • Network performance analysis
    • RDMA technologies

    HPC & AI Workloads

    • MPI (OpenMPI, Client MPI)
    • GPU Computing
    • CUDA
    • NVIDIA GPU Platforms
    • AI/ML infrastructure deployment
    • Scientific and engineering applications

    Virtualization & Cloud

    • VMware vSphere
    • Kubernetes
    • OpenShift
    • Hybrid cloud HPC architectures
    • Azure and AWS HPC services (preferred)

    Automation & Scripting

    • Python
    • Bash
    • Ansible
    • Terraform (preferred)
    • Git

    Monitoring & Management

    • Grafana
    • Prometheus
    • Nagios
    • Dell OpenManage Enterprise
    • Performance monitoring and capacity reporting

    Required Qualifications

    • Bachelor's degree in Computer Science, Engineering, Information Systems, or related field.
    • 5-10+ years of experience supporting enterprise Linux and infrastructure environments.
    • Minimum 3-5 years of hands-on HPC deployment and administration experience.
    • Strong troubleshooting skills across servers, storage, networking, and operating systems.
    • Experience working directly with enterprise customers in a consulting or professional services capacity.
    • Excellent verbal, written, and presentation skills.
    • Ability to manage multiple customer projects simultaneously.

    Preferred Qualifications

    • Experience with Dell HPC environments.
    • NVIDIA certification or GPU deployment experience.
    • Red Hat Certification (RHCSA/RHCE).
    • Dell Technologies Certifications.
    • Experience supporting AI and machine learning infrastructure.
    • Experience with research, life sciences, manufacturing, financial services, or government HPC environments.

    Success Factors

    The successful candidate will:

    • Be viewed as a trusted advisor by strategic customers.
    • Independently lead complex HPC deployments.
    • Resolve challenging performance and infrastructure issues.
    • Communicate effectively with both engineers and executives.
    • Thrive in a fast-paced, customer-facing environment.

    Travel Requirement: Up to 50% travel throughout the United States for customer meetings, solution deployments, health checks, and executive briefings.

    This role is ideal for a senior infrastructure engineer, Linux architect, or HPC specialist looking to work with highly visible enterprise customers and cutting-edge Dell HPC and AI technologies.












    Similar Jobs