Milestone Technologies Inc. logo

Data Center Technician L2 - GPU Specialist

Milestone Technologies Inc.

  • Reno, NV
  • 30 days ago

    Highlights

    The Data Center Operations Technician II - GPU Specialist is responsible for supporting and maintaining highly available GPU-based compute environments, engineering labs, and data center infrastructure. The ideal candidate combines strong data center operations experience with a deep understanding of GPU technologies, server hardware, Linux/Windows administration, and large-scale test infrastructure.

    Numbers & Facts

    LocationReno, NV
    IndustryOther/Not Classified
    Company Size2,000 to 2,499 employees
    Year Founded1997
    Websitehttp://www.milestone.tech

    Description

    The Data Center Operations Technician II - GPU Specialist is responsible for supporting and maintaining highly available GPU-based compute environments, engineering labs, and data center infrastructure. This role partners closely with hardware, software, QA, and systems engineering teams to deploy, troubleshoot, and optimize next-generation computing platforms. The ideal candidate combines strong data center operations experience with a deep understanding of GPU technologies, server hardware, Linux/Windows administration, and large-scale test infrastructure.

    Key Responsibilities

    Compute Farm & Infrastructure Operations

    • Manage and maintain a high-performance compute farm consisting of builders, packagers, testers, and supporting infrastructure.
    • Monitor system health, availability, and performance to ensure operational excellence and SLA compliance.
    • Lead system recovery efforts and incident response activities to minimize downtime and restore services quickly.
    • Support deployment, configuration, and lifecycle management of GPU servers, workstations, and test systems.
    • Perform rack, stack, cabling, hardware installation, and equipment decommissioning activities within the data center.

    Engineering Support

    • Collaborate closely with system architects, hardware engineers, software engineers, QA teams, and platform operations teams to develop, test, debug, and release next-generation products.
    • Troubleshoot hardware, software, networking, and infrastructure issues impacting engineering and validation environments.
    • Provide technical support for GPU systems, PCBs, servers, storage systems, and network-connected devices.
    • Assist engineering teams with validation, benchmarking, and deployment activities for new technologies and platforms.

    Process Improvement & Documentation

    • Gather operational metrics and performance data to identify trends, risks, and improvement opportunities.
    • Develop, maintain, and enhance Standard Operating Procedures (SOPs), runbooks, and technical documentation.
    • Drive continuous improvement initiatives that increase availability, throughput, operational efficiency, and test accuracy.
    • Participate in change management activities and ensure documentation is kept current.

    Systems Administration & Automation

    • Support and troubleshoot Linux, Windows, and macOS environments.
    • Utilize scripting and automation tools to streamline operational tasks and improve scalability.
    • Maintain accurate asset and infrastructure records using DCIM systems.
    • Assist with infrastructure automation and configuration management initiatives.

    Required Qualifications

    • Associate's degree or Bachelor's degree in Engineering, Information Technology, Computer Science, or a related technical field; equivalent experience will be considered.
    • 5+ years of experience supporting data center operations, engineering labs, high-performance computing environments, or related technical infrastructure.
    • Experience working with GPU-based systems, PCBs, servers, and large-scale system deployments.
    • Proficiency with DCIM platforms such as Nautobot or similar infrastructure management tools.
    • Experience with scripting and automation technologies including Shell, Python, and Ansible.
    • Working knowledge of networking fundamentals and protocols including: TCP/IP DNS NFS SSL/TLS
    • Experience administering and troubleshooting: Linux Windows macOS
    • Strong troubleshooting and problem-solving skills across hardware, operating systems, networking, and infrastructure.
    • Excellent written and verbal communication skills with the ability to present technical concepts to non-technical audiences.
    • Strong teamwork skills and the ability to work effectively in cross-functional engineering environments.

    Preferred Qualifications

    • Experience managing High Performance Computing (HPC) environments.
    • Experience utilizing cluster management and workload scheduling platforms such as: Bright Cluster Manager (BCM) Slurm
    • Industry certifications such as CCNA or equivalent networking certifications.
    • Advanced Windows and Linux systems administration experience.
    • Understanding of modern data center architecture, including: Compute infrastructure Storage platforms Networking systems
    • Knowledge of data center facilities infrastructure with emphasis on liquid-cooled environments.
    • Experience supporting AI, machine learning, or GPU-intensive workloads.
    • Strong mechanical aptitude and comfort performing hands-on hardware installation, maintenance, and repair tasks.

    Core Competencies

    • Data Center Operations
    • GPU Infrastructure Management
    • Linux & Windows Administration
    • Hardware Troubleshooting
    • Network Fundamentals
    • Automation & Scripting
    • HPC Cluster Support
    • Incident Response
    • Documentation & SOP Development
    • Cross-Functional Collaboration
    • Continuous Improvement
    • Customer & Engineering Partner Support

    This position may require access to hardware, software, technology, or technical data subject to U.S. export control laws, including the Export Administration Regulations (EAR) and, where applicable, the International Traffic in Arms Regulations (ITAR). Any offer, assignment, or continued access to controlled items is contingent upon the company's determination that the individual is legally authorized to access such items or that any required government authorization can be obtained.

    #LI-TK3

    About Company

    At Milestone, we know IT, and we’re consistently driving innovation in infrastructure operations to improve the overall customer experience. As a Managed Services Provider (MSP), we use technology intelligently to make IT infrastructures smarter, streamlined, and ultimately, more successful.

    Our seasoned professionals deliver services based on Milestone’s best practices and service delivery framework. By leveraging our vast knowledge base to execute initiatives, we deliver both short-term and long-term value to your company and apply continuous service improvement to deliver transformational benefits to IT. With Intelligent Automation, Milestone helps businesses further accelerate their IT transformation. The result is a sharper focus on business objectives and a dramatic improvement in employee productivity. Through our key technology partnerships and our people-first approach, Milestone continues to deliver industry-leading innovation to our clients.

    Since our inception in 1997, our clients have benefited from more efficient operations and a renewed focus on employee development and business innovation. When founder, Prem Chand, started Milestone Technologies, Inc. he aimed to solve the growing problem of IT Relocation for Silicon Valley businesses. Today, with more than 2,000 employees serving a substantial client base of over 200 companies worldwide, we are following our mission of revolutionizing the way IT is deployed around the globe.

    Similar Jobs

    See more jobs