Highlights

KEY RESPONSIBILITIES Liquid Cooling Operations Operate, monitor, and maintain Coolant Distribution Units (CDUs) including flow rate management (target: ~385 L/min per rack system), temperature setpoints, and alarm response. Networking & Fabric Infrastructure Patch and manage fiber/copper connections within the Broadcom Tomahawk 6-based switch fabric supporting single-hop, multi-plane GPU-to-GPU interconnect.

Numbers & Facts

LocationAUSTIN, TX

Description

KEY RESPONSIBILITIES
Liquid Cooling Operations
  • Operate, monitor, and maintain Coolant Distribution Units (CDUs) including flow rate management (target: ~385 L/min per rack system), temperature setpoints, and alarm response.
  • Install, replace, and pressure-test blind-mate quick-disconnect (QD) connectors on compute and switch trays without interrupting adjacent rack cooling circuits.
  • Perform leak detection, containment, and remediation in accordance with site EHS protocols.
  • Maintain secondary loop plumbing (supply/return manifolds, flexible hose assemblies, isolation valves); coordinate with facilities for primary loop support.
  • Support rack-level thermal assessments; contribute to coolant quality monitoring (pH, conductivity, biocide levels).
High-Density Rack & Server Hardware
  • Install and service AMD Helios-class compute trays (~77 kg each; 576 differential connections) using approved lift equipment and torque specifications.
  • Install and service switch trays (1,728 connections; ~310 kg insertion force) using lever-assist handles per OEM guidance.
  • Replace field-replaceable units (FRUs): AMD Instinct MI450/MI455X GPU modules, HBM4 memory, EPYC CPU+DIMM assemblies, NVMe drives, power modules, and fan trays.
  • Execute rack-level hot-swap operations on busbar and power shelf components following LOTO and high-voltage DC (HVDC) safety procedures.
  • Cable-manage high-speed optical interconnects (QSFP-DD / OSFP) and DAC/Client assemblies for UALink, Ethernet, and out-of-band management fabrics.
Networking & Fabric Infrastructure
  • Patch and manage fiber/copper connections within the Broadcom Tomahawk 6-based switch fabric supporting single-hop, multi-plane GPU-to-GPU interconnect.
  • Validate optical power budgets and verify high-speed links following tray moves, adds, or changes.
  • Coordinate with network engineering for fabric topology changes and UALink / UEC configuration validation.
Monitoring, Diagnostics & Firmware
  • Use BMC/Redfish/IPMI tools to monitor GPU and CPU health, review system event logs, and identify hardware faults proactively.
  • Execute firmware update workflows (BIOS/UEFI, BMC, NIC, GPU firmware) following approved change management processes.
  • Run AMD ROCm-based diagnostic utilities (rocm-smi, GPU benchmarks) to validate compute readiness post-maintenance.
  • Interface with DCIM and Client platforms (e.g., Schneider EcoStruxure, Motivair CDU telemetry) to track PUE, coolant flow, and rack-level power draw (up to 246 kW/rack).
Documentation & Compliance
  • Maintain accurate asset inventory, maintenance logs, and incident records in the CMDB/ticketing system.
  • Follow EHS, physical security, and datacenter access procedures at all times.
  • Contribute to runbooks, SOP updates, and lessons-learned documentation after significant incidents or deployments.
  • Participate in on-call rotation for after-hours critical hardware incidents.
 
EXPERIENCE REQUIREMENTS
Minimum: 4–5 years in a datacenter infrastructure or datacenter operations (DCO) technician role
Must Have
  • 4+ years hands-on experience with Direct Liquid Cooling (DLC) systems in production datacenter environments, including CDU operation and coolant loop maintenance.
  • 4+ years experience with enterprise server hardware installation and break/fix servicing at component level (CPU, DIMM, GPU, storage, NIC).
  • Demonstrated experience with high-density rack platforms (&Client;20 kW/rack) including cabling, power distribution, and liquid-cooling integration.
  • Proven ability to follow and enforce LOTO, EHS, and physical security procedures in a critical facility environment.
  • Experience with out-of-band server management (BMC, IPMI, Redfish, iDRAC, iLO, or equivalent).
  • Familiarity with high-speed fiber optics (LC/MPO, QSFP-DD/OSFP) including cleaning, inspection, and patch panel management.
Strongly Preferred
  • 2+ years experience with GPU compute infrastructure (AMD Instinct, NVIDIA HGX, or equivalent OAM/SXM form factors).
  • Experience with AMD EPYC server platforms (1P/2P) including BIOS, memory population, and diagnostic procedures.
  • Familiarity with OCP Open Rack / ORW standards and open datacenter architectures.
  • Exposure to AI or HPC datacenter deployments with cluster-scale interconnects (InfiniBand, Ethernet RoCE, or UALink).
  • Working knowledge of AMD ROCm software stack for hardware validation tasks.
Nice to Have
  • BICSI RCDD, DCIS, or equivalent datacenter infrastructure certification.
  • CompTIA Server+, Linux+, or similar vendor/industry certifications.
  • Experience with Schneider Electric EcoStruxure, Motivair, or Vertiv CDU/DCIM platforms.
  • Familiarity with change management frameworks (ITIL) and CMDB tooling (ServiceNow or equivalent).
PHYSICAL REQUIREMENTS
  • Ability to lift and maneuver equipment up to 50 lbs (23 kg) unassisted; team-lift required for compute trays (~77 kg) and switch trays (~120 kg with tooling).
  • Ability to stand, bend, kneel, and work in confined rack spaces for extended periods.
  • Comfort working in raised-floor and hot/cold-aisle-contained datacenter environments.
  • Manual dexterity sufficient to handle fine connectors, optical fibers, and small form-factor electronics.

Similar Jobs

See more jobs