| Location | North Carolina, NC |
Contractor 2 - AI Cluster Validation Engineer (Tester & Programmer) Key Responsibilities AI Cluster Infrastructure Validation: Build, deploy, and maintain hardware and software test environments for AI cluster validation, qualification, and performance benchmarking. Execute test plans, reproduce complex issues, collect and analyze logs, and perform first-level root cause analysis across hardware, software, networking, and system components. Develop and enhance automated test frameworks, scripts, and tools to improve validation efficiency, coverage, and repeatability. Collaborate with software, hardware, networking, and system engineering teams to investigate issues, validate fixes, and improve overall cluster stability, scalability, and performance. Conduct functional, performance, stress, and reliability testing for AI/HPC cluster solutions. Document test methodologies, configurations, results, troubleshooting procedures, and operational best practices. Basic Qualifications Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field. 1+ years of experience in deploying, testing, validating, or supporting data center hardware and software systems, with expertise in one or more of the following areas: Server systems, Networking, Storage Strong understanding of Linux operating systems, system administration, and troubleshooting. Familiarity with data center cluster management platforms, distributed computing environments, and infrastructure validation methodologies. Programming or scripting experience with Python, Bash, or similar languages. Strong analytical, debugging, problem-solving, and troubleshooting skills. Proven ability to quickly learn new technologies and adapt in a fast-paced engineering environment. Self-motivated team player with strong communication and collaboration skills. Preferred Qualifications Experience operating, maintaining, or supporting data center, cloud, or laboratory environments. Familiarity with AI/HPC clusters, GPU-based systems, and high-speed interconnect technologies such as InfiniBand, RoCE, NVLink, and Ethernet fabrics. Experience developing test automation tools, validation frameworks, or CI/CD pipelines. Knowledge of cluster orchestration and management technologies, including Kubernetes, Slurm, virtualization platforms, or cloud infrastructure. Experience with performance analysis, benchmarking, workload characterization, and system optimization. Understanding of storage technologies, distributed file systems, and AI workload deployment environments.