Location: Orlando, Florida (Hybrid; on-site support required as needed)
Program: AI as a Service (AIaaS)
Clearance: Ability to obtain and maintain a DoD Secret clearance; active clearance preferred
Position Overview
Join a forward-looking engineering team building and sustaining a centralized AI as a Service environment for mission-critical government applications. In this role, you will help operationalize advanced AI capabilities by designing, integrating, administering, and optimizing the secure infrastructure that powers them.
This is an infrastructure-focused positionnot an AI model-development role. You will work across GPU compute, enterprise storage, virtualization, operating systems, containers, networking, and supporting platform services. You will also collaborate closely with cybersecurity engineers, AI engineers, network engineers, and platform architects to integrate mission workflows into the AETHER AI ecosystem.
The position may be filled at the Systems Engineer III or Systems Engineer IV level based on the selected candidates experience and qualifications.
Key Responsibilities
- Integrate the AETHER AI platform into existing VMware ESXi and virtualized infrastructure environments.
- Provision, configure, administer, and maintain Linux- and Windows-based infrastructure hosts.
- Integrate AI services with enterprise storage, networking, authentication, and identity-management systems.
- Configure storage backends supporting large-scale simulation artifacts, AI data repositories, and workflow execution.
- Implement and optimize secure remote-access solutions for distributed users and administrators.
- Support infrastructure capacity planning, sustainment, modernization, and lifecycle-management activities.
- Perform system administration across compute, storage, virtualization, GPU, and AI platform resources.
- Deploy, configure, validate, and maintain NVIDIA GPU infrastructure supporting AI workloads.
- Install and manage NVIDIA drivers, CUDA, NCCL, and related GPU runtime environments.
- Configure GPU allocation and multi-tenant resource-sharing capabilities.
- Deploy and support containerized services using Docker, Kubernetes, K3s, Rancher, OpenShift, or comparable technologies.
- Monitor GPU utilization, performance, availability, and capacity.
- Troubleshoot complex interactions among GPU hardware, drivers, containers, operating systems, and platform services.
- Optimize GPU resources for concurrent AI inference and workflow execution.
- Deploy and configure AI workloads within the platform environment.
- Integrate AI models and workflows into secure operational environments.
- Support workflow execution, automation, and orchestration.
- Evaluate and validate platform performance under operational loads.
- Assist with production deployments, acceptance testing, and operational-readiness activities.
- Partner with cybersecurity personnel to implement secure configurations and Risk Management Framework requirements.
- Support system testing, verification, validation, accreditation, and technical reviews.
- Lead or contribute to issue resolution, root-cause analysis, and system troubleshooting.
- Develop and maintain system diagrams, technical documentation, configuration records, and standard operating procedures.
Required Qualifications
Systems Engineer III
- Bachelors degree in computer science, information systems, engineering, or a related technical discipline.
- Five or more years of experience in systems administration, infrastructure engineering, systems engineering, or a related field.
- Hands-on experience with VMware ESXi, VMware vCenter, or comparable virtualization platforms.
- Experience administering Windows Server and Linux operating systems.
- Experience with enterprise storage technologies such as SAN, NAS, iSCSI, and NFS.
- Experience supporting and troubleshooting complex enterprise infrastructure environments.
- Strong understanding of networking fundamentals, protocols, and services.
- Experience supporting containerized environments using technologies such as Docker, Kubernetes, K3s, OpenShift, or Rancher.
- Ability to obtain and maintain a DoD Secret security clearance.
Systems Engineer IV
- Bachelors degree in computer science, information systems, engineering, or a related technical discipline.
- Eight or more years of experience in systems administration, infrastructure engineering, systems engineering, or a related field.
- Advanced experience with virtualization, Windows and Linux administration, enterprise storage, networking, and containerized environments.
- Demonstrated experience leading infrastructure integration, modernization, or platform-deployment initiatives.
- Ability to investigate and resolve complex, cross-platform infrastructure issues.
- Ability to provide technical leadership during engineering reviews, planning activities, testing, and production deployments.
- Ability to obtain and maintain a DoD Secret security clearance.
Desired Qualifications
- Current or recently active DoD Secret security clearance.
- Experience administering NVIDIA GPU infrastructure in enterprise or multi-user environments.
- Working knowledge of CUDA, NCCL, NVIDIA drivers, GPU runtimes, and GPU monitoring tools.
- Experience supporting AI, machine-learning, high-performance-computing, or simulation workloads.
- Experience operating Kubernetes-based platforms and multi-tenant container environments.
- Familiarity with infrastructure automation, configuration management, or scripting.
- Experience supporting systems within secure government or DoD environments.
- Familiarity with the DoD Risk Management Framework, security hardening, system accreditation, and technical control implementation.
- Experience with performance benchmarking, capacity planning, and infrastructure lifecycle management.
- Strong technical documentation, communication, and cross-functional collaboration skills.