AI Data Center Architect

Intellectt INC

  • Plano, TX
  • 2 days ago

    Highlights

    The ideal candidate will have deep expertise in GPU infrastructure, AI Factories, GPU-as-a-Service (GPUaaS), NVIDIA reference architectures, high-performance networking, AI storage, Kubernetes, and large-scale AI/LLM workloads. The successful candidate will be responsible for designing end-to-end AI infrastructure solutions, including GPU compute, high-performance networking, storage, orchestration, multi-tenancy, and AI platform services.

    Numbers & Facts

    LocationPlano, TX

    Description

    Job Title: AI Data Center Architect
    Location: Plano, TX

    Job Summary

    We are seeking a highly skilled AI Data Center Architect with hands-on experience designing and architecting modern AI infrastructure at enterprise or hyperscale scale.

    This is not a traditional enterprise data center architecture role. The ideal candidate will have deep expertise in GPU infrastructure, AI Factories, GPU-as-a-Service (GPUaaS), NVIDIA reference architectures, high-performance networking, AI storage, Kubernetes, and large-scale AI/LLM workloads.

    The successful candidate will be responsible for designing end-to-end AI infrastructure solutions, including GPU compute, high-performance networking, storage, orchestration, multi-tenancy, and AI platform services. The individual should be comfortable working directly with customers and technical stakeholders to translate AI workload requirements into scalable infrastructure architectures.

    Key Responsibilities
    • Design and architect enterprise and hyperscale AI Factory environments from the ground up.
    • Develop scalable GPU infrastructure and GPU-as-a-Service (GPUaaS) architectures.
    • Design infrastructure for large-scale AI training, inference, LLM, and GenAI workloads.
    • Develop solutions based on NVIDIA Enterprise Reference Architectures.
    • Architect NVIDIA GPU platforms including DGX, HGX, H200, GB200, and Blackwell systems.
    • Design high-performance AI networking using InfiniBand, RoCE, Spectrum-X, NVLink, and BlueField DPUs.
    • Perform GPU cluster sizing, capacity planning, performance optimization, and scalability analysis.
    • Design high-performance storage architectures capable of supporting large-scale GPU and AI workloads.
    • Architect Kubernetes-based AI platforms, including GPU scheduling, workload orchestration, resource management, and multi-tenancy.
    • Design distributed training and inference infrastructure for modern LLM and GenAI workloads.
    • Develop architectures for AI Cloud, GPU Cloud, and on-premises AI Factory deployments.
    • Evaluate infrastructure requirements across compute, networking, storage, orchestration, and AI platform layers.
    • Work with engineering, infrastructure, cloud, networking, and AI/ML teams to develop end-to-end solutions.
    • Participate in customer-facing architecture discussions, technical workshops, solution design, and presentations.
    • Create high-level and low-level architecture designs, technical documentation, reference architectures, and solution proposals.
    • Evaluate emerging NVIDIA technologies and AI infrastructure trends and incorporate them into future architectures.

    Required Qualifications
    • 4+ years of experience designing AI-focused infrastructure, accelerated computing platforms, or GPU-based data center environments.
    • Proven experience architecting AI Factories, GPU clusters, or GPUaaS platforms.
    • Experience designing environments supporting 100+ GPUs is strongly preferred.
    • Deep understanding of NVIDIA AI infrastructure and Enterprise Reference Architectures.
    Hands-on architecture experience with one or more of:
    • NVIDIA DGX
    • NVIDIA HGX
    • H200
    • GB200
    • NVIDIA Blackwell platforms
    Strong understanding of:
    • InfiniBand
    • Spectrum-X
    • RoCE
    • NVLink / NVSwitch
    • BlueField DPUs
    • 400G/800G networking
    • Experience with GPU cluster sizing, performance optimization, scaling, and capacity planning.
    • Strong understanding of LLM training and inference infrastructure.
    • Experience with Kubernetes for AI/GPU workloads.
    • Knowledge of GPU scheduling, workload orchestration, and multi-tenant GPU environments.
    • Experience designing distributed AI/ML training infrastructure.
    • Strong knowledge of high-performance storage architectures for AI workloads.
    • Understanding of MLOps, AI platforms, and GenAI infrastructure.
    • Ability to design complete AI infrastructure solutions spanning compute, networking, storage, orchestration, and platform services.

    Preferred Qualifications
    • Experience designing on-premises AI Factories, Sovereign AI infrastructure, or GPU Cloud platforms.
    • Experience with Azure, AWS, or Google Cloud AI infrastructure.
    • Experience with NVIDIA AI Enterprise and the broader NVIDIA AI software ecosystem.
    • Experience with large-scale distributed training frameworks and AI workload orchestration.
    • Experience with AI infrastructure benchmarking and performance optimization.
    • Customer-facing architecture, consulting, solution engineering, or technical pre-sales experience.
    • Experience developing technical proposals, architecture diagrams, bills of materials, and solution designs.

    Similar Jobs

    See more jobs