Infrastructure Architect

Kasmo Inc

  • Bay Area, Charlotte, Dallas, Phoenix, NY, NC
  • 11 days ago

    Highlights

    The successful candidate will demonstrate deep expertise in cloud-native architectures, AI/ML platforms, distributed systems, and platform engineering, along with the ability to influence technical roadmaps and lead cross-functional initiatives. This role will provide architectural leadership for large-scale AI and cloud platform initiatives, partnering with engineering, infrastructure, security, product, and business stakeholders to deliver scalable, resilient, and secure AI capabilities.

    Numbers & Facts

    LocationBay Area, Charlotte, Dallas, Phoenix, NY, NC

    Description



    Description:
    Local candidates preferred.
    Job Summary
    We are seeking an experienced AI Platform Architect to lead the technical architecture, strategy, and evolution of enterprise AI platforms. This role will provide architectural leadership for large-scale AI and cloud platform initiatives, partnering with engineering, infrastructure, security, product, and business stakeholders to deliver scalable, resilient, and secure AI capabilities.
    The successful candidate will demonstrate deep expertise in cloud-native architectures, AI/ML platforms, distributed systems, and platform engineering, along with the ability to influence technical roadmaps and lead cross-functional initiatives.
    Key Responsibilities
    Architecture Leadership
    Lead architecture and technical direction for enterprise AI platform capabilities, including:
    Enterprise Generative AI platforms
    Agentic AI platforms and agent runtime environments
    Model serving and inference infrastructure
    Prompt engineering, evaluation, and testing frameworks
    AI governance, risk management, and guardrails
    AI observability, monitoring, and operations
    Multi-cloud AI platform strategy and architecture
    Platform Architecture & Design
    Design, evaluate, and guide architecture across AI and cloud technologies such as:
    Red Hat OpenShift AI (RHOAI)
    Google Cloud Vertex AI
    Gemini models
    Azure AI Foundry
    Amazon Bedrock
    Anthropic Claude
    OpenAI services and models
    Cloud-native platform scalability and resiliency solutions
    Strategic Initiatives
    Lead and collaborate on initiatives involving:
    Capacity planning and performance optimization
    GPU infrastructure and platform strategy
    Large-scale NVIDIA-based AI infrastructure architectures
    Multi-region and multi-cloud resiliency
    Disaster recovery planning and cloud DR strategies
    Active-active platform architectures
    High availability and business continuity solutions
    Required Qualifications
    Bachelor's degree in Computer Science, Information Systems, Engineering, or a related field; or equivalent combination of education and relevant experience.
    Minimum of 10 years of experience in software engineering, infrastructure engineering, systems architecture, or related technical disciplines.
    Minimum of 5 years of experience designing and implementing large-scale cloud-native platforms.
    Experience designing, deploying, and supporting highly available, mission-critical production systems.
    Experience with:
    Kubernetes and/or OpenShift
    Distributed systems architectures
    Public cloud platforms and cloud-native architectures
    AI/ML platforms and infrastructure
    API and integration platforms
    Data platforms and data-intensive applications
    Knowledge of:
    Generative AI systems and architectures
    Retrieval-Augmented Generation (RAG)
    Agentic AI frameworks and platforms
    Model serving and inference architectures
    MLOps practices and tooling
    Demonstrated ability to lead technical initiatives across multiple teams and stakeholders.
    Ability to work from or relocate to one of the following approved locations: San Francisco Bay Area, CA; Charlotte, NC; Dallas, TX; Phoenix, AZ; or New York, NY.
    Preferred Qualifications
    Experience with one or more major cloud providers, including AWS, Azure, or Google Cloud Platform.
    Experience architecting GPU-accelerated AI infrastructure.
    Experience implementing AI governance, security, risk management, and compliance controls.
    Experience building or supporting multi-region, highly resilient enterprise platforms.
    Relevant industry certifications in cloud, AI/ML, Kubernetes, or architecture disciplines.

    Similar Jobs