Inference Lead

Kasmo Inc

  • Charlotte, NC
  • 10 days ago

    Highlights

    The Real-Time Inference Engineering Lead will design and industrialize low-latency, resilient model-serving services for predictive AI use cases. Develop scalable APIs, microservices, and deployment patterns for predictive models.

    Numbers & Facts

    LocationCharlotte, NC

    Description



    Description:
    Hybrid Onsite - local candidates preferred.
    Real-Time Services Real-Time Inference Engineering Lead
    Experience: 8 12 years
    Role Summary
    The Real-Time Inference Engineering Lead will design and industrialize low-latency, resilient model-serving services for predictive AI use cases. The role will define deployment patterns, capacity controls, monitoring, performance standards, and operational practices across cloud and on-premises environments.
    Key Responsibilities
    Architect low-latency online inference and real-time model-serving solutions.
    Develop scalable APIs, microservices, and deployment patterns for predictive models.
    Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
    Conduct benchmarking, performance tuning, capacity planning, and load testing.
    Optimize latency, throughput, resource consumption, availability, and cost.
    Define monitoring, alerting, SLOs, runbooks, and incident-response practices.
    Build CI/CD pipelines for repeatable model and service releases.
    Design resilience, failover, rollback, disaster recovery, and graceful-degradation patterns.
    Lead technical reviews and mentor inference and platform engineers.
    Required Skills
    Online inference and real-time model-serving architecture.
    REST/gRPC APIs and distributed microservices.
    Kubernetes, containers, autoscaling, and traffic management.
    Performance engineering, latency optimization, and load testing.
    Monitoring, SLOs, capacity planning, and production operations.
    CI/CD and progressive-deployment approaches.
    Resilience and high-availability engineering.
    Cloud and on-premises deployment experience.
    Preferred Qualifications
    Degree in computer science, engineering, or a related discipline.
    Experience with enterprise model-serving platforms and inference runtimes.
    Cloud, Kubernetes, SRE, or ML engineering certification.

    Similar Jobs