Description:
Hybrid Onsite - local candidates preferred.
Real-Time Services Real-Time Inference Engineering Lead
Experience: 8 12 years
Role Summary
The Real-Time Inference Engineering Lead will design and industrialize low-latency, resilient model-serving services for predictive AI use cases. The role will define deployment patterns, capacity controls, monitoring, performance standards, and operational practices across cloud and on-premises environments.
Key Responsibilities
Architect low-latency online inference and real-time model-serving solutions.
Develop scalable APIs, microservices, and deployment patterns for predictive models.
Implement Kubernetes-based deployment, autoscaling, load balancing, and traffic-management strategies.
Conduct benchmarking, performance tuning, capacity planning, and load testing.
Optimize latency, throughput, resource consumption, availability, and cost.
Define monitoring, alerting, SLOs, runbooks, and incident-response practices.
Build CI/CD pipelines for repeatable model and service releases.
Design resilience, failover, rollback, disaster recovery, and graceful-degradation patterns.
Lead technical reviews and mentor inference and platform engineers.
Required Skills
Online inference and real-time model-serving architecture.
REST/gRPC APIs and distributed microservices.
Kubernetes, containers, autoscaling, and traffic management.
Performance engineering, latency optimization, and load testing.
Monitoring, SLOs, capacity planning, and production operations.
CI/CD and progressive-deployment approaches.
Resilience and high-availability engineering.
Cloud and on-premises deployment experience.
Preferred Qualifications
Degree in computer science, engineering, or a related discipline.
Experience with enterprise model-serving platforms and inference runtimes.
Cloud, Kubernetes, SRE, or ML engineering certification.