| Location | Charlotte, NC |
| Job Type | Contractor |
| Salary | $60–$65 Per Hour |
| Company Size | 51 - 200 |
| Year Founded | 1996 |
| Headquarters | Edison, NJ, US |
| Website | https://www.neotechusa.com |
Role : Cortex Real-Time Inference Engineering Lead
Location : Charlotte NC (Onsite)
Contract
Role Overview
We are seeking a highly skilled Cortex Real-Time Inference Engineering Lead to design, build, and industrialize our next-generation, low-latency model serving platform. In this role, you will bridge the gap between data science and production engineering. You will be responsible for creating resilient, autoscaling deployment patterns and robust operational frameworks that support real-time predictive use cases at scale.
The ideal candidate has deep expertise in online inference architectures, container orchestration, and performance optimization, ensuring our predictive models deliver high availability and sub-second latency.
Key Responsibilities
Inference Platform Engineering: Design, build, and industrialize low-latency, highly available model-serving services and architecture for real-time predictive workflows.
Deployment Patterns: Establish standardized deployment patterns (e.g., canary, blue/green, shadow deployments) to safely deploy and update machine learning models in production without downtime.
Capacity & Autoscaling Control: Define capacity controls and implement advanced autoscaling strategies to handle highly fluctuating traffic patterns efficiently while minimizing cloud spend.
Performance & Latency Optimization: Conduct continuous performance profiling, load testing, and optimization to meet strict service-level agreements (SLAs) for model execution and API response times.
Monitoring & SLOs: Design and implement comprehensive monitoring, alerting, and logging systems to track model drift, data quality, system health, and Service Level Objectives (SLOs).
CI/CD Automation: Build robust CI/CD pipelines to automate the testing, validation, packaging, and deployment of models and inference code.
Resilience & Fault Tolerance: Build self-healing systems and implement fallback mechanisms to ensure high operational resilience and disaster recovery across both cloud and on-premises environments.
Required Skills & Qualifications
Core MLOps & Architecture
Online Inference Architecture: Deep understanding of real-time model architectures, feature stores, and the lifecycle of online model evaluation.
Model Serving Frameworks: Hands-on experience with production model servers such as Triton Inference Server, TorchServe, TF Serving, vLLM, Seldon Core, or KServe.
API Development: High proficiency in designing and consuming high-performance APIs using gRPC, REST, or GraphQL.
Infrastructure & Operations
Kubernetes Mastery: Strong experience deploying, managing, and scaling containerized workloads on Kubernetes (including microservices architecture).
Autoscaling & Orchestration: Expertise in configuring horizontal pod autoscaling (HPA), cluster autoscaling, and custom metrics-driven scaling.
Hybrid Operations: Proven track record managing deployments across both cloud providers (AWS, GCP, or Azure) and on-premises infrastructure.
Performance & Reliability
Load Testing & Profiling: Experience using benchmarking tools like Locust, JMeter, or K6 to run performance, stress, and load testing.
Latency Optimization: Knowledge of techniques to reduce inference latency, including model quantization, pruning, hardware acceleration (GPUs/TPUs), and optimized serialization.
Observability: Experience building dashboards and alerts using Prometheus, Grafana, ELK stack, or Datadog to enforce tight SLOs/SLIs.
CI/CD & Automation: Experience with GitOps and automation tools such as GitHub Actions, GitLab CI, ArgoCD, or Jenkins.
Preferred Qualifications
Bachelor’s or Master’s degree in Computer Science, Software Engineering, or a related technical field.
5+ years of experience in Software Engineering or DevOps, with at least 3 years dedicated to MLOps and production ML pipelines.
Strong programming skills in Python, Go, C++, or Java.
Familiarity with machine learning frameworks like PyTorch, TensorFlow, or Scikit-Learn.