Responsibilities include eliminating hardware bottlenecks through CUDA kernel tuning and GPU parallel computing, ensuring deep learning models and CV algorithms seamlessly processing massive, high-bandwidth streaming data at production scale. High-Performance Computing Pipeline Architecture Design, implement, and optimize high-throughput, low-latency image processing pipelines for real-time optical inspection and machine vision systems.
Numbers & Facts
Location
Phoenix, AZ
Salary
$130,000–$150,000 Per Year
Description
Sr. Machine Learning Engineer Salary Range: $130k to $150k
Our client is seeking a Sr. Machine Learning Engineering for a direct hire role to sit in North Phoenix, AZ or Hillsboro, OR. This role will be onsite 4 days a week and 1 remote day.
JOB SUMMARY
The role of Senior Machine Learning Engineer will architect and optimize real-time, high-throughput, and ultra-low latency image pipelines for next-generation Mask Inspection Tools. Responsibilities include eliminating hardware bottlenecks through CUDA kernel tuning and GPU parallel computing, ensuring deep learning models and CV algorithms seamlessly processing massive, high-bandwidth streaming data at production scale.
ESSENTIAL DUTIES AND RESPONSIBILITIES
High-Performance Computing Pipeline Architecture
Design, implement, and optimize high-throughput, low-latency image processing pipelines for real-time optical inspection and machine vision systems.
Develop scalable architectures capable of processing large volumes of imaging data while meeting stringent latency and reliability requirements.
Profile and optimize system performance across CPU, GPU, memory, and I/O subsystems.
GPU Acceleration
Design, develop, and optimize CUDA kernels to accelerate deep learning inference and classical computer vision algorithms.
Maximize GPU utilization through efficient memory management, kernel optimization, and parallel programming techniques.
Evaluate and implement performance improvements using NVIDIA GPU technologies and profiling tools.
Model Deployment & Optimization
Optimize, quantize, and deploy machine learning models using TensorRT, ONNX Runtime, or similar inference frameworks.
Integrate AI models into production-grade C++ and Python applications.
Improve inference throughput, latency, and resource utilization while maintaining model accuracy.
Develop automated deployment and validation pipelines for machine learning models.
Concurrency & Systems Optimization
Architect and implement multi-threaded, high-concurrency software components for data acquisition, buffering, streaming, and real-time processing.
Design robust synchronization and communication mechanisms between hardware interfaces and AI processing pipelines.
Optimize end-to-end system performance for deterministic, real-time execution.
Cross-Functional Collaboration
Partner with machine learning scientists, computer vision engineers, hardware engineers, and software developers to deliver integrated AI solutions.