Job Title: Principal Machine Learning Engineer
Position Description: Protingent Staffing has an exciting direct hire Principal Machine Learning Engineer
with our client that is fully remote.
Job Description:- As a Principal Machine Learning Engineer, you are a deep technical authority responsible for designing and evolving the most critical ML systems in the company.
- You operate across training, inference, evaluation, and infrastructure, solving the hardest architectural and performance problems.
- While Technical Leads may own execution at the team level, you set the technical standard and shape how ML systems are built across the organization.
- This is a hands-on, high-impact role focused on depth.
Job Responsibilities:- Architect and build large-scale ML systems spanning data, training, evaluation, inference, and deployment.
- Design reproducible, high-performance training pipelines across GPU infrastructure.
- Architect inference systems that balance latency, throughput, cost, and reliability at scale.
- Design and maintain data systems for high-quality synthetic and real-world training data.
- Implement evaluation pipelines covering performance, robustness, safety, and bias, in partnership with research leadership.
- Own production deployment, including GPU optimization, memory efficiency, latency reduction, and scaling policies.
- Collaborate closely with application engineering to integrate ML systems cleanly into backend, mobile, and desktop products
- Make pragmatic trade-offs and ship improvements quickly, learning from real usage.
- Work under real production constraints: latency, cost, reliability, and safety
- ML systems (training, inference, evaluation) are reliable, scalable, and meet defined performance targets.
- Models deployed to production achieve measurable quality improvements and meet user-impact goals.
- Production issues are proactively monitored, debugged, and resolved with clear root-cause analysis.
- Team and cross-functional collaborators benefit from clear guidance, best practices, and scalable ML solutions.
- Research-to-production cycles are efficient, safe, and continuously improve the product experience.
Job Qualifications:- Strong background in deep learning and transformer-based architectures.
- Hands-on experience training, fine-tuning, or deploying large-scale ML models in production.
- Proficiency with at least one modern ML framework (e.g. PyTorch, JAX), and ability to learn others quickly.
- Experience with distributed training and inference frameworks (e.g. DeepSpeed, FSDP, Megatron, ZeRO, Ray).
- Strong software engineering fundamentals – you write robust, maintainable, production-grade systems.
- Experience with GPU optimization, including memory efficiency, quantization, and mixed precision.
- Comfort owning ambiguous, zero-to-one ML systems end-to-end.
- A bias toward shipping, learning fast, and improving systems through iteration.
Must Have:- Experience with LLM inference frameworks such as vLLM, TensorRT-LLM, or FasterTransformer.
- Contributions to open-source ML or systems libraries.
- Background in scientific computing, compilers, or GPU kernels.
- Experience with RLHF pipelines (PPO, DPO, ORPO).
- Experience training or deploying multimodal or diffusion models.
- Experience with large-scale data processing (Apache Arrow, Spark, Ray).
Job Details:- Job Type: Direct Hire
- Pay Range: Market Rate
- Location: Fully Remote.
About Protingent: Protingent is an
Award-Winning provider of top-tier Engineering and IT talent, trusted by companies at the forefront of innovation — from
Software and Aerospace to
AI, Clean Tech, Medical Devices, and Connected Technologies. We’re passionate about making a positive impact by connecting exceptional talent with meaningful opportunities and helping our clients build the future.