Remote | ML Infrastructure & Kernel Optimization Engineer — $65–$105/hour

24-Mag

  • New York, New York
  • 12 days ago
  • Remote

    Highlights

    We are sharing a specialised full-time consulting opportunity for US-based MLOps and ML systems engineers with production experience in JAX, PyTorch, distributed training infrastructure, and custom GPU kernel development using Pallas or Triton. This role supports a high-impact generative AI initiative focused on developing and evaluating advanced ML infrastructure tasks for frontier model training.

    Numbers & Facts

    LocationNew York, New York (
    Remote
    )
    Website4-mag.com/privacy-policy

    Description

    We are sharing a specialised full-time consulting opportunity for US-based MLOps and ML systems engineers with production experience in JAX, PyTorch, distributed training infrastructure, and custom GPU kernel development using Pallas or Triton.

    This role supports a high-impact generative AI initiative focused on developing and evaluating advanced ML infrastructure tasks for frontier model training. Selected engineers will design technically challenging problems, produce rigorous solutions, assess model-generated outputs, and help establish evaluation standards across training pipelines, distributed systems, framework-level optimisation, and GPU kernel performance.

    Key Responsibilities

    ML Infrastructure & Training Systems

    • Analyse and improve machine learning training infrastructure, deployment workflows, and model-development systems
    • Guide research and engineering teams on MLOps, distributed training, and ML framework-level challenges
    • Evaluate training-pipeline architecture, scalability, reliability, and performance
    • Identify technical gaps affecting model training, experimentation, and infrastructure efficiency

    Technical Task & Solution Development

    • Design challenging tasks covering MLOps, ML systems, training infrastructure, and framework-level engineering
    • Write accurate, technically rigorous, and well-structured solutions
    • Develop realistic scenarios involving distributed systems, accelerator utilisation, and production ML workflows
    • Ensure tasks reflect practical engineering challenges encountered in advanced AI environments

    JAX, PyTorch & GPU Kernel Evaluation

    • Evaluate technical work involving JAX and PyTorch at production scale
    • Review custom GPU kernels written or optimised using Pallas or Triton
    • Assess kernel correctness, memory access patterns, computational efficiency, and hardware utilisation
    • Analyse framework-level implementation choices and identify opportunities for performance improvement

    Evaluation Frameworks & Technical Feedback

    • Compare alternative technical solutions and determine which approach is more accurate and effective
    • Provide clear written feedback on correctness, system design, scalability, and optimisation quality
    • Develop detailed rubrics for evaluating training pipelines, distributed systems reasoning, and kernel-level implementations
    • Collaborate with other technical specialists to maintain consistency across evaluation standards and training data

    Ideal Profile

    Strong candidates may have:

    • At least 2 years of dedicated professional experience in MLOps, ML infrastructure, or ML systems engineering
    • Production experience with JAX, PyTorch, or both at meaningful scale
    • Hands-on experience writing or optimising custom GPU kernels using Pallas or Triton
    • Strong knowledge of model-training pipelines, distributed systems, accelerators, and performance optimisation
    • Experience working within a recognised technology, AI research, or high-performance engineering organisation
    • Demonstrable professional growth and increasing technical responsibility
    • Strong written communication and the ability to explain complex engineering decisions clearly
    • Reliable availability for a full-time, 40-hour weekday schedule

    Educational Background

    • A degree in computer science, machine learning, electrical engineering, applied mathematics, or a related technical field is highly relevant
    • Graduate-level education in machine learning systems, distributed computing, or high-performance computing may be helpful
    • Equivalent professional experience in production ML infrastructure may also be considered
    • Advanced technical work involving GPU programming, compiler systems, or large-scale model training is especially valuable

    Nice to Have

    • Experience supporting large language model or generative AI training environments
    • Familiarity with distributed training frameworks, accelerator orchestration, and multi-host systems
    • Knowledge of XLA, CUDA, compiler optimisation, or low-level performance engineering
    • Experience benchmarking GPU workloads and diagnosing training-performance bottlenecks
    • Familiarity with model-evaluation pipelines, technical annotation, or structured training-data development
    • Previous involvement in technical review, engineering mentorship, or rubric development
    • Experience collaborating with research scientists and infrastructure engineering teams

    Why This Opportunity

    • Contribute to advanced generative AI and large-scale model-training initiatives
    • Apply deep expertise in JAX, PyTorch, Pallas, Triton, and ML infrastructure
    • Work on challenging problems spanning training systems, distributed computing, and GPU optimisation
    • Influence the quality of technical training data used in frontier AI development
    • Join a full-time remote engagement with competitive hourly compensation

    Contract Details

    • Full-time W-2 contingent employment arrangement
    • Fully remote role available to candidates based in the United States
    • Expected commitment of 40 hours per week during weekdays
    • This engagement requires full professional availability without conflicting employment or external commitments
    • Competitive rates between $65–$105 per hour depending on expertise and project scope
    • Immediate availability is preferred
    • Work may include onboarding, technical calibration, and ongoing quality-review activities
    • Project scope and duration may be adjusted according to programme requirements and performance

    About the Platform

    This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

    By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.

    Similar Jobs