Remote | AWS Trainium Kernel Engineer (NKI) — $60–$80/hour

24-Mag

  • New York, New York
  • 7 days ago
  • Remote

    Highlights

    We are sharing a specialised part-time consulting opportunity for experienced kernel engineers with hands-on expertise in the Neuron Kernel Interface (NKI), AWS Trainium/Inferentia2 hardware, low-level performance optimisation, and CUDA-to-NKI migration. Selected experts will review Trainium-specific implementations, migration decisions, memory-management strategies, profiling results, and cross-platform numerical behaviour while providing clear, rubric-based technical feedback.

    Numbers & Facts

    LocationNew York, New York (
    Remote
    )
    Website4-mag.com/privacy-policy

    Description

    We are sharing a specialised part-time consulting opportunity for experienced kernel engineers with hands-on expertise in the Neuron Kernel Interface (NKI), AWS Trainium/Inferentia2 hardware, low-level performance optimisation, and CUDA-to-NKI migration.

    This role focuses on evaluating NKI kernel-development tasks for technical correctness, hardware appropriateness, numerical fidelity, and performance quality. Selected experts will review Trainium-specific implementations, migration decisions, memory-management strategies, profiling results, and cross-platform numerical behaviour while providing clear, rubric-based technical feedback.

    Key Responsibilities

    NKI Kernel Development Review

    • Evaluate kernels developed using the Neuron Kernel Interface (NKI)
    • Assess whether implementations appropriately target AWS Trainium and Inferentia2 hardware
    • Review low-level computation patterns for technical correctness
    • Identify inefficient, incorrect, or hardware-inappropriate implementation choices
    • Apply practical judgement grounded in hands-on NKI development experience

    CUDA-to-NKI Migration

    • Review migrations of existing CUDA kernels to NKI
    • Assess whether computational semantics are preserved across platforms
    • Identify translation errors, unsupported assumptions, or inefficient migration strategies
    • Evaluate whether NKI implementations appropriately account for Trainium architecture
    • Distinguish faithful migrations from implementations that merely reproduce surface-level CUDA structure

    Tile-Based Computation

    • Assess tile decomposition and computation strategies
    • Review partitioning decisions against NKI execution constraints
    • Evaluate whether kernels make effective use of available compute resources
    • Identify inefficient tiling or data-movement patterns
    • Assess whether implementation choices align with NKI programming requirements

    Memory Hierarchy Management

    • Review use of SBUF, PSUM, and HBM
    • Assess data placement and movement across Trainium memory hierarchies
    • Evaluate memory-bandwidth utilisation and locality
    • Identify unnecessary transfers or memory bottlenecks
    • Review implementation decisions affecting on-chip memory efficiency

    DMA & Data Movement

    • Evaluate DMA orchestration within NKI kernels
    • Review sequencing of computation and data-transfer operations
    • Identify stalls, inefficient transfer patterns, or synchronisation issues
    • Assess whether data movement appropriately overlaps with computation
    • Evaluate implementation choices affecting pipeline utilisation

    Trainium Performance Optimisation

    • Review Trainium-specific optimisation strategies
    • Assess NeuronCore pipeline utilisation, tensor-engine throughput, and memory-bandwidth behaviour
    • Identify performance bottlenecks within kernel implementations
    • Evaluate whether optimisation decisions are supported by profiling evidence
    • Review trade-offs affecting throughput, latency, and resource utilisation

    Numerical Correctness

    • Evaluate numerical consistency between GPU and Trainium implementations
    • Review differences caused by accumulation order, rounding behaviour, and mixed-precision semantics
    • Assess appropriate tolerances for cross-platform comparisons
    • Identify numerical discrepancies that indicate implementation defects
    • Distinguish expected hardware-level variation from substantive correctness problems

    Precision & Data Types

    • Review kernels using supported formats such as FP32, BF16, FP8, and INT8
    • Assess precision choices against computational requirements
    • Evaluate mixed-precision behaviour and numerical stability
    • Identify inappropriate casting or accumulation strategies
    • Review whether performance gains are achieved without compromising required correctness

    AWS Neuron Ecosystem

    • Evaluate implementations using the AWS Neuron SDK
    • Review interactions between kernel code, compilation, and Trainium execution
    • Assess compiler-related behaviours where relevant
    • Apply familiarity with NKI kernel libraries and Neuron tooling
    • Identify implementation issues arising from platform-specific constraints

    Benchmarking & Validation

    • Review benchmark results for Trainium workloads
    • Assess performance comparisons and experimental methodology
    • Evaluate workloads running on Trn1 or Trn2 instances where applicable
    • Determine whether claimed performance improvements are supported by evidence
    • Identify benchmarking methodologies that could produce misleading conclusions

    Rubric-Based Technical Evaluation

    • Assess assigned kernel-development tasks against structured technical criteria
    • Provide clear written explanations supporting evaluation decisions
    • Reference specific implementation, profiling, or numerical evidence
    • Apply evaluation standards consistently across assignments
    • Distinguish valid optimisation alternatives from technically flawed approaches

    Ideal Profile

    • 2+ years of hands-on experience developing or optimising kernels using the Neuron Kernel Interface (NKI)
    • Professional experience targeting AWS Trainium or Inferentia2 hardware
    • Strong understanding of tile-based computation
    • Deep familiarity with SBUF, PSUM, and HBM memory management
    • Strong knowledge of partition-dimension constraints and DMA orchestration
    • Demonstrated experience evaluating or performing CUDA-to-NKI migrations
    • Familiarity with Trainium-specific performance profiling
    • Experience assessing NeuronCore pipeline utilisation, tensor-engine throughput, and memory-bandwidth bottlenecks
    • Strong understanding of cross-platform numerical correctness and mixed-precision behaviour
    • Direct experience with the AWS Neuron SDK, Neuron Compiler internals, or NKI kernel libraries is preferred
    • Prior CUDA or Triton kernel development experience is advantageous
    • Familiarity with NeuronCore-v2 architecture and supported numerical formats is preferred
    • Experience benchmarking ML workloads on Trn1 or Trn2 instances is advantageous
    • Strong written communication and ability to provide precise technical feedback

    Engagement Details

    • Part-time independent contractor engagement
    • Fully remote within the United States
    • Flexible scheduling based on project requirements
    • Compensation: $60–$80/hour
    • Work focuses on NKI kernel development, Trainium performance optimisation, CUDA migration, numerical correctness, and technical quality evaluation
    • Projects may be extended, shortened, or concluded based on project needs and performance
    • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
    • H1-B and STEM OPT support is unavailable for this engagement

    About the Platform

    This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

    By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.

    Similar Jobs

    See more jobs