GPU Infrastructure NOC Engineer

Orion Placement

  • Pittsburgh, Pennsylvania
  • 12 days ago
  • Remote

    Highlights

    Note: Must have 2+ years of NOC, network operations, or infrastructure monitoring experience, hands-on GPU or HPC infrastructure experience, and Python or Bash scripting experience. Work alongside technical teams, OEMs, data center operators, and network providers to solve real infrastructure problems.

    Numbers & Facts

    LocationPittsburgh, Pennsylvania (
    Remote
    )

    Description

    Pay: $75,000.00 - $140,000.00 per year

    Why This Is a Great Opportunity

    • Monitor and support cutting-edge GPU and AI infrastructure powering demanding enterprise workloads.
    • Go beyond traditional NOC monitoring by building automation, improving tooling, and making operations smarter.
    • Get hands-on exposure to GPU, HPC, networking, data center infrastructure, and AI-enabled operations.
    • Play a direct role in protecting customer uptime and meeting demanding SLA commitments.
    • Help create and continuously improve runbooks, SOPs, monitoring processes, and incident-response workflows.
    • Work alongside technical teams, OEMs, data center operators, and network providers to solve real infrastructure problems.
    • Join a fast-growing environment where your ideas can directly improve how the NOC operates.
    • Receive bonus and equity opportunities in addition to competitive base compensation.

    Location: Remote nationwide. This is a fully remote role supporting a 24/7 global operations environment through rotating shifts, including nights and weekends.

    Note: Must have 2+ years of NOC, network operations, or infrastructure monitoring experience, hands-on GPU or HPC infrastructure experience, and Python or Bash scripting experience. Candidates must be willing to work rotating 24/7 shifts, including nights and weekends. These are hard requirements.

    About Us

    We are building next-generation AI infrastructure designed to give enterprise customers fast, flexible access to high-performance GPU compute. Our operations team plays a critical role in keeping these environments reliable, responsive, and continuously improving. Confidential Employer.

    Job Description

    • Monitor live GPU cluster health, power, cooling, networking, and infrastructure status across production deployments.
    • Triage, troubleshoot, and resolve incidents while maintaining SLA requirements.
    • Identify infrastructure issues early and take proactive action before they become customer-impacting incidents.
    • Escalate appropriate issues to OEMs, data center operators, network providers, or other Tier 3 partners.
    • Build and improve internal monitoring tools, scripts, and automation to reduce repetitive manual work.
    • Use Python, Bash, or similar scripting tools to automate monitoring, triage, reporting, and operational workflows.
    • Explore and implement AI-enabled workflows that improve NOC speed, accuracy, and efficiency.
    • Create, maintain, and continuously improve technical runbooks and standard operating procedures.
    • Participate in incident reviews and turn recurring problems into permanent tooling, process, or automation improvements.
    • Track SLA and incident metrics and identify opportunities to improve reliability and response times.
    • Communicate clearly and proactively with customers and internal stakeholders during incidents.
    • Coordinate with data center operators, OEMs, network providers, and other third parties to resolve customer-impacting issues.
    • Support a 24/7 rotating operations schedule, including nights, weekends, and other assigned shifts.

    Qualifications

    • 2+ years of experience in a NOC, network operations, infrastructure monitoring, or related technical operations environment.
    • Direct experience supporting or monitoring GPU, HPC, AI infrastructure, or comparable high-performance computing environments.
    • Hands-on Python or Bash scripting experience.
    • Experience with infrastructure monitoring and alerting tools such as Datadog, Grafana, PagerDuty, or similar platforms.
    • Strong troubleshooting and incident-response skills.
    • Experience working with network, compute, storage, or data center infrastructure.
    • Ability to understand technical issues quickly and communicate effectively during incidents.
    • Demonstrated interest in automation, scripting, tool-building, and continuous operational improvement.
    • Must be comfortable working rotating 24/7 shifts, including nights and weekends.
    • Ability to work independently in a remote environment while collaborating effectively with global technical teams.
    • Additional languages beyond English are a plus.

    Why You Will Love Working Here

    • Work directly with advanced GPU and HPC infrastructure rather than generic IT systems.
    • Build technical depth across AI compute, networking, data centers, monitoring, and automation.
    • Have a voice in how the NOC operates and help improve processes rather than simply following them.
    • Work with modern monitoring, automation, and AI-enabled operations tools.
    • Gain exposure to complex enterprise infrastructure and high-availability environments.
    • Remote nationwide flexibility with multiple shift options.
    • Medical, dental, and vision insurance.
    • 401(k).
    • Paid maternity and paternity leave.
    • Bonus and equity opportunities.

    JPC-1916

    Benefits:

    • Dental insurance
    • Paid time off
    • Retirement plan
    • Vision insurance

    Similar Jobs