NOC Engineer / NOC Lead

STN

  • San Francisco Bay Area, California
  • 8 days ago
  • Remote

    Highlights

    The role triages alerts, executes documented runbooks, and coordinates with on-call specialists during incidents to protect customer SLAs. The NOC Engineer operates STN's 24/7 monitoring and first-response capability for GPU One (GPUaaS) infrastructure.

    Numbers & Facts

    LocationSan Francisco Bay Area, California (
    Remote
    )

    Description

    NOC Engineer / NOC Lead

    Infrastructure operations · shared across customers

    Reports to: Manager, NOC (or Director, Service Operations)

    Location: Remote (US) with assigned shift; rotating coverage

    Department: Infrastructure & DC Operations / Network Engineering

    Position summary

    The NOC Engineer operates STN's 24/7 monitoring and first-response capability for GPU One (GPUaaS) infrastructure. The role triages alerts, executes documented runbooks, and coordinates with on-call specialists during incidents to protect customer SLAs.

    Key responsibilities

    • Monitor infrastructure alerts, customer SLA dashboards, and system health on a 24/7 basis

    • Triage incidents and engage on-call SREs, Network, Hardware, or Field Engineering as needed

    • Execute documented runbooks for common platform, network, and hardware issues

    • Manage the incident lifecycle including initial customer notification and status updates

    • Coordinate planned maintenance windows and change windows with internal teams and customers

    • Update status pages and customer-facing communications during incidents

    • Maintain shift handoff documentation and active-incident logs

    • Support ticket queue handling including Tier 1 ticket resolution

    • Contribute to continuous improvement of monitoring coverage, alert quality, and runbooks

    • Work rotating shifts including nights, weekends, and holidays

    Required qualifications

    • 3+ years in a NOC, SOC, or IT operations function

    • Hands-on experience with monitoring tools (Datadog, Prometheus, Grafana, PagerDuty, or equivalent)

    • Strong Linux and basic networking fundamentals

    • Excellent written and verbal communication, particularly under pressure

    • Willingness and ability to work rotating shifts including overnight coverage

    Preferred qualifications

    • GPU, HPC, or large-scale cloud infrastructure background

    • ITIL Foundations certification

    • Demonstrated on-call and major-incident response experience

    • Scripting skills (Python, Bash) for runbook automation

    Similar Jobs