Pay rate range - $65/hr. to $68/hr.
100% Remote
The Operating Engineer will serve as the Problem Management Engineer responsible for leading and maturing the Problem Management function to prevent incident recurrence, reduce operational risk, and improve service resiliency.
The role will oversee the end-to-end Problem Management lifecycle, lead and validate root cause analysis efforts, partner with cross-functional teams to implement permanent fixes, and drive continuous improvement aligned with ITIL best practices.
Key Responsibilities
Lead the end-to-end Problem Management lifecycle
Ensure quality, evidence-based root cause analysis and validation of permanent fixes
Partner with Incident Management, Change Management, Reliability Engineering, Service Owners, and vendors
Establish governance, metrics, reporting, and quality standards for Problem Management
Maintain and improve the Known Error Database (KEDB) and related knowledge articles
Drive proactive problem management, automation, and monitoring improvements
Required Qualifications
Strong understanding of ITIL Problem Management processes and best practices
Experience leading Root Cause Analysis in complex technical environments
Technical background in infrastructure, applications, cloud, or enterprise platforms
Experience with ServiceNow or comparable ITSM platforms
Strong communication and stakeholder management skills
Ability to identify systemic risk and drive proactive improvements
Preferred Qualifications
ITIL Foundation certification or higher
Experience in large-scale enterprise environments
Experience supporting Major Incident or executive outage review forums
Familiarity with automation, observability, and proactive problem management techniques
Experience working with vendors and external service providers