About the Business:
LexisNexis Risk Solutions is the essential partner in the assessment of risk. We help customers improve operational efficiency, manage risk, and drive growth through innovative technology solutions and data-driven insights.
About the Role:
We are seeking a Director of Cloud Operations & Reliability to lead enterprise-wide AI operational excellence for large-scale cloud-native platforms. This leader will be responsible for driving availability, reliability, incident management, observability, automation, technical resiliency and operational maturity while leading high-performing operations and reliability teams. The role partners closely with Engineering, Product, Infrastructure, and Client Engagement teams to ensure resilient, customer-focused service delivery.
Responsibilities:
- Lead 24x7 Global cloud operations for mission-critical applications and platforms, ensuring high availability, reliability, and performance.
- Execute on operational reliability strategies aligned with service level objectives (SLOs) and customer outcomes.
- Serve as a hands-on leader during major incidents, driving incident response, stakeholder communications, root cause analysis, and corrective actions.
- Drive operational excellence initiatives through AI-driven operations, self-healing capabilities, and continuous improvement programs.
- Establish enterprise observability standards and playbooks, practices across applications, and infrastructure for Operations surveillance
- Develop and track key operational metrics, including availability, MTTR, incident recurrence, change success rate, and customer impact reduction.
- Build strong cross-functional partnerships across Engineering, Product, Client Engagement, Security, and Infrastructure teams.
- Lead, mentor, and develop high-performing operations leaders and technical teams while fostering a culture of accountability and continuous improvement.
Requirements:
- 9+ years of technology operations experience, including 5+ years managing managers and senior technical professionals.
- Bachelor's degree in computer science, Engineering, or related field; advanced degree preferred.
- Deep expertise with public cloud platforms such as Azure, AWS, or Google Cloud.
- Proven experience supporting large-scale, highly available, mission-critical platforms.
- Strong background in Service Reliability Engineering (SRE), platform operations, incident management, and cloud infrastructure.
- Hands-on experience with observability and monitoring tools such as Grafana and ServiceNow.
- Strong understanding of AIDLC/DevSecOps practices, automation, and AI-driven operational solutions.
- Demonstrated ability to lead major incidents, operational governance, risk management, and executive communications.
- Excellent leadership, stakeholder management, strategic planning, and communication skills.
- Strong experience with process improvement, operational governance, budget management, and organizational change initiatives.
Working for you:
We know that your wellbeing and happiness are key to a long and successful career. These are some of the benefits we are delighted to offer:
- Health Benefits: Comprehensive, multi-carrier program for medical, dental and vision benefits
- Retirement Benefits: 401(k) with match and an Employee Share Purchase Plan
- Wellbeing: Wellness platform with incentives, Headspace app subscription, Employee Assistance and Time-off Programs
- Short-and-Long Term Disability, Life and Accidental Death Insurance, Critical Illness, and Hospital Indemnity
- Family Benefits, including bonding and family care leaves, adoption and surrogacy benefits
- Health Savings, Health Care, Dependent Care and Commuter Spending Accounts
- In addition to annual Paid Time Off, we offer up to two days of paid leave each to participate in Employee Resource Groups and to volunteer with your charity of choice