Senior IT Triage Engineer

First Citizens Bank

Raleigh, North Carolina

JOB DETAILS
JOB TYPE
Full-time
SKILLS
Analysis Skills, Application Programming Interface (API), Automation, Business Services, Cloud Applications, Cloud Computing, Communication Skills, Compensation and Benefits, Continuous Improvement, Corrective Action, Cross-Functional, Customer Support/Service, Event Correlation, Event Management, Hardware Virtualization, High School Diploma, ITIL (IT Infrastructure Library), Identify Issues, Incident Management, Information Technology & Information Systems, Information/Data Security (InfoSec), Linux Administration, Mentoring, Microsoft Hyper-V, Microsoft Windows System Administration, Operational Improvement, Operational Support, Operations Management, Operations Processes, Performance Analysis, Presentation/Verbal Skills, Problem Solving Skills, Process Improvement, Records Management, Reliability Engineering, Reporting Dashboards, Risk Management, Root Cause Analysis, Software Development, System Architecture, Systems Administration/Management, Systems Analysis, Technical Leadership, Technical Operations, Technical Support, Telemetry, VMWare, Writing Skills
LOCATION
Raleigh, North Carolina
POSTED
4 days ago
Overview:

The Senior IT Triage Engineer serves as a critical technical leader responsible for the rapid diagnosis, restoration, and resolution of complex technology incidents across enterprise infrastructure, cloud platforms, applications, and end-user services. This role combines deep technical troubleshooting expertise with a proactive focus on eliminating recurring issues through structured problem management, root cause analysis, and continuous service improvement.

The ideal candidate excels in high-pressure operational environments, drives technical resolution efforts across multiple teams, and leverages automation, observability, and reliability practices to improve service stability and reduce operational risk.

Responsibilities:

Key Responsibilities

  • Incident Management, Service Restoration, and System Analysis.
  • Lead technical triage efforts for high-priority incidents and service disruptions.
  • Coordinate cross-functional teams to restore critical business services as quickly as possible.
  • Analyze alerts, logs, monitoring data, and telemetry to identify the source of issues.
  • Serve as a senior escalation point for complex infrastructure, cloud, network, application, and platform incidents.
  • Drive incident bridges, facilitate technical discussions, and maintain clear communication with stakeholders.
  • Problem Management & Root Cause Elimination
  • Lead root cause investigations for recurring or significant incidents.
  • Develop and drive corrective and preventive action plans across technology teams.
  • Track and manage problem records through resolution.
  • Identify systemic issues and technical debt that impact service stability.
  • Attend post-incident reviews and ensure lessons learned are translated into operational improvements.

 Reliability & Operational Excellence

  • Continuously improve service availability, resiliency, and operational performance.
  • Partner with engineering teams to improve monitoring, alerting, telemetry, and observability capabilities.
  • Reduce alert noise through event correlation, automation, and process optimization.
  • Drive efforts to improve mean time to detect (MTTD) and mean time to restore service (MTTR).

Automation & Process Improvement

  • Identify opportunities to automate operational workflows and repetitive support activities.
  • Develop runbooks, playbooks, and operational procedures.
  • Collaborate with engineering teams to implement self-healing and automated recovery mechanisms.

Technical Leadership

  • Provide mentorship and guidance to engineers and operational support teams.
  • Influence technical decision-making related to operational readiness and supportability.
  • Act as a trusted advisor for service reliability, supportability, and operational risk management.
  • Participate in change reviews to ensure production readiness and minimize operational impact.
Qualifications:

Bachelor's Degree and 8 years of experience in Technical work in Application Development, Server Administration, Information Security, or Engineering OR High School Diploma or GED and 12 years of experience in Technical work in Application Development, Server Administration, Information Security, or Engineering

  • 7+ years of experience in enterprise IT operations, infrastructure engineering, platform operations, or technical support environments.
  • Proven experience managing and resolving critical production incidents across multiple system architectures and infrastructures.
  • Strong background in problem management and root cause analysis methodologies.
  • Experience supporting large-scale enterprise environments.
  • Strong understanding of ITIL service management practices.

Technical Skills

  • Modern application architecture patterns and operations
  • Infrastructure & Platforms
  • Windows and Linux administration
  • Virtualization platforms (VMware, Hyper-V, etc.)
  • Storage and backup technologies
  • Containers and orchestration platforms (OpenShift)

Monitoring & Observability

  • Enterprise monitoring platforms
  • Log aggregation and analysis tools
  • Application performance monitoring (APM)
  • Operational dashboards and telemetry systems
  • Event management solutions

Automation

  • PowerShell, Python, Bash, or equivalent scripting
  • Workflow automation platforms
  • API integration and orchestration
  • Runbook development

Key Competencies

  • Exceptional troubleshooting and analytical skills.
  • Strong sense of ownership and accountability.
  • Ability to perform effectively during major incidents and high-pressure situations.
  • Excellent collaboration and stakeholder management skills.
  • Strong written and verbal communication.
  • Data-driven decision making.
  • Systems thinking and problem-solving mindset.
  • Continuous improvement orientation.

Benefits are an integral part of total rewards and First Citizens Bank is committed to providing a competitive, thoughtfully designed and quality benefits program to meet the needs of our associates. More information can be found at https://jobs.firstcitizens.com/benefits.

About the Company

F

First Citizens Bank

First Citizens Bank helps personal, business, commercial and wealth clients build financial strength that lasts. As the largest family-controlled bank in the United States, First Citizens is continuing a unique legacy of strength, stability, and long-term thinking that has spanned generations. Founded in 1898 and headquartered in Raleigh, N.C., First Citizens also operates a nationwide direct bank and a network over 550 branches in 22 states. Industry specialists bring a depth of expertise that helps businesses and individuals meet their specific goals at every stage of their financial journey. First Citizens Bank brings together personal service and powerful tools to help customers do more with their money – and make more of their future.  

Looking for a career with CIT? CIT is now a division of First Citizens Bank.

First Citizens Bank. Forever First®

COMPANY SIZE
10,000 employees or more
INDUSTRY
Banking
EMPLOYEE BENEFITS
Paid Sick Days, Prescription Drug Coverage, Professional Development, 401K, Flexible Spending Accounts, Retirement / Pension Plans, Life Insurance
FOUNDED
1898
WEBSITE
https://www.firstcitizens.com