Wells Fargo & Co logo

Lead Site Reliability Engineer

Wells Fargo & Co

  • Des Moines, IA
  • 4 days ago

    Highlights

    Summary: We are seeking a Senior Site Reliability Engineer (SRE) to join a growing team responsible for the reliability, observability, automation, and operational excellence of critical enterprise applications and AI-powered platforms. This role is ideal for an experienced SRE who thrives in complex production environments and is passionate about driving platform stability, automation, self-healing capabilities, and continuous improvement.

    Numbers & Facts

    LocationDes Moines, IA
    IndustryFinancial Services
    Company Size10,000 employees or more
    Year Founded1852

    Description

    Title: Lead Site Reliability Engineer

    Location: Des Moines, IA

    Alternate Location: Minneapolis, MN / Irving, TX

    Duration: 18 months

    Work Engagement: W2

    Work Schedule: Hybrid 3 days in office / 2 days remote

    Benefits on offer for this contract position: Health Insurance, Life insurance, 401K and Voluntary Benefits

    Summary:

    We are seeking a Senior Site Reliability Engineer (SRE) to join a growing team responsible for the reliability, observability, automation, and operational excellence of critical enterprise applications and AI-powered platforms. This role is ideal for an experienced SRE who thrives in complex production environments and is passionate about driving platform stability, automation, self-healing capabilities, and continuous improvement.

    The successful candidate will bring deep expertise in production support, incident management, observability, and automation while helping the team evolve its reliability engineering practices. Experience supporting AI/ML and LLM-based systems, including emerging Agentic AI solutions, is highly desired.

    Experience within banking, financial services, or other highly regulated industries is strongly preferred.

    Key Responsibilities:

    • Lead and support large-scale production environments using SRE and ITIL best practices.

    • Manage incident, problem, and change management processes to ensure service reliability and operational excellence.

    • Monitor, troubleshoot, and optimize business-critical applications and platforms.

    • Design and implement observability solutions to improve system visibility and performance.

    • Automate operational processes and develop self-healing solutions to reduce manual intervention.

    • Support application deployments and continuous delivery pipelines.

    • Drive root cause analysis and remediation efforts for complex production issues.

    • Partner with development, infrastructure, and platform engineering teams to improve system resilience and scalability.

    • Support AI/ML and LLM-based applications in production environments while ensuring reliability, performance, and operational governance.

    • Contribute to platform modernization and operational maturity initiatives.

    Required Qualifications:

    • Applicants must be authorized to work for ANY employer in the U.S. This position is not eligible for visa sponsorship.

    • 8+ years of experience in Site Reliability Engineering (SRE), Production Support, Platform Engineering, or a related discipline.

    • Proven ability to lead and support large-scale, mission-critical production environments using ITIL-based incident, problem, and change management practices.

    • Strong SRE mindset with a passion for improving reliability, operational excellence, automation, and team maturity.

    • Experience supporting AI/ML and Large Language Model (LLM) platforms in production environments, with a solid understanding of Agentic AI concepts, architectures, and operational considerations.

    • Hands-on experience managing monitoring, observability, deployments, automation, incident response, and self-healing application capabilities.

    • Strong background supporting hybrid technology environments spanning on-premises infrastructure, public cloud, and hybrid cloud platforms.

    • Experience providing L2/L3 production support for enterprise applications, including troubleshooting complex system and application issues.

    About Company

    We believe in our vision and values just as strongly today as we did the first time we put them on paper more than 20 years ago. Staying true to them will guide us toward continued growth and success for decades to come. As you read more about our vision and values, you will learn about who we are, where we’re headed and how every Wells Fargo team member can help us get there.

    Similar Jobs

    See more jobs