Senior Site Reliability Engineer ( SRE)

The Charles Schwab Corp

Austin, TX

JOB DETAILS
SKILLS
Agile Programming Methodologies, Amazon Web Services (AWS), Apache Kafka, Artificial Intelligence (AI), Automation, Bash Scripting, Business Operations, Capacity Management, Cloud Applications, Cloud Computing, Computer Science, Continuous Deployment/Delivery, Continuous Improvement, Continuous Integration, Cross-Functional, Customer Support/Service, DHCP (Dynamic Host Configuration Protocol), DNS (Domain Name System), Database Technology, Distributed Computing, Finance, Financial Services, Firewalls, GCP (Good Clinical Practices), GitHub, High Availability Software, IBM WebSphere MQ (Message Queue), Identify Issues, Incident Response, Information Technology & Information Systems, Java, Jenkins, Large-Scale Systems, Linux Administration, Machine Tool, Microsoft .NET, Microsoft SQL Server, Microsoft Windows Azure, Microsoft Windows Server, MongoDB, Network Routing, NoSQL, On Call, Operational Improvement, Operational Strategy, Oracle Database, Performance Tuning/Optimization, Problem Solving Skills, Process Improvement, Production Systems, Python Programming/Scripting Language, RabbitMQ, Relational Databases (RDBMS), Reliability Engineering, Reporting Dashboards, Risk Analysis, Scripting (Scripting Languages), Scrum Project Management and Software Development, Software Development Lifecycle (SDLC), Software Engineering, Splunk, Streaming Technology, System Operations, Systems Engineering, Systems Reliability, Team Player, Technical Operations, Technical Strategy, Validation Plan, Windows PowerShell
LOCATION
Austin, TX
POSTED
5 days ago

Your Opportunity

At Schwab, you're empowered to make an impact on your career. Here, innovative thought meets creative problem solving, helping us challenge the status quo and transform the finance industry together. We believe in the importance of in-office collaboration and fully intend for the selected candidate for this role to work on site in the specified location(s).

As a Senior Reliability Engineer, you will help shape the reliability, scalability, and operational excellence of mission-critical enterprise platforms that support our clients and business operations. This role blends software engineering, systems engineering, and operational expertise to drive resilient, high-performing technology solutions in complex distributed environments. Working closely with engineering, product, scrum, and operations teams, you will apply Site Reliability Engineering (SRE) principles to improve system availability, accelerate delivery, reduce operational complexity, and strengthen platform observability.

You will play a key role in designing and implementing automation solutions that minimize manual effort, improve operational efficiency, and enhance system reliability at scale. Leveraging modern observability practices, AI/ML-enabled operational capabilities, and proactive monitoring strategies, you will help identify risks, detect anomalies, improve incident response, and support data-driven decision making. Your work will influence platform stability through automation, predictive insights, deployment optimization, rollout validation, capacity planning, and continuous improvement initiatives.

Success in this role requires strong problem-solving skills, sound technical judgment, and the ability to collaborate across teams to address complex operational challenges. You will contribute to the evolution of reliability practices, champion automation-first approaches, support continuous delivery initiatives, and help build systems that enable teams to operate with greater confidence, agility, and efficiency. As part of a highly collaborative environment, you will also participate in incident response and on-call support while helping drive long-term improvements that enhance customer and business outcomes.

What you have

Required Qualifications

  • 8+ years of experience supporting and administering enterprise-scale applications, platforms, or infrastructure environments.
  • 6+ years of experience developing automation solutions, operational tooling, monitoring dashboards, and alerting frameworks.
  • 6+ years of experience applying Software Development Life Cycle (SDLC) practices and driving process improvement initiatives.
  • Experience with Site Reliability Engineering (SRE), production operations, system monitoring, deployment management, and operational excellence practices.
  • Experience administering and supporting Linux and Windows Server environments, including troubleshooting, performance tuning, and system optimization.
  • Experience deploying, supporting, configuring, or migrating cloud-based applications and platforms.
  • Knowledge of networking fundamentals including DNS, DHCP, firewalls, routing, and related infrastructure technologies.
  • Experience supporting large-scale distributed systems and highly available application architectures.
  • Proficiency in one or more programming or scripting languages such as Python, Java, PowerShell, Bash, or .NET.
  • Experience working with relational or NoSQL database technologies including SQL Server, Oracle, or MongoDB.
  • Knowledge of messaging and event-streaming technologies such as Kafka, RabbitMQ, IBM MQ, or Solace.
  • Experience with observability and monitoring platforms such as Splunk, AppDynamics, or similar tools.
  • Experience applying AI/ML-powered operational practices, including anomaly detection, predictive alerting, AIOps, or ML-assisted observability capabilities.
  • Bachelor's degree in computer science, Information Technology, Engineering, or a related field.

Preferred Qualifications

  • 8+ years of experience in enterprise technology operations, platform engineering, or reliability engineering.
  • Experience within the financial services industry.
  • Experience working in Agile environments and cross-functional product teams.
  • Hands-on experience with AIOps platforms, intelligent automation solutions, or ML-driven observability tools.
  • Experience integrating AI/ML capabilities into operational automation, deployment workflows, or continuous delivery processes.
  • Experience with CI/CD technologies such as Jenkins, Harness, GitHub Actions, or similar platforms.
  • Experience implementing GitOps practices and infrastructure automation strategies.
  • Experience with containerization and orchestration platforms such as Kubernetes or OpenShift.
  • Experience supporting cloud platforms such as Google Cloud Platform (GCP), AWS, or Microsoft Azure.
  • Demonstrated ability to influence technical strategy, drive operational improvements, and lead reliability-focused initiatives across teams.

In addition to the salary range, this role is eligible for bonus or incentive opportunities.

About the Company

T

The Charles Schwab Corp

The Charles Schwab Corporation is a leading provider of financial services, with more than 300 offices. Through its operating subsidiaries, the company provides a full range of securities brokerage, banking, money management and financial advisory services to individual investors and independent investment advisors. Named "Highest in Investor Satisfaction with Self-Directed Services" by J.D. Power and Associates in 2009, its broker-dealer subsidiary, Charles Schwab & Co., Inc. (member SIPC) affiliates offer a complete range of investment services and products including an extensive selection of mutual funds; financial planning and investment advice; retirement plan and equity compensation plan services; referrals to independent fee-based investment advisors; and custodial, operational and trading support for independent, fee-based investment advisors through Schwab Advisor Services.

The Charles Schwab Bank (member FDIC) provides banking and mortgage services and products. To meet the needs of our clients, we are actively recruiting people with the desire, drive and creativity to find solutions that help meet our clients' needs; who want the chance to learn, grow with the company and explore their career opportunities; who will strive for excellence in achieving our clients' and our company's goals; who have the highest ethical standards - individuals who take pride in making a difference in people's lives.
COMPANY SIZE
1,000 to 1,499 employees
INDUSTRY
Security and Surveillance
FOUNDED
1971
WEBSITE
http://www.aboutschwab.com/careers