We're seeking a future team member for the role of SVP. Site Reliability Engineer to join our Technology team. This role is located in Lake Mary, FL and Pittsburgh, PA.
In this role, you'll make an impact in the following ways:
- Design and implement end-to-end observability (logs, metrics, traces) across distributed systemsBuild Observability & Monitoring
- Integrate and optimize tools such as AppDynamics, Dynatrace, Grafana, and Splunk
- Develop dashboards, alerts, and telemetry frameworks to provide real-time visibility
- Identify gaps in monitoring and drive adoption of best practices
Drive Automation & Reduce Toil
- Identify repetitive operational work and automate it using code and tooling
- Build self-healing and auto-remediation solutions
- Enable scalable, reliable processes through automation and engineering rigor
- Improve operational efficiency across production environments
Support Production & Incident Triage
- Troubleshoot and resolve complex production issues across distributed systems
- Participate in incident management, triage, and root cause analysis
- Improve monitoring and automation based on recurring incident patterns
- Collaborate with support and engineering teams to improve system stability
Improve Reliability & Performance
- Define and measure service health using SLIs/SLOs and key performance metrics
- Identify system bottlenecks and reliability risks
- Contribute to performance optimization and capacity planning
- Provide input into system architecture to improve resilience and scalability
To be successful in this role, we're seeking the following:
- 9+ years of experience in Site Reliability Engineering, Software Engineering
- Strong programming background in Java (preferred) or another modern language
- Experience with at least one observability platform:
•AppDynamics, Dynatrace, Grafana, or Splunk
- Hands-on experience supporting and troubleshooting production systems
- Strong analytical and problem-solving skills
- Ability to identify inefficiencies and drive automation
Preferred Qualifications
- Experience with distributed systems or microservices architectures
- Familiarity with CI/CD pipelines and DevOps practices
- Exposure to cloud platforms and/or Kubernetes
- Experience scripting (Python, Bash, etc.) for automation
- Knowledge of SRE concepts like observability, incident management, and reliability engineering
At BNY, our culture allows us to run our company better and enables employees' growth and success. As a leading global financial services company at the heart of the global financial system, we influence nearly 20% of the world's investible assets. Every day, our teams harness cutting-edge AI and breakthrough technologies to collaborate with clients, driving transformative solutions that redefine industries and uplift communities worldwide.
Recognized as a top destination for innovators, BNY is where bold ideas meet advanced technology and exceptional talent. Together, we power the future of finance - and this is what #LifeAtBNY is all about. Join us and be part of something extraordinary.