Senior Site Reliability Engineer
A leading investment management firm is seeking a Senior Site Reliability Engineer to help scale and support critical workflow orchestration and automation platforms across the organization. This role sits within a Platform Engineering team responsible for delivering highly available, resilient, and scalable infrastructure that powers business-critical workloads and data processes.
The ideal candidate will have a strong background in SRE, DevOps, or Platform Engineering and enjoy balancing hands-on operational support with long-term engineering improvements. You'll work closely with engineering, data, infrastructure, and security teams to enhance platform reliability, automate manual processes, and drive modernization efforts across the environment.
Responsibilities
- Provide operational ownership of workflow orchestration and enterprise scheduling platforms, including Apache Airflow and similar technologies.
- Act as an escalation point for platform-related incidents, troubleshooting complex production issues and driving root cause analysis through resolution.
- Partner with application and data teams to resolve workflow failures, dependency issues, scheduling conflicts, and performance bottlenecks.
- Develop and maintain reliability standards, service objectives, monitoring strategies, and operational best practices.
- Build automation and tooling that reduce manual effort and improve the overall user experience for engineering teams.
- Design and implement observability solutions utilizing metrics, dashboards, alerting, logging, and performance monitoring.
- Support platform lifecycle management, including upgrades, patching, configuration management, and security remediation.
- Contribute to infrastructure modernization initiatives involving cloud services, containerization, platform migrations, and deployment automation.
- Develop and maintain Infrastructure-as-Code solutions using tools such as Terraform, Helm, Ansible, and related technologies.
- Participate in architectural discussions, platform roadmap planning, and engineering standards development.
- Maintain operational documentation, procedures, and on-call readiness for supported environments.
Required Experience
- Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent experience.
- 5+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, or DevOps-focused roles.
- Strong production experience supporting Apache Airflow environments.
- Experience with distributed Airflow deployments, including Celery and/or Kubernetes executors.
- Experience supporting enterprise workload automation and job scheduling platforms such as Automic/UC4, Control-M, or comparable technologies.
- Strong Linux administration skills with working knowledge of Windows-based environments.
- Experience supporting cloud infrastructure, preferably within AWS environments.
- Proficiency with Python and scripting for automation, tooling, and operational efficiency.
- Experience working with Kubernetes, Docker, CI/CD pipelines, and modern deployment methodologies.
- Strong understanding of monitoring, logging, tracing, and observability concepts.
- Experience with tools such as Grafana, Prometheus, ELK, or comparable monitoring platforms.
- Proven ability to manage production incidents and communicate effectively during high-priority situations.
- Strong automation mindset with a focus on improving efficiency and reducing operational overhead.
Preferred Qualifications
- Hands-on experience administering Broadcom Automic/UC4.
- Experience with managed Airflow platforms such as AWS MWAA, Cloud Composer, or Astronomer.
- Exposure to modern data ecosystems, including technologies such as Kafka, dbt, and Snowflake.
- Experience migrating workloads from legacy scheduling platforms to cloud-native orchestration solutions.
- Familiarity with SLOs, SLIs, error budgets, and reliability engineering best practices.