Senior SRE

Selby Jennings Ltd

  • Austin, TX
  • 8 days ago

    Highlights

    This role sits within a Platform Engineering team responsible for delivering highly available, resilient, and scalable infrastructure that powers business-critical workloads and data processes. A leading investment management firm is seeking a Senior Site Reliability Engineer to help scale and support critical workflow orchestration and automation platforms across the organization.

    Numbers & Facts

    LocationAustin, TX

    Description

    Senior Site Reliability Engineer

    A leading investment management firm is seeking a Senior Site Reliability Engineer to help scale and support critical workflow orchestration and automation platforms across the organization. This role sits within a Platform Engineering team responsible for delivering highly available, resilient, and scalable infrastructure that powers business-critical workloads and data processes.

    The ideal candidate will have a strong background in SRE, DevOps, or Platform Engineering and enjoy balancing hands-on operational support with long-term engineering improvements. You'll work closely with engineering, data, infrastructure, and security teams to enhance platform reliability, automate manual processes, and drive modernization efforts across the environment.

    Responsibilities

    • Provide operational ownership of workflow orchestration and enterprise scheduling platforms, including Apache Airflow and similar technologies.
    • Act as an escalation point for platform-related incidents, troubleshooting complex production issues and driving root cause analysis through resolution.
    • Partner with application and data teams to resolve workflow failures, dependency issues, scheduling conflicts, and performance bottlenecks.
    • Develop and maintain reliability standards, service objectives, monitoring strategies, and operational best practices.
    • Build automation and tooling that reduce manual effort and improve the overall user experience for engineering teams.
    • Design and implement observability solutions utilizing metrics, dashboards, alerting, logging, and performance monitoring.
    • Support platform lifecycle management, including upgrades, patching, configuration management, and security remediation.
    • Contribute to infrastructure modernization initiatives involving cloud services, containerization, platform migrations, and deployment automation.
    • Develop and maintain Infrastructure-as-Code solutions using tools such as Terraform, Helm, Ansible, and related technologies.
    • Participate in architectural discussions, platform roadmap planning, and engineering standards development.
    • Maintain operational documentation, procedures, and on-call readiness for supported environments.

    Required Experience

    • Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent experience.
    • 5+ years of experience in Site Reliability Engineering, Platform Engineering, Infrastructure Engineering, or DevOps-focused roles.
    • Strong production experience supporting Apache Airflow environments.
    • Experience with distributed Airflow deployments, including Celery and/or Kubernetes executors.
    • Experience supporting enterprise workload automation and job scheduling platforms such as Automic/UC4, Control-M, or comparable technologies.
    • Strong Linux administration skills with working knowledge of Windows-based environments.
    • Experience supporting cloud infrastructure, preferably within AWS environments.
    • Proficiency with Python and scripting for automation, tooling, and operational efficiency.
    • Experience working with Kubernetes, Docker, CI/CD pipelines, and modern deployment methodologies.
    • Strong understanding of monitoring, logging, tracing, and observability concepts.
    • Experience with tools such as Grafana, Prometheus, ELK, or comparable monitoring platforms.
    • Proven ability to manage production incidents and communicate effectively during high-priority situations.
    • Strong automation mindset with a focus on improving efficiency and reducing operational overhead.

    Preferred Qualifications

    • Hands-on experience administering Broadcom Automic/UC4.
    • Experience with managed Airflow platforms such as AWS MWAA, Cloud Composer, or Astronomer.
    • Exposure to modern data ecosystems, including technologies such as Kafka, dbt, and Snowflake.
    • Experience migrating workloads from legacy scheduling platforms to cloud-native orchestration solutions.
    • Familiarity with SLOs, SLIs, error budgets, and reliability engineering best practices.

    Similar Jobs