Senior SRE Workflow Automation

Selby Jennings Ltd

  • Austin, TX
  • 3 days ago

    Highlights

    The successful candidate will partner closely with infrastructure, application, and data-focused teams to maintain platform stability while driving continuous improvements across tooling, observability, and automation. A growing engineering organization is seeking an experienced Platform Reliability Engineer to support and enhance critical workflow automation and job scheduling infrastructure.

    Numbers & Facts

    LocationAustin, TX

    Description

    A growing engineering organization is seeking an experienced Platform Reliability Engineer to support and enhance critical workflow automation and job scheduling infrastructure. This position combines hands-on operational ownership with platform engineering initiatives focused on improving reliability, scalability, automation, and user experience.

    The role offers a balanced mix of production support responsibilities and engineering project work. The successful candidate will partner closely with infrastructure, application, and data-focused teams to maintain platform stability while driving continuous improvements across tooling, observability, and automation.

    Key Responsibilities

    Reliability & Operations

    • Provide advanced support for enterprise workflow orchestration and scheduling platforms.
    • Investigate production incidents, perform root cause analysis, and implement preventative solutions.
    • Define and maintain service reliability metrics, operational objectives, and performance standards.
    • Monitor platform health, availability, and capacity, proactively addressing potential issues.
    • Troubleshoot workflow execution failures, scheduling problems, and platform performance concerns.
    • Manage platform lifecycle activities, including upgrades, patching, and configuration management.
    • Partner with security and infrastructure teams to address vulnerabilities and maintain operational best practices.
    • Participate in incident response and on-call support rotations.
    • Develop and maintain operational documentation, procedures, and knowledge resources.

    Platform Engineering & Automation

    • Design and implement automation solutions that reduce manual effort and improve platform operations.
    • Contribute to infrastructure modernization efforts, deployment improvements, and platform enhancements.
    • Develop and maintain Infrastructure-as-Code solutions for platform provisioning and management.
    • Build monitoring, alerting, logging, and reporting capabilities that improve visibility and operational insights.
    • Establish platform standards, governance, and best practices to support engineering teams at scale.
    • Participate in architecture discussions, technical reviews, and long-term platform planning.
    • Improve developer and user experience through self-service tools and operational efficiencies.

    Required Qualifications

    • Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent experience.
    • 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or related disciplines.
    • Strong experience supporting workflow orchestration, automation, or enterprise scheduling platforms.
    • Experience operating distributed applications in cloud-based environments.
    • Solid Linux administration skills; Windows experience is a plus.
    • Proficiency with scripting or programming languages used for automation and operational tooling.
    • Experience with containerization technologies and modern deployment practices.
    • Familiarity with CI/CD pipelines and Infrastructure-as-Code methodologies.
    • Strong understanding of monitoring, observability, logging, and performance analysis principles.
    • Proven ability to manage production incidents and prioritize issues in fast-paced environments.
    • Strong communication and cross-functional collaboration skills.
    • Passion for automation, operational excellence, and continuous improvement.

    Preferred Qualifications

    • Experience supporting enterprise workload automation or scheduling systems.
    • Background working with managed cloud platform services.
    • Familiarity with modern data engineering or large-scale analytics ecosystems.
    • Experience with platform migrations, modernization efforts, or large-scale infrastructure transformations.
    • Knowledge of reliability engineering concepts, including service objectives and operational performance frameworks.

    Similar Jobs