A growing engineering organization is seeking an experienced Platform Reliability Engineer to support and enhance critical workflow automation and job scheduling infrastructure. This position combines hands-on operational ownership with platform engineering initiatives focused on improving reliability, scalability, automation, and user experience.
The role offers a balanced mix of production support responsibilities and engineering project work. The successful candidate will partner closely with infrastructure, application, and data-focused teams to maintain platform stability while driving continuous improvements across tooling, observability, and automation.
Key Responsibilities
Reliability & Operations
- Provide advanced support for enterprise workflow orchestration and scheduling platforms.
- Investigate production incidents, perform root cause analysis, and implement preventative solutions.
- Define and maintain service reliability metrics, operational objectives, and performance standards.
- Monitor platform health, availability, and capacity, proactively addressing potential issues.
- Troubleshoot workflow execution failures, scheduling problems, and platform performance concerns.
- Manage platform lifecycle activities, including upgrades, patching, and configuration management.
- Partner with security and infrastructure teams to address vulnerabilities and maintain operational best practices.
- Participate in incident response and on-call support rotations.
- Develop and maintain operational documentation, procedures, and knowledge resources.
Platform Engineering & Automation
- Design and implement automation solutions that reduce manual effort and improve platform operations.
- Contribute to infrastructure modernization efforts, deployment improvements, and platform enhancements.
- Develop and maintain Infrastructure-as-Code solutions for platform provisioning and management.
- Build monitoring, alerting, logging, and reporting capabilities that improve visibility and operational insights.
- Establish platform standards, governance, and best practices to support engineering teams at scale.
- Participate in architecture discussions, technical reviews, and long-term platform planning.
- Improve developer and user experience through self-service tools and operational efficiencies.
Required Qualifications
- Bachelor's degree in Computer Science, Engineering, Information Systems, or equivalent experience.
- 5+ years of experience in Site Reliability Engineering, DevOps, Infrastructure Engineering, or related disciplines.
- Strong experience supporting workflow orchestration, automation, or enterprise scheduling platforms.
- Experience operating distributed applications in cloud-based environments.
- Solid Linux administration skills; Windows experience is a plus.
- Proficiency with scripting or programming languages used for automation and operational tooling.
- Experience with containerization technologies and modern deployment practices.
- Familiarity with CI/CD pipelines and Infrastructure-as-Code methodologies.
- Strong understanding of monitoring, observability, logging, and performance analysis principles.
- Proven ability to manage production incidents and prioritize issues in fast-paced environments.
- Strong communication and cross-functional collaboration skills.
- Passion for automation, operational excellence, and continuous improvement.
Preferred Qualifications
- Experience supporting enterprise workload automation or scheduling systems.
- Background working with managed cloud platform services.
- Familiarity with modern data engineering or large-scale analytics ecosystems.
- Experience with platform migrations, modernization efforts, or large-scale infrastructure transformations.
- Knowledge of reliability engineering concepts, including service objectives and operational performance frameworks.