Site Reliability Engineer

Iconma LLC

  • AL
  • 22 days ago

    Highlights

    Collaboration: Work closely with developers, scientists, and infrastructure teams to deliver reliable platform services and translate operational needs into sustainable engineering solutions. GitOps and Deployment Automation: Support and improve GitOps workflows using ArgoCD to manage cluster and application configuration in a consistent, auditable manner.

    Numbers & Facts

    LocationAL

    Description

    Our Client, a Business Maunufacturing and Supply company, is looking for a Site Reliability Engineer for their Remote location.

    Responsibilities:

    • Platform Operations: Maintain and enhance Kubernetes platforms across on-premises and cloud environments, ensuring reliability, scalability, and operational efficiency.
    • Cluster Management: Support provisioning, upgrades, troubleshooting, and lifecycle management of Kubernetes clusters managed through Rancher.
    • Linux Systems Administration: Provide deep technical expertise in Linux-based systems, including performance tuning, troubleshooting, automation, and operational support.
    • Infrastructure as Code: Develop and maintain infrastructure-as-code solutions to standardize and automate platform deployment and management, with a preference for Cluster API (CAPI)-based approaches.
    • GitOps and Deployment Automation: Support and improve GitOps workflows using ArgoCD to manage cluster and application configuration in a consistent, auditable manner.
    • Collaboration: Work closely with developers, scientists, and infrastructure teams to deliver reliable platform services and translate operational needs into sustainable engineering solutions.
    • Continuous Improvement: Identify opportunities to improve platform resilience, observability, security, and maintainability through automation and modern SRE practices.

    Requirements:

    • BS in Computer Science, Software Engineering, Information Technology, or related field preferred; or equivalent professional experience.
    • 5+ years of professional experience in site reliability engineering, platform engineering, DevOps, or systems engineering roles.
    • Hands-on experience operating and supporting Kubernetes platforms in production environments.
    • Strong experience managing Kubernetes clusters in both on-premises and cloud-based environments.
    • Strong Linux systems administration skills, including troubleshooting, scripting, networking, and system performance analysis.
    • Experience with Rancher for Kubernetes cluster management and platform operations.
    • Experience implementing infrastructure-as-code solutions for platform provisioning and lifecycle management.
    • Demonstrated success working in Agile teams (Scrum, Kanban).
    • Cluster operations, upgrades, networking, storage, troubleshooting, and workload support.
    • Platform Management: Rancher or similar Kubernetes management platforms.
    • Linux: Advanced administration of Linux/Unix systems.
    • Infrastructure as Code: Strong IaC experience; Cluster API (CAPI) preferred.
    • GitOps/CI-CD: ArgoCD, Git version control, and deployment automation practices.
    • Scripting/Automation: Bash, Python, or similar scripting languages for automation and operational tooling.
    • Experience with hybrid infrastructure spanning on-premises and public cloud platforms (AWS, Azure, GCP).
    • Experience with Kubernetes ecosystem tooling for observability, logging, monitoring, and alerting.
    • Familiarity with security best practices for Kubernetes and Linux platforms.
    • Experience supporting scientific research environments, high-performance computing, or computational science workflows.
    • Knowledge of CI/CD pipeline development and platform automation patterns.

    Why Should You Apply?

    • Health Benefits
    • Referral Program
    • Excellent growth and advancement opportunities

    Similar Jobs

    See more jobs