SRE - Site Reliability Engineer - Mid-Level

Mindlance

  • Southlake, TX
  • 1 day ago

    Highlights

    Build and maintain dashboards, alerts, and monitoring solutions using Splunk, Grafana, Prometheus, GCP Operations Suite, or similar tools. This role focuses on improving reliability, observability, operational efficiency, and automation across on-premises and cloud platforms.

    Numbers & Facts

    LocationSouthlake, TX

    Description

    ** Onsite 4 days weekly in Southlake or Austin**

    SRE Contractor - Mandatory Qualifications
    Absolute Must-Haves

    3-5 years of hands-on Site Reliability Engineering or Production Engineering experience
    Must have supported production systems at scale.
    Operations ownership, incident response, and reliability engineering experience required.
    Pure DevOps, build/release, or CI/CD-only backgrounds are not a fit.
    Strong Python and operation automation Development Skills
    Demonstrated experience building automation tools, scripts, frameworks, or operational solutions.
    Candidate should be able to provide examples of automation they personally developed.
    Python must be a primary skill, not just basic scripting.
    Production Operations & Incident Management Experience
    Experience troubleshooting critical production incidents.
    Root Cause Analysis (RCA) participation and problem remediation.
    Experience reducing operational toil through automation.

    Vendor Screening Questions
    Before submitting a candidate, please verify:
    What automation solutions has the candidate personally built?
    What operational toil have they automated?
    Describe a major production incident they helped resolve.
    How much of their current role is SRE versus DevOps/build-release work?


    Job Summary
    We are seeking a motivated Site Reliability Engineer (Contractor) with 3 to 5 years of experience in automation, cloud infrastructure, and production operations. The ideal candidate has a strong automation-first mindset and enjoys solving complex operational challenges. This role focuses on improving reliability, observability, operational efficiency, and automation across on-premises and cloud platforms.
    Key Responsibilities
    Automation & Platform Engineering
    • Develop Python-based automation solutions to reduce manual operational effort.
    • Automate infrastructure management across Linux, Windows, Kubernetes, GCP, and cloud-native environments.
    • Integrate tools and platforms through APIs and client libraries.
    • Assist in implementing infrastructure automation using Ansible, Terraform, or similar technologies.
    • Support CI/CD automation and deployment reliability initiatives.
    Reliability & Operations
    • Monitor and maintain production systems to meet reliability and availability objectives.
    • Participate in incident response, troubleshooting, and root cause analysis activities.
    • Develop automation and operational improvements to prevent recurring issues.
    • Support disaster recovery, failover testing, and operational readiness activities.
    • Perform performance analysis and system health reviews.
    Observability & Monitoring
    • Build and maintain dashboards, alerts, and monitoring solutions using Splunk, Grafana, Prometheus, GCP Operations Suite, or similar tools.
    • Improve visibility into application and infrastructure health through metrics, logs, and traces.
    • Investigate alerts and identify opportunities to reduce noise and improve detection.
    AIOps & Intelligence
    • Explore AI/ML-driven operational improvements such as anomaly detection, intelligent alerting, and log analytics.
    • Assist in developing automation solutions that leverage AI to improve operational efficiency.
    • Participate in evaluating emerging AIOps capabilities and observability technologies.
    Required Qualifications
    • Bachelor's degree in Computer Science, Engineering, or related field, or equivalent experience.
    • 3 to 5 years of experience in Site Reliability Engineering, DevOps, Systems Engineering, or Platform Engineering.
    • Strong programming skills in Python for automation and tooling development.
    • Experience supporting Kubernetes and cloud platforms (GCP, AWS, or Azure).
    • Familiarity with infrastructure automation and configuration management tools.
    • Experience with monitoring and observability platforms such as Splunk, Grafana, Prometheus, Datadog, or similar.
    • Understanding of Linux systems, networking, and distributed applications.
    • Strong analytical, troubleshooting, and problem-solving skills.
    • Ability to work effectively in fast-paced, mission-critical environments.
    Preferred Qualifications
    • Experience with Terraform, Ansible, or Infrastructure as Code solutions.
    • Exposure to OpenTelemetry and modern observability practices.
    • Experience with CI/CD pipelines and deployment automation.
    • Knowledge of AI/ML, AIOps, or intelligent operational tooling.
    • Experience supporting highly available production systems in regulated or enterprise environments.

    “Mindlance is an Equal Opportunity Employer and does not discriminate in employment on the basis of – Minority/Gender/Disability/Religion/LGBTQI/Age/Veterans.”

    Similar Jobs

    See more jobs