Only W2: SRE / Production Reliability Engineer, Woonsocket, RI

Tror AI for everyone

  • Woonsocket, RI
  • 14 days ago

    Highlights

    The ideal candidate should have strong experience in SRE/DevOps, Incident Management, Observability, Kubernetes, GCP, and production monitoring , along with hands-on experience with time-series anomaly detection . Location: Woonsocket, RI Job Summary We are looking for a Senior SRE / Production Reliability Engineer to improve the reliability, performance, and availability of critical production systems.

    Numbers & Facts

    LocationWoonsocket, RI

    Description

    Role: Senior SRE / Production Reliability Engineer
    Experience: 8+ Years
    Location: Woonsocket, RI
    Job Summary
    We are looking for a Senior SRE / Production Reliability Engineer to improve the reliability, performance, and availability of critical production systems.
    The ideal candidate should have strong experience in SRE/DevOps, Incident Management, Observability, Kubernetes, GCP, and production monitoring, along with hands-on experience with time-series anomaly detection.
    What You'll Do
    • Own reliability and performance of critical production applications.
    • Act as an Incident Commander (IC) during P1/P2 production incidents.
    • Lead incident response, root-cause analysis, postmortems, and reliability improvements.
    • Define and manage SLIs, SLOs, and error budgets.
    • Build and improve monitoring, alerting, and observability solutions.
    • Tune and validate time-series anomaly detection models for production monitoring.
    • Develop automation and operational tools using Python, Java, and React.
    • Troubleshoot Kubernetes, cloud, batch processing, and data pipeline issues.
    • Work with engineering and operations teams to improve system reliability and reduce manual work.
    • Support large-scale deployments and manage production risks such as configuration drift and blast radius.
    Must-Have Skills
    • 8+ years of experience in SRE, DevOps, Platform Engineering, or Production Engineering.
    • Hands-on experience as an Incident Commander for P1/P2 incidents.
    • Strong experience with time-series anomaly detection models in production observability - mandatory.
    • Strong production-level programming skills in:
      • Python
      • Java
      • React
    • Strong experience with SLI, SLO, and error budgets.
    • Hands-on observability experience with:
      • Prometheus
      • Grafana
      • OpenTelemetry
      • At least 2 log platforms such as Loki, Splunk, or Elasticsearch
    • Strong GCP experience.
    • Strong Kubernetes operational experience.
    • Experience with Rancher K3s.
    • Experience troubleshooting Apache Airflow and Tidal workflows/batch jobs.
    • Experience with production-scale distributed systems and on-call support.
    Preferred Skills
    • Production Readiness Reviews / service launch experience.
    • BigQuery and PostgreSQL.
    • Chaos Engineering / fault injection.
    • TIC/Technical Incident Commander certification.
    • Healthcare, pharmacy, retail, or other high-availability environments.
    • LLM/GenAI for incident management, alert summarization, or runbook recommendations.
    • Kafka.
    • Istio / Envoy.
    • Terraform / Ansible.
    Important: The Incident Commander experience and production time-series anomaly detection should be treated as hard requirements. A candidate who only has general monitoring/observability experience but has never worked with anomaly detection models would not be a strong fit.

    Similar Jobs

    See more jobs