Site Reliability Engineer (OpenSearch) - Remote

Information Consulting Services

  • Herndon, VA
  • 5 days ago
  • Remote

    Highlights

    Experience building/implementing/supporting cloud monitoring and observability; solid knowledge of cloud computing, infrastructure operations, databases, web services, networking, virtualization, and internet protocols. Deep OpenSearch administration: cluster architecture, performance tuning, scaling, upgrades, and troubleshooting; index/shard/replica strategy; sizing; snapshot/restore; backup/DR.

    Numbers & Facts

    LocationHerndon, VA (
    Remote
    )

    Description

    Contract Details

    • Work Mode: 100% Remote (US-based)
    • Location: Herndon, VA
    • Schedule: 40 hours/week
    • Duration: 08/17/2026 08/16/2027
    • Type: Contract with potential to convert to full-time after ~12 months (not guaranteed)

    About the Opportunity

    Seeking a Site Reliability Engineer to ensure availability, performance, scalability, and security for mission-critical, cloud-hosted search and analytics services built on OpenSearch. You will focus on reliability engineering, operations, automation, and continuous improvement for distributed platforms, working within a diverse, globally distributed team.

    Key Responsibilities

    • Provision, build, deploy, monitor, operate, and support cloud services in a global team environment.
    • Architect, build, deploy, and maintain high performance OpenSearch clusters and platforms from the ground up.
    • Optimize OpenSearch for high availability, resiliency, scalability, security, and performance.
    • Monitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation/replication, and storage utilization.
    • Analyze and resolve operational issues across infrastructure, platform, and application layers; lead incident response, RCA, and remediation.
    • Maintain integrity and security of servers, systems, and OpenSearch platform infrastructure.
    • Support lifecycle activities: installation, configuration, upgrades/patching, backup/restore, and disaster recovery.
    • Develop and maintain monitoring policies, alerting standards, runbooks, and support procedures.
    • Automate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and related cloud services.
    • Plan capacity for compute, memory, storage, and network; partner with engineering to enhance reliability and operational readiness.
    • Support log ingestion, index management, lifecycle/retention, and search performance tuning.
    • Participate in an on-call rotation; support occasional weekend/after-hours needs.

    Required Qualifications

    • US citizenship required; dual citizenship not permitted.
    • 8 years of experience in SRE/DevOps/cloud operations with distributed systems.
    • Proven, hands-on experience designing, building, deploying, operating, and optimizing OpenSearch clusters from scratch in production.
    • Expert-level Kubernetes experience (operations, troubleshooting, management, configuration of complex services).
    • Deep OpenSearch administration: cluster architecture, performance tuning, scaling, upgrades, and troubleshooting; index/shard/replica strategy; sizing; snapshot/restore; backup/DR.
    • Strong Linux expertise (SUSE and Ubuntu).
    • Expertise with Git and Concourse (pipeline setup, management, troubleshooting).
    • Experience with Kafka and Zookeeper; strong automation for testing, deployment, scalability, and cloud service management.
    • Experience building/implementing/supporting cloud monitoring and observability; solid knowledge of cloud computing, infrastructure operations, databases, web services, networking, virtualization, and internet protocols.
    • Security fundamentals for SaaS multi-tenant application systems; excellent communication and prioritization skills; ability to multitask.

    Preferred Qualifications

    • AWS experience (e.g., Route 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, VPC); experience deploying/operating OpenSearch in AWS.
    • Experience with Cloud Foundry environments.
    • Experience with Jenkins, Chef, and/or Terraform.
    • Experience with Prometheus and Grafana.
    • Background with log ingestion pipelines, index lifecycle management, retention strategies, and search platform security controls.
    • Familiarity with capacity forecasting, performance benchmarking, and resilience testing for distributed search platforms.

    Work Environment

    • Collaborative, globally distributed team with cross-training opportunities.
    • Participation in an on-call rotation and occasional after-hours/weekend support.
    #ZR

    Similar Jobs

    See more jobs