Site Reliability Engineer OpenSearch - Remote

The Dignify Solutions, LLC

  • Newtown Square, PA
  • 1 day ago
  • Remote

    Highlights

    Deep experience in site reliability engineering, DevOps, cloud operations, automation, observability, and distributed systems, with proven hands-on expertise architecting, building, deploying, operating, and optimizing high-performance OpenSearch clusters and platforms from the ground up in production environments. 10+ years of experience in Site Reliability Engineer – OpenSearch to help ensure the highest levels of availability, performance, scalability, and Quality of Service (QoS) for mission-critical cloud services.

    Numbers & Facts

    LocationNewtown Square, PA (
    Remote
    )

    Description

    • 10+ years of experience in Site Reliability Engineer – OpenSearch to help ensure the highest levels of availability, performance, scalability, and Quality of Service (QoS) for mission-critical cloud services.
    • Deep experience in site reliability engineering, DevOps, cloud operations, automation, observability, and distributed systems, with proven hands-on expertise architecting, building, deploying, operating, and optimizing high-performance OpenSearch clusters and platforms from the ground up in production environments.
    • Expert with Kubernetes, including troubleshooting, operations, management, and configuration of complex Kubernetes services.
    • Proven hands-on expertise designing, building, deploying, supporting, and maintaining OpenSearch clusters and platforms from scratch in production environments
    • Strong experience with OpenSearch administration, cluster architecture, performance tuning, scaling, upgrades, and troubleshooting
    • Experience with index design, shard and replica strategy, cluster sizing, node management, snapshot/restore, backup, and disaster recovery
    • Strong understanding of distributed systems, search platforms, indexing pipelines, query optimization, and high-availability architectures
    • Expertise with Git
    • Expertise with Concourse, including setup, management, and troubleshooting of new pipelines
    • Expertise with Linux, specifically SUSE and Ubuntu
    • Expertise with Kafka, Zookeeper, and Big Data technologies
    • Expert in development of automation for testing, deployment, scalability, and management of cloud services
    • Expertise with building, implementing, and/or supporting cloud monitoring tools
    • Expert knowledge of cloud computing, infrastructure operations, and databases
    • Expert understanding of web services, networking, virtualization, and internet protocols
    • Ability to multitask and handle various projects, deadlines, and changing priorities
    • Excellent communication and prioritization skills
    • Expertise with security fundamentals as they pertain to SaaS multi-tenant application systems
    • Experience with AWS services including Route 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, and VPC
    • Experience deploying and operating OpenSearch in AWS-based environments
    • Experience with Cloud Foundry-based environments
    • Experience with Jenkins, Chef, and/or Terraform
    • Exposure to and understanding of troubleshooting IP networks and application stacks
    • Experience with observability tools such as Prometheus and Grafana
    • Experience with log ingestion pipelines, index lifecycle management, retention strategies, and search platform security controls
     

     

    Similar Jobs

    See more jobs