Site Reliability Engineer OpenSearch

TechWish

  • Princeton, NJ
  • 4 days ago
  • Remote

    Highlights

    The ideal candidate brings deep experience in site reliability engineering, DevOps, cloud operations, automation, observability, and distributed systems, with proven hands-on expertise architecting, building, deploying, operating, and optimizing high-performance OpenSearch clusters and platforms from the ground up in production environments . SAP NS2 is seeking a Site Reliability Engineer OpenSearch to help ensure the highest levels of availability, performance, scalability, and Quality of Service (QoS) for mission-critical cloud services.

    Numbers & Facts

    LocationPrinceton, NJ (
    Remote
    )

    Description

    Description:

    MUST be a US Citizen & ONLY hold US Citizenship (No Dual Citizens)

    Fully remote position

    Possible conver to hire after 1 year

    PLEASE only submit candidates that have OpenSearch deployment ON KUBERNETES experience

    Site Reliability Engineer OpenSearch

    SAP NS2 is seeking a Site Reliability Engineer OpenSearch to help ensure the highest levels of availability, performance, scalability, and Quality of Service (QoS) for mission-critical cloud services. This role will focus on the reliability, operations, automation, and continuous improvement of distributed search and analytics platforms built on OpenSearch, while working in a diverse, globally distributed team environment.

    The ideal candidate brings deep experience in site reliability engineering, DevOps, cloud operations, automation, observability, and distributed systems, with proven hands-on expertise architecting, building, deploying, operating, and optimizing high-performance OpenSearch clusters and platforms from the ground up in production environments .

    General Responsibilities

    • Provision, build, deploy, monitor, operate, and support cloud services in a globally distributed team environment
    • Architect, build, deploy, and maintain high-performance OpenSearch clusters and platforms from the ground up
    • Administer and optimize OpenSearch environments for high availability, resiliency, scalability, security, and performance
    • Monitor and troubleshoot cluster health, node performance, indexing throughput, search latency, shard allocation, replication, and storage utilization
    • Analyze and resolve operational issues, platform instability, and production incidents across infrastructure, platform, and application layers
    • Conduct incident response, root cause analysis, and post-incident remediation to drive continuous improvement
    • Maintain the integrity and security of servers, systems, and OpenSearch platform infrastructure
    • Support platform lifecycle activities including installation, configuration, upgrades, patching, hotfixes, backup, restore, and disaster recovery
    • Develop and maintain monitoring policies, alerting standards, operational runbooks, and support procedures
    • Automate testing, deployment, scaling, recovery, and operational workflows for OpenSearch and related cloud services
    • Ensure proper resource allocation and capacity planning across compute, memory, storage, and network resources
    • Partner with product development and engineering teams to design and enhance service reliability and operational readiness
    • Develop and implement testing strategies and document results for platform changes and operational improvements
    • Support log ingestion, index management, retention policies, lifecycle management, and search performance tuning
    • Work in a diverse environment and cross-train with other global team members
    • Participate in an on-call rotation and support weekend or after-hours operational needs as required

    Requirements

    • Expert with Kubernetes , including troubleshooting, operations, management, and configuration of complex Kubernetes services.
    • Proven hands-on expertise designing, building, deploying, supporting, and maintaining OpenSearch clusters and platforms from scratch in production environments
    • Strong experience with OpenSearch administration, cluster architecture, performance tuning, scaling, upgrades, and troubleshooting
    • Experience with index design, shard and replica strategy, cluster sizing, node management, snapshot/restore, backup, and disaster recovery
    • Strong understanding of distributed systems, search platforms, indexing pipelines, query optimization, and high-availability architectures
    • Expertise with Git
    • Expertise with Concourse , including setup, management, and troubleshooting of new pipelines
    • Expertise with Linux , specifically SUSE and Ubuntu
    • Expertise with Kafka, Zookeeper, and Big Data technologies
    • Expert in development of automation for testing, deployment, scalability, and management of cloud services
    • Expertise with building, implementing, and/or supporting cloud monitoring tools
    • Expert knowledge of cloud computing, infrastructure operations, and databases
    • Expert understanding of web services, networking, virtualization, and internet protocols
    • Ability to multitask and handle various projects, deadlines, and changing priorities
    • Excellent communication and prioritization skills
    • Expertise with security fundamentals as they pertain to SaaS multi-tenant application systems
    • Strong interpersonal, presentation, and customer service skills

    Desired Qualifications

    • Experience with AWS services including Route 53, EC2, S3, CloudWatch, DynamoDB, RDS, IAM, ACM, KMS, and VPC
    • Experience deploying and operating OpenSearch in AWS-based environments
    • Experience with Cloud Foundry -based environments
    • Experience with Jenkins , Chef , and/or Terraform
    • Exposure to and understanding of troubleshooting IP networks and application stacks
    • Experience with observability tools such as Prometheus and Grafana
    • Experience with log ingestion pipelines, index lifecycle management, retention strategies, and search platform security controls
    • Familiarity with capacity forecasting, performance benchmarking, and resilience testing for distributed search platforms

    Education

    • BS/BA degree in Computer Science, Management Information Systems, or related IT discipline preferred
    • Allowable substitution: An additional four (4) years of experience may be substituted for a BS/BA degree
    • 8+ years of experience

    Additional Requirements

    • Participation in an on-call rotation for handling P1 incidents is required
    • Flexible schedule which may include weekend or after-hours work
    • Ability to work effectively in a diverse, collaborative, and globally distributed team environment

    Similar Jobs

    See more jobs