Design and support resilient multi-region AWS architectures and containerized deployments using Docker and Amazon ECR. Own reliability measures including SLIs, SLOs, SLAs, MTTR improvement, and service reliability KPIs.
Numbers & Facts
Location
Charlotte, NC
Salary
$70,000–$120,000 Per Year
Description
Must Have Technical/Functional Skills
SRE / Application Reliability Engineering (ARE) and 24x7 production operations
Major Incident Management (P1/P2), ServiceNow, and Incident / Problem / Change Management
SLI, SLO, SLA governance; MTTR reduction; service reliability KPIs
AWS: CloudWatch, Route 53, S3, CloudFront, Lambda, ECR, and Bedrock
Observability: Dynatrace, Grafana, Splunk, and CloudWatch
Control-M and enterprise batch operations
GitLab CI/CD, REST APIs, Personal Access Tokens, security scanning, and DevSecOps
Python GitLab library and API-based automation
Automation, AIOps, event correlation, self-healing, and stakeholder management
Roles & Responsibilities
Lead 24x7 SRE operations and coordinate P1/P2 major incident response through service restoration and follow-up.
Own reliability measures including SLIs, SLOs, SLAs, MTTR improvement, and service reliability KPIs.
Drive end-to-end observability using Dynatrace, Grafana, CloudWatch, and Splunk.
Lead automation, AIOps, self-healing, event correlation, and AI-driven operations initiatives.
Oversee AWS platform operations, batch processing, and Control-M environments.
Integrate Claude AI on AWS Bedrock with GitLab using APIs, PATs, and custom workflows.
Develop AI-driven analysis of GitLab project data, vulnerabilities, pipelines, and security findings.
Design GitLab API automation, custom workflows, DevSecOps controls, and CI/CD pipeline improvements.
Build and maintain Python-based GitLab integrations and REST API solutions.
Support vulnerability remediation and onboarding/configuration of security scanning tools.
Manage AWS Lambda, ECR, and Bedrock for deployment and automation; optimize Lambda configuration, concurrency, and scaling.
Design and support resilient multi-region AWS architectures and containerized deployments using Docker and Amazon ECR.
Generic Managerial Skills
Lead geographically distributed operations teams and coordinate effectively during critical incidents.
Communicate reliability risks, service performance, and remediation plans to technical and business stakeholders.
Drive governance, prioritization, continuous improvement, and cross-team collaboration.
Mentor engineers and promote automation-first, blameless, and reliability-focused ways of working.