Sr AWS DevOps Engineer

Abode Techzone LLC

  • Scottsdale, AZ
  • 2 days ago

    Highlights

    Lead the design, implementation, and continuous improvement of DevOps best practices, including Continuous Integration (CI), Continuous Deployment (CD), automated testing, and Test-Driven Development (TDD). • Design and implement automated backup and rollback strategies, achieving a Recovery Time Objective (RTO) of under 15 minutes, a Recovery Point Objective (RPO) of under 5 minutes, and a 30-day backup retention policy.

    Numbers & Facts

    LocationScottsdale, AZ

    Description

    Job Title: Sr AWS DevOps Automation Engineer

    Location: Scottsdale, AZ

    Project Duration: 12 months Contract

    Job Type: Contract

    Work Arrangement: Work from Client Office – 3 days a week

    Interview: Three rounds of Video interview

     

    Key Responsibilities

    DevOps Strategy & Automation

    • Lead the design, implementation, and continuous improvement of DevOps best practices, including Continuous Integration (CI), Continuous Deployment (CD), automated testing, and Test-Driven Development (TDD).

    • Design and maintain highly resilient, scalable, secure, and software-defined infrastructure platforms supporting telecom-grade availability and performance requirements.

    • Automate operational processes to improve deployment speed, consistency, and platform reliability.

    • Drive continuous improvement initiatives across infrastructure, deployments, monitoring, and incident response processes.

    CI/CD Pipeline Architecture

    • Design and implement end-to-end automated deployment pipelines using GitHub, GitHub Actions, and Microsoft SQL Server (MSSQL).

    • Build and manage environment promotion workflows across Development, Staging, and Production environments with minimal to no manual intervention.

    • Optimize deployment processes to enable rapid and reliable software releases.

    Self-Hosted Runner Management

    • Install, configure, and maintain Windows Server-based GitHub Actions self-hosted runners.

    • Ensure high availability (99.9% uptime), secure network connectivity, performance monitoring, and operational health checks.

    PowerShell Automation & Release Engineering

    • Develop and maintain production-grade PowerShell automation scripts for:

    Deployment

    Backup

    Rollback

    Health checks

    Validation

    • Implement robust error handling, transaction management, auditing, and logging capabilities.

    Jira, GitHub & Workflow Automation

    • Integrate GitHub repositories with Jira to automatically link code branches and deployments to Jira tickets.

    • Automate ticket lifecycle management, including status updates throughout the deployment pipeline.

    • Enrich Jira tickets with deployment logs, build data, commit hashes, and deployment URLs.

    Cloud & Database Technologies

    • Design and support cloud-native solutions leveraging AWS, DynamoDB, Amazon Aurora, and Amazon Kinesis.

    • Support database deployment automation and governance for enterprise applications.

    Observability & Monitoring

    • Implement and maintain monitoring, observability, and alerting solutions using Prometheus, Grafana, Kubernetes Monitoring, OpenTelemetry, Splunk, Zabbix, and Dynatrace.

    • Proactively monitor production systems and rapidly troubleshoot performance, infrastructure, and deployment issues.

    Notifications & Incident Response

    • Configure real-time deployment notifications via Slack and email.

    • Integrate critical alerts with PagerDuty for rapid incident response and escalation management.

    • Participate in on-call support rotations and production incident management.

    Security & Governance

    • Implement DevOps security best practices, including branch protection policies, environment approval gates, GitHub Secrets management, GPG-signed commits, and role-based access controls.

    • Ensure compliance with enterprise security and audit requirements.

    Monitoring, Logging & Audit Trail

    • Develop deployment auditing mechanisms, including deployment tracking tables within MSSQL.

    • Build dashboards and reports to monitor deployment frequency, success rates, failure trends, and operational KPIs.

    • Maintain comprehensive documentation for infrastructure, automation solutions, and deployment processes.

    Disaster Recovery & Business Continuity

    • Design and implement automated backup and rollback strategies, achieving a Recovery Time Objective (RTO) of under 15 minutes, a Recovery Point Objective (RPO) of under 5 minutes, and a 30-day backup retention policy.

    • Develop disaster recovery runbooks and automate restoration procedures.

    Performance Optimization & Operational Excellence

    • Optimize deployment processes to reduce deployment times from approximately 30 minutes to under 5 minutes through automation and parallel execution strategies.

    • Improve storage efficiency by implementing incremental backup solutions and reducing backup costs.

    • Document operational runbooks, conduct knowledge transfer sessions, and lead continuous improvement reviews.

    Required Qualifications

    • Bachelor's degree in Computer Science, Engineering, Information Technology, or a related technical field; equivalent industry experience will also be considered.

    • 10+ years of hands-on experience in DevOps, Infrastructure Automation, Site Reliability Engineering (SRE), or Platform Engineering.

    • Proven expertise designing, implementing, and managing enterprise-scale CI/CD pipelines.

    • Strong scripting and automation experience with PowerShell, Python, Bash, and Groovy.

    • Experience with configuration management tools such as Ansible, Puppet, or Chef.

    • Hands-on experience with GitHub and GitHub Actions, Windows Server administration, Microsoft SQL Server, AWS cloud services, Docker, and Kubernetes.

    • Strong knowledge of monitoring, observability, and logging platforms.

    • Experience integrating DevOps solutions with Jira, Slack, PagerDuty, and other enterprise collaboration tools.

    • Understanding of database deployment automation and release management methodologies.

    • Experience supporting highly available production environments with strict uptime requirements.

    • Excellent troubleshooting, analytical, and problem-solving skills.

    • Strong communication and stakeholder management abilities.

    Preferred Qualifications

    • Experience within the telecommunications, wireless networks, or network infrastructure domain.

    • Exposure to large-scale enterprise deployment and automation frameworks.

    • Knowledge of Infrastructure as Code (Terraform, CloudFormation, or similar technologies).

    • Familiarity with SRE principles and platform engineering practices.

    • Relevant cloud or DevOps certifications (AWS, Kubernetes, Azure, GitHub, HashiCorp Terraform, etc.).

    Success Metrics

    The successful candidate will help the organization achieve:

    • Less than 2% deployment failure rate.

    • Automated deployment recovery and rollback capabilities.

    • Deployment cycle time reduced to under 5 minutes.

    • 99.9% platform and deployment infrastructure availability.

    • Fully auditable deployment processes and compliance controls.

    • Independent team deployment capability through automation and self-service tooling.

     

     

    Similar Jobs

    See more jobs