Sr DevOps Engineer

HCL Global Systems Inc.

  • Denver, CO
  • 30+ days ago

    Highlights

    The role requires deep expertise in deployment automation, HPA tuning, release reliability, networking, observability, and rollback strategies for large-scale microservices platforms handling high throughput workloads. Do you have deep expertise in deployment automation, HPA tuning, release reliability, networking, observability, and rollback strategies for large-scale microservices platforms handling high throughput workloads?.

    Numbers & Facts

    LocationDenver, CO

    Description

    Sr DevOps Engineer
    Local to Denver only - 4-5 days a week at client site office

    Request for 4 Sr. Devops Engineers Local to Denver only - 4-5 days a week at client site office

    Dev Ops
    Amazon EKS
    ISTIO-Service Mesh
    Tuning CPU / Memory
    ArgoCD & Rollouts
    Kubernetes

    Are you local to Denver, CO?
    Are you able and willing to be at the client site office location 4-5 days a week?
    Do you have strong experience in AWS, Kubernetes (EKS), ArgoCD, autoscaling, and cloud-native deployments?
    Do you have deep expertise in deployment automation, HPA tuning, release reliability, networking, observability, and rollback strategies for large-scale microservices platforms handling high throughput workloads?
    Are you able to manage and optimize Amazon EKS clusters?
    Do you have strong experience with Istio service mesh?


    We are looking for a Senior DevOps Engineer with strong experience in AWS, Kubernetes (EKS), ArgoCD, autoscaling, and cloud-native deployments.
    The role requires deep expertise in deployment automation, HPA tuning, release reliability, networking, observability, and rollback strategies for large-scale microservices platforms handling high throughput workloads. This role will support mission-critical systems requiring safe deployments, dynamic scaling, performance tuning, and production stability.

    Key Responsibilities
    Manage and optimize Amazon EKS clusters
    Configure and tune:
    Horizontal Pod Autoscaler (HPA)
    Cluster Autoscaler
    Resource limits / requests
    Manage Istio service mesh
    Traffic routing
    Circuit breaker
    Retry / timeout policies
    VirtualService / DestinationRule
    Gateway configuration
    Implement safe deployment strategies:
    Blue-Green
    Canary
    Rolling update
    Traffic shifting using Istio / Argo Rollouts
    Automated rollback
    Automate infrastructure using:
    Terraform / CDK / CloudFormation
    Bash / Python / Shell scripting
    Manage AWS services:
    EKS
    DynamoDB
    S3
    IAM
    Route53 / DNS
    Lambda
    CloudWatch
    SNS / SQS / EventBridge
    Configure logging, monitoring, and alerting
    Implement release readiness validation checks
    Support production deployments and rollback
    Optimize infrastructure for:
    Performance
    Cost
    Scalability
    Implement automation for:
    Cluster health checks
    Deployment validation
    Environment readiness
    Rollback triggers
    Work with development teams to support:
    Microservices deployments
    Containerization
    Performance tuning
    Failure recovery
    Required Skills
    Kubernetes / EKS / Istio
    Strong experience with Amazon EKS
    Strong experience with Istio service mesh
    Experience configuring:
    VirtualService
    DestinationRule
    Gateway
    Sidecar
    Circuit breaker
    Retry / timeout
    Experience with Horizontal Pod Autoscaler (HPA)
    Experience with Cluster Autoscaler
    Experience with Helm / Kustomize
    Experience with ArgoCD / Argo Rollouts / GitOps
    Deployment & Release
    Blue-Green deployments
    Canary deployments
    Traffic shifting using Istio
    Rollback strategies
    GitOps workflows
    Release automation
    AWS Services
    DynamoDB
    S3
    IAM
    Route53 / DNS
    CloudWatch
    Lambda
    ELB / ALB / NLB
    Automation / Scripting
    Bash
    Python
    Shell scripting
    YAML / JSON
    Terraform / CDK / CloudFormation
    Networking
    TCP/IP fundamentals
    DNS
    VPC / Subnets
    Security groups
    Load balancers
    TLS / SSL certificates
    Service mesh networking concepts
    Observability / Logging
    Datadog / Splunk / Prometheus / Grafana
    Metrics dashboards
    Alerting setup
    Log aggregation
    Production Support
    Deployment troubleshooting
    Rollback handling
    Incident response
    Capacity planning
    Performance tuning
    Preferred / Nice to Have
    Kafka / streaming systems
    High TPS / real-time systems
    DynamoDB capacity tuning
    Cost optimization in AWS
    Multi-cluster / multi-region EKS
    Release automation tooling
    Service mesh performance tuning

    Similar Jobs

    See more jobs