Senior Site Reliability Engineer

EPAM Systems Inc

  • 30+ days ago

    Highlights

    Working closely with AI and full-stack engineers, you will shape how the team builds, ships, and operates, with particular focus on the reliability challenges unique to LLM workloads: cost, latency, and non-deterministic failure modes. Build and maintain CI/CD pipelines (GitHub Actions) for automated build, test and deployment across multiple microservices, including Docker image management, registries and deployment config.

    Numbers & Facts

    Location

    Description

    Back to Search

    Senior Site Reliability Engineer

    Remote in Türkiye

    DevOps& 6 others

    apply

    FacebookLinkedInSend via email

    Looking for something else?

    Find a vacancy that works for you. Send us your CV to receive a personalized offer.

    Find me a job

    We are looking for a Senior Site Reliability Engineer ready to own the reliability and operational maturity of a production AI platform. You will be the engineering foundation that keeps agentic content workflows running at scale, ensuring services are observable, deployments are automated, and infrastructure is reproducible. Working closely with AI and full-stack engineers, you will shape how the team builds, ships, and operates, with particular focus on the reliability challenges unique to LLM workloads: cost, latency, and non-deterministic failure modes. This is a role for someone who takes pride in building systems that others depend on.

    Responsibilities

    • Own the reliability, scalability and performance of the platforms services running on Azure Container Apps and AWS ECS

    • Build and maintain CI/CD pipelines (GitHub Actions) for automated build, test and deployment across multiple microservices, including Docker image management, registries and deployment config

    • Implement and manage infrastructure as code (Terraform, Bicep or ARM) across Azure and AWS

    • Set up and maintain observability - monitoring, alerting, logging and dashboards (New Relic, Langfuse, CloudWatch)

    • Manage Azure Service Bus, Blob Storage, Key Vault and Container Apps configurations

    • Ensure security best practices - secret management, image scanning and vulnerability remediation

    • Implement auto-scaling, load balancing and cost optimization for AI workloads

    • Support incident response and establish runbooks for production services

    • Collaborate with AI engineers to optimize LLM API usage, token costs and latency

    Requirements

    • 3-5+ years of SRE, DevOps or platform engineering experience

    • Strong hands-on expertise in Azure (Container Apps, Service Bus, Key Vault, Blob Storage) and Azure OpenAI resource management

    • Familiarity with AWS (ECS, S3, Aurora, CloudWatch)

    • Proficiency in infrastructure as code using Terraform, Bicep or ARM templates

    • Background in CI/CD pipeline design and maintenance (GitHub Actions preferred)

    • Knowledge of Docker and container orchestration, with Kubernetes experience as a strong plus

    • Skills in monitoring and observability with New Relic or equivalent

    • Competency in security practices including secret management, vulnerability scanning and image hardening

    • Capability to script in Python or Bash for automation

    • Daily user of AI development tools (Cursor, Claude Code, Copilot)

    • English: B2+

    Nice to have

    • Familiarity with LLM observability tooling

    • Understanding of Kubernetes in production environments

    Similar Jobs