DevOps Engineer (Operations/SRE)
Location: Santa Clara Valley, California, United States
Minimum Experience: 6+ years
SUMMARY
We are seeking an experienced DevOps/SRE Engineer to support the availability, performance, and
automation of mission-critical enterprise platforms and infrastructure. You will work across cloud
environments, routing infrastructure, database systems, and internal tooling — contributing to operational
excellence for high-volume, globally distributed systems.
This role requires hands-on technical execution: building automation, managing infrastructure-as-code,
supporting production deployments, and collaborating closely with application engineers, DBAs, and crossfunctional
teams.
KEY QUALIFICATIONS
Core (Required)
Cloud — Solid working knowledge of cloud platforms (compute, storage, networking, IAM, serverless,
container services); ability to architect, deploy, and troubleshoot cloud-hosted workloads
Linux — Filesystem, process management, user/permission management, systemd, package
management, performance troubleshooting
DevOps & CI/CD — CI/CD pipelines (Jenkins, GitHub Actions, or similar); release management and
deployment automation
Networking — DNS, TCP/IP, load balancing, firewalls, VPNs, proxies, routing
Scripting — Proficient in Bash and/or Python; production-quality automation scripts for operational
workflows
Configuration Management — Hands-on Ansible (playbooks, roles, inventories, vault); infrastructure
automation at scale
Containers & Orchestration — Docker and Kubernetes basics
Database — Experience with relational and NoSQL databases — Oracle, Cassandra, PostgreSQL, or
similar
SRE Practices — Capacity planning, failover procedures, backup/recovery, and monitoring
Monitoring & Observability — Splunk for log analysis, alerting, and operational troubleshooting
Reverse Proxy — nginx configuration, traffic management, and performance tuning
Problem Solving — Strong troubleshooting and incident diagnosis in production environments
Preferred (Strong Plus)
Generative AI — Experience leveraging AI tools (LLMs, code assistants, prompt engineering) to
accelerate automation, troubleshooting, and documentation
Java Development — Experience writing or maintaining Java/J2EE applications; familiarity with Spring,
React, or similar frameworks
Familiarity with additional observability tools (Datadog, Prometheus, Grafana, or similar)
Infrastructure-as-code tools (Terraform, CloudFormation)
Security fundamentals — secrets management, access control, certificate management
Prior experience in large-scale enterprise IT operations or platform engineering environments
RESPONSIBILITIES
Build and maintain automation for infrastructure provisioning, deployments, and operational runbooks
Manage and optimize cloud environments — compute, networking, storage, and security configurations
Develop and maintain Ansible playbooks for configuration management across multiple environments
Support production systems through incident triage, root cause analysis, and on-call participation
Collaborate with DBAs on database infrastructure — capacity testing, failover readiness, and monitoring
Implement and improve CI/CD pipelines for application and infrastructure deployments
Participate in change management and release coordination processes
Manage service accounts, access provisioning, and security configurations
Maintain routing and proxy infrastructure — configuration updates, performance tuning, troubleshooting
Document operational procedures, runbooks, and architecture decisions
Work cross-functionally with application engineers, security teams, and platform owners
EDUCATION & EXPERIENCE
BS/MS in Computer Science, Engineering, or equivalent practical experience
5+ years in a DevOps, SRE, Systems Engineering, or Cloud Engineering role
Demonstrated experience operating in cloud environments
Track record of delivering automation that reduced manual operational work
Experience supporting critical production systems with 24x7 availability requirements
ADDITIONAL INFORMATION
Comfortable working independently while coordinating with multiple teams
Strong written communication skills — ability to document processes and contribute to team knowledge
bases
Familiarity with Agile/Scrum workflows is a plus