Job Description: Senior DevOps Engineer Kubernetes, Kafka & Python Automation
Location: Massachusetts- Onsite
Experience: 7+ years
Job Summary
We are seeking an experienced Senior DevOps Engineer with strong expertise in Kubernetes, Apache Kafka, Python automation, and CI/CD. The ideal candidate will be responsible for designing, implementing, automating, and supporting highly available cloud-native infrastructure and event-driven platforms.
The candidate must have strong hands-on experience with Kubernetes administration and troubleshooting, Kafka platform operations, and Python-based automation. This role requires close collaboration with development, infrastructure, and application teams to improve deployment velocity, platform reliability, scalability, and operational efficiency.
Work from the office is mandatory 4 days per week. Candidates local to Massachusetts are strongly preferred.
Key Responsibilities
- Design, deploy, configure, and manage Kubernetes clusters in production environments.
- Perform Kubernetes administration, including deployments, services, ingress, namespaces, RBAC, ConfigMaps, Secrets, StatefulSets, persistent volumes, and autoscaling.
- Troubleshoot complex Kubernetes issues involving pods, nodes, networking, resource utilization, deployments, and application failures.
- Manage and support Apache Kafka infrastructure, including brokers, topics, partitions, replication, consumer groups, and Kafka Connect.
- Monitor and troubleshoot Kafka consumer lag, broker failures, under-replicated partitions, throughput, connectivity, and performance issues.
- Develop Python automation scripts to automate infrastructure provisioning, application deployment, monitoring, health checks, operational tasks, and remediation.
- Build and maintain robust CI/CD pipelines using Jenkins, GitLab CI/CD, GitHub Actions, or similar tools.
- Automate infrastructure provisioning and configuration using Terraform and Ansible.
- Containerize applications using Docker and deploy workloads using Kubernetes and Helm.
- Implement monitoring, alerting, and observability using tools such as Prometheus, Grafana, Datadog, Splunk, or ELK.
- Develop dashboards and alerts for Kubernetes cluster health, Kafka performance, application availability, and infrastructure metrics.
- Implement secure Kafka environments using SSL/TLS, SASL, ACLs, authentication, and authorization.
- Participate in production incident management, troubleshooting, root-cause analysis, and preventive remediation.
- Support Kubernetes upgrades, application releases, rollback procedures, disaster recovery, and capacity planning.
- Work closely with software engineers to optimize applications for cloud-native and event-driven architectures.
- Establish DevOps best practices around automation, infrastructure-as-code, release management, monitoring, and operational reliability.