We are seeking a mid-level Infrastructure Engineer (5+ years of relevant experience) to support the build-out of Flexential's enterprise Observability platform. This is a hands-on technical role working alongside Platform Engineering to provision, configure, and maintain the infrastructure underpinning a modern LGTM-based observability stack across 40+ data center facilities.
Responsibilities
Provision and manage virtual infrastructure within VMware VCF/vSphere/vCenter environments supporting the Observability platform • Build and maintain Terraform modules for infrastructure provisioning including VMs, storage, and networking • Develop and maintain Ansible playbooks for OS configuration, service deployment, and environment hardening • Design, deploy, and manage Kubernetes clusters (RKE2 or equivalent) hosting the LGTM stack workloads • Manage Helm chart deployments for Grafana, Mimir, Loki, Tempo, and Prometheus Operator • Build and maintain GitLab CI/CD pipelines and toolchain supporting end-to-end GitOps workflows • Configure and maintain network connectivity for Observability project to all infrastructure and application components including virtualization, LAN, and WAN • Collaborate with the Platform Engineering team to establish standards for Infrastructure-as-Code (IaC), environment promotion, and operational runbooks
Required Skills • Python and YAML — 3+ years of automation scripting and configuration authoring • Terraform — 3+ years writing and maintaining infrastructure modules; experience with vSphere or cloud providers preferred • Ansible — 2+ years of playbook development; AWX/automation controller experience preferred • Kubernetes and Helm — 3+ years of cluster administration and workload deployment, including day-2 operations • GitLab — 2+ years of CI/CD pipeline development and GitOps workflow management • VMware VCF, vSphere, and vCenter — 2+ years of hands-on experience required; strong working knowledge of VM lifecycle, resource pools, and storage policies • Linux — 3+ years of admin-level proficiency required; systemd, networking, storage, and security hardening • Networking — 2+ years working with LAN/WAN, DNS, load balancing, and network virtualization in an enterprise environment
Preferred Skills • Any hands-on exposure to LGTM stack technologies (Loki, Grafana, Tempo, Mimir) or OpenTelemetry • Familiarity with AIOps concepts and observability platform operations nincluding ITSM Major incident management workflows and tools • Experience with secrets management tools such as CyberArk, Conjur, or HashiCorp Vault • Experience with Postgres or other RDBMS