Site Reliability Engineer

vaaridatech

  • Dallas, TX, TX
  • 3 days ago
  • Temporary
  • Contractor
  • Full-time

Highlights

p>Position - Senior/Lead Site Reliability Engineer Observability

Location - 100% Remote

Experience - 8+ Years

Type - Full Time

Technology Stack - Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.

Job Description -

Must Have Technical/Functional Skills:

• 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.

• Hands-on experience administering Splunk Enterprise or Splunk Cloud.

• Strong knowledge of Splunk SPL.

• Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.

• Experience implementing metrics, logs, and traces as part of a modern observability strategy.

• Experience with Terraform and Infrastructure as Code.

• Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and service mesh technologies.

• Experience supporting FedRAMP or regulated environments.

Numbers & Facts

LocationDallas, TX, TX
Job TypeTemporary, Contractor, Full-time

Description

Position - Senior/Lead Site Reliability Engineer Observability

Location - 100% Remote

Experience - 8+ Years

Type - Full Time

 

Technology Stack - Splunk Enterprise, Splunk Cloud, Elasticsearch, ELK, Kibana, Prometheus, Grafana, Grafana Tempo, OpenTelemetry, Distributed Tracing, Kafka, Terraform, Kubernetes, Docker, Linux, Python, Go, Ruby, Bash, AWS, Ansible, Consul.

 

Job Description -

Must Have Technical/Functional Skills:

• 7+ years in Site Reliability Engineering, Platform Engineering, or DevOps.

• Hands-on experience administering Splunk Enterprise or Splunk Cloud.

• Strong knowledge of Splunk SPL.

• Experience with Elasticsearch/ELK, Prometheus, Grafana, Grafana Tempo, distributed tracing, OpenTelemetry, and Kafka.

• Experience implementing metrics, logs, and traces as part of a modern observability strategy.

• Experience with Terraform and Infrastructure as Code.

• Programming experience in Python, Go, Ruby, or Bash. 

• Splunk certification.

• Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and servicemesh technologies.

• Experience supporting FedRAMP or regulated environments. 

 

Roles & Responsibilities:

• Design, deploy, and operate enterprise observability platforms.

• Build and maintain Splunk Enterprise/Splunk Cloud infrastructure including Indexers, SearchHead Clusters, Heavy Forwarders, and Deployment Servers.

• Deploy and operate large-scale Elasticsearch clusters for log analytics and search.

• Design, deploy, and support distributed tracing platforms using Grafana Tempo and OpenTelemetry.

• Build and maintain end-to-end tracing pipelines, instrumentation standards, and trace retention strategies.

• Scale Prometheus, Grafana, Kafka, Tempo, and OpenTelemetry-based monitoring solutions.

• Develop dashboards, alerts, analytics, and trace visualizations using Splunk SPL, Grafana, Kibana, and Tempo.

• Automate infrastructure using Terraform and configuration management tools. 

 

Nice to have skills:

• Splunk certification.

• Experience with Kubernetes, AWS/Azure/GCP, Ansible, Consul, CI/CD pipelines, and service mesh technologies.

• Experience supporting FedRAMP or regulated environments.

Similar Jobs

See more jobs