Site Reliability Engineer (SRE) - AI & Automation

Iconma LLC

  • Dallas, TX
  • 20 days ago

    Highlights

    Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies. API & Microservices Engineering - Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.

    Numbers & Facts

    LocationDallas, TX

    Description

    Our Client, an IT Services and Consultant company, is looking for a Site Reliability Engineer (SRE) - AI & Automation for their Dallas, TX / Scottsdale, AZ / Hybrid location.

    Responsibilities:

    • Build automation and operational tools using Java, Python, and Node.js to improve efficiency, scalability, and platform operations.
    • Leverage AI and Generative AI technologies (Gemini, Llama, Mistral, Qwen, etc.) to automate alert analysis, incident response, operational workflows, and runbook execution.
    • Implement API and microservices reliability solutions using Apigee/Apigee X, REST APIs, GraphQL gateways, traffic routing, canary deployments, and failover strategies.
    • Manage Kubernetes platforms across GKE and Rancher RKE2, including cluster administration, performance tuning, and troubleshooting.
    • Ensure platform reliability and high availability by supporting active-active deployments, disaster recovery readiness, and multi-datacenter Kubernetes environments.
    • Develop observability and monitoring capabilities using tools such as Splunk, Grafana, Datadog, and AppDynamics to meet reliability and performance objectives.
    • Drive SRE best practices and operational excellence by partnering with cross-functional teams to improve reliability, security, incident management, and continuous improvement.

    Requirements:

    • Site Reliability Engineering (SRE) - Reliability, availability, incident management, SLO/SLI monitoring, and operational excellence.
    • Kubernetes Platform Engineering - 5+ years of Strong hands-on experience with GKE and Rancher RKE2, multi-cluster management, troubleshooting, and performance optimization.
    • Cloud & Infrastructure Automation - Strong experience in GCP, Terraform, Helm, GitHub, CI/CD, and production-grade automation.
    • Software Development - 5+ years of Advanced programming skills in Python and Java (Node.js preferred for integrations and automation workflows).
    • Observability & Monitoring - Splunk, Grafana, Datadog, AppDynamics, alerting, and platform health monitoring.
    • API & Microservices Engineering - Apigee/Apigee X, REST APIs, GraphQL, traffic routing, canary deployments, and failover strategies.
    • AI-Driven Operations (AIOps) - Applying LLMs such as Gemini, Llama, Mistral, and Qwen for alert analysis, incident triage, automation, and operational workflows.
    • Years of Experience: 14.00 Years of Experience

    Why Should You Apply?

    • Health Benefits
    • Referral Program
    • Excellent growth and advancement opportunities

    Similar Jobs