SRE Lead - Messaging Services

Diverse Lynx, LLC

Chandler, AZ

JOB DETAILS
SKILLS
Apache Kafka, Applications Security, Artificial Intelligence (AI), Automation, Best Practices, Brokerage, Budgeting, Cryptography, Distributed Computing, Failover, Financial Services, High Availability, High Reliability, IBM WebSphere MQ (Message Queue), Identify Issues, Incident Management, Large-Scale Systems, Leadership, Linux Operating System, Messaging Middleware, Messaging Technology, Microsoft Windows Operating System, On Call, Production Support, Production Systems, Python Programming/Scripting Language, Realtime Operating System, Reliability Engineering, Risk Management, Root Cause Analysis, SSL-TLS (Secure Socket Layer - Transport Layer Security), Software Patches, Splunk, Systems Administration/Management, Systems Reliability, Systems Scalability, Unix Operating Systems, Unix Shell Programming
LOCATION
Chandler, AZ
POSTED
8 days ago

Primary Skills: Site Reliability Engineering / Production Engineering, IBM MQ, Kafka

Secondary Skills: Observability (Splunk/Dynatrace), Linux/Unix, Python/Shell Automation, Incident & Problem Management, Messaging Platform

SUMMARY:

We are seeking an experienced Site Reliability Engineer (SRE) Lead - Messaging Services to drive platform reliability, observability, and operational excellence across IBM MQ and Kafka environments.

This role combines:

  • Production engineering and reliability leadership for messaging platforms
  • Platform security, resilience engineering, and vulnerability remediation
  • Ownership of large-scale, distributed messaging runtimes

Key responsibilities include:

  • Leading reliability engineering for high-scale messaging platforms supporting tens of thousands of runtimes and high-volume message throughput
  • Driving EOL remediation, patching, and stabilization across MQ queue managers and Kafka clusters
  • Implementing SRE best practices:

o SLIs / SLOs focused on message delivery, latency, and availability

o Incident management, escalation, and postmortem culture

  • Enhancing observability and monitoring for messaging flows, queue depths, lag, and throughput
  • Designing proactive fault detection and auto-remediation strategies (e.g., DLQ handling, backlog mitigation, failover recovery)
  • Building resilient messaging platforms capable of supporting real-time, event-driven workloads
  • Supporting global production messaging environments with on-call rotation and escalation ownership
  • Partnering with engineering, application, and security teams to ensure reliability, scalability, and secure message transport

Required Skill

  • Strong experience in Site Reliability Engineering / Production Engineering
  • Hands-on expertise with:

o IBM MQ (queue managers, clustering, channels, DLQ management)

o Kafka / Confluent platform (topics, brokers, partitions, consumer groups)

o Large-scale distributed messaging systems and runtime management

  • Deep understanding of:

o System reliability, scalability, and high availability design

o Messaging reliability patterns (guaranteed delivery, retry handling, replay, ordering)

o Incident management, root cause analysis, and problem management

  • Experience with:

o Observability tools (Dynatrace, Splunk, Prometheus, Grafana) for messaging platforms

o Event and anomaly detection in high-volume systems

  • Strong scripting/automation skills:

o Shell, Python, PowerShell

  • Experience managing Linux/Unix and Windows production environments
  • Knowledge of:

o Event-driven architecture and messaging-based integration patterns

  • Understanding of:

o Messaging platform security (TLS, certificates, channel auth, encryption)

o Vulnerability remediation and risk mitigation in production systems

  • Excellent troubleshooting skills in high-pressure, real-time environments (e.g., message backlog, latency spikes, connection failures

Desired Skill

  • Experience implementing SRE frameworks (SLIs, SLOs, error budgets) specifically for messaging workloads
  • Familiarity with:

oKubernetes / containerized messaging platforms

  • Experience with:

o Kafka ecosystem components (Schema Registry, Connect, Streams)

o IBM MQ advanced features (Native HA, clustering)

  • Exposure to:

o AI-driven operations (AIOps), anomaly detection, or automated remediation

o Large-scale messaging modernization or migration programs

  • Messaging or middleware certifications (IBM MQ, Kafka, or equivalent)
  • Experience in regulated environments (e.g., financial services)

Diverse Lynx LLC is an Equal Employment Opportunity employer. All qualified applicants will receive due consideration for employment without any discrimination. All applicants will be evaluated solely on the basis of their ability, competence and their proven capability to perform the functions outlined in the corresponding role. We promote and support a diverse workforce across all levels in the company.

About the Company

D

Diverse Lynx, LLC