Site Reliability Engineer – Network Observability

Team Red Dog

  • Redmond, WA
  • 2 days ago
  • Remote
  • $56 Per Hour

Highlights

This Site Reliability Engineer will primarily support syslog-ng and trapd environments running on Linux, along with Azure-hosted Windows and Linux VMs and observability platforms such as IBM SevOne Network Performance Manager and Broadcom AppNeta. This role will operate and maintain enterprise-scale network observability platforms supporting Azure infrastructure, with a strong focus on Linux administration, syslog-ng, SNMP telemetry, security compliance, patching, and platform reliability.

Numbers & Facts

LocationRedmond, WA (
Remote
)

Description

Team Red Dog is hiring a Site Reliability Engineer – Network Observability for our client, a leading cloud and software provider and intelligent cloud leader. This role will operate and maintain enterprise-scale network observability platforms supporting Azure infrastructure, with a strong focus on Linux administration, syslog-ng, SNMP telemetry, security compliance, patching, and platform reliability. You will automate operational workflows using Ansible, Bash, PowerShell, and Python while troubleshooting complex network and systems issues and supporting highly available observability services at global scale. This is an opportunity to work with large-scale network telemetry and help connect observability data with AI-driven workflows for performance, capacity management, and security.

 

Top Required Skills (Must Haves):

  1. Syslog-ng – 3+ years of hands-on experience operating, configuring, troubleshooting, patching, and maintaining syslog-ng in enterprise Linux environments.
  2. Linux Systems Administration – 5+ years administering Linux systems, including security and OS patching, configuration management, capacity planning, troubleshooting, and production support.
  3. Network Engineering & Observability – 3+ years of enterprise networking experience with strong knowledge of network monitoring and telemetry protocols including SNMP, SNMP Traps, NetFlow, and gNMI.
  4. Automation & Configuration Management – Working knowledge of Ansible and playbooks, with experience automating operational tasks using Bash, PowerShell, or Python.

 

Opportunity Overview:

Join the Service Health Platforms team supporting enterprise network observability systems that provide critical telemetry to network and security engineers operating at global scale. This Site Reliability Engineer will primarily support syslog-ng and trapd environments running on Linux, along with Azure-hosted Windows and Linux VMs and observability platforms such as IBM SevOne Network Performance Manager and Broadcom AppNeta. The role offers exposure to sophisticated cloud and network infrastructure while helping advance the use of network telemetry within AI workflows for performance monitoring, capacity management, and security.

How you will make an impact:

  • Administer and operate Windows and Linux virtual machines hosted in Azure, maintaining compliance with security and configuration standards.
  • Operate and maintain network observability platforms, with a primary focus on syslog-ng and trapd running on Linux.
  • Plan and execute security, operating system, and application patching and upgrades.
  • Maintain observability systems through capacity planning, configuration management, functional audits, and regular maintenance.
  • Author and maintain rules using regular expressions to support telemetry processing and platform operations.
  • Support IBM SevOne Network Performance Manager and Broadcom AppNeta observability platforms.
  • Investigate automated monitoring alerts and customer-reported incidents involving network observability platforms.
  • Troubleshoot complex network observability configurations, applications, operating systems, and infrastructure issues.
  • Assist network and security engineers with identifying traffic patterns, resource utilization, and performance trends.
  • Automate operational and administrative tasks using Bash, PowerShell, Python, and configuration management tools.
  • Deploy, configure, and manage Azure cloud services with an emphasis on scalability, reliability, security, and cost effectiveness.
  • Apply DevOps practices including CI/CD pipelines, infrastructure as code, and source control to streamline deployment and operational processes.
  • Perform business continuity and disaster recovery failover testing as appropriate.
  • Manage assigned projects and program components to deliver services against established objectives and timelines.
  • Participate in the team's on-call DRI rotation and support timely incident resolution.

The expertise you bring:

  • Bachelor's degree in computer science, computer engineering, a related technical field, or equivalent professional experience.
  • 5–7 years of enterprise experience in IT systems, network engineering, site reliability engineering, or a closely related role.
  • 5+ years of Linux systems administration experience.
  • 3+ years of hands-on syslog-ng experience.
  • 3+ years of enterprise network engineering experience.
  • Strong understanding of network observability and telemetry technologies including SNMP, SNMP Traps, NetFlow, and gNMI.
  • Strong knowledge of enterprise networking and routing and switching protocols.
  • Hands-on experience with Microsoft Azure or a comparable cloud platform.
  • Experience with system capacity planning, functional configuration, auditing, and capacity analysis tools.
  • Working knowledge of Ansible and Ansible playbooks.
  • Experience automating tasks with Bash, PowerShell, and/or Python.
  • Proficiency with regular expressions.
  • Experience with IBM SevOne Network Performance Manager, Broadcom AppNeta, or similar enterprise observability platforms is preferred.
  • Experience with source control platforms and DevOps practices.
  • Intermediate knowledge of data retrieval and query languages such as KQL and T-SQL is preferred.
  • Exposure to commercially available AI platforms is preferred.

What makes a candidate highly successful in this role:

A highly successful candidate will bring deep, hands-on experience operating network monitoring and observability systems rather than solely working with general cloud infrastructure. Strong Linux administration skills combined with practical syslog-ng and trapd expertise will be particularly valuable, as will a solid understanding of enterprise routing, switching, and network management protocols. Candidates who can pair this operational depth with Ansible playbooks and scripting automation will be especially well positioned to improve reliability and efficiency across the environment. Experience with IBM SevOne, Broadcom AppNeta, Azure, and large-scale enterprise network telemetry environments will further strengthen a candidate's ability to contribute quickly.

Why Work with Team Red Dog?

At Team Red Dog, people are at the heart of everything we do. Our commitment to personalized service and our deep experience in matching talented professionals with meaningful roles at some of the world’s most inspiring companies is what sets us apart. We take the time to understand your unique skills, strengths, and passions—because we believe your career should reflect who you are.

Whether you're looking to grow, pivot, or simply find a place where your work truly matters, we offer opportunities that empower you to make a positive impact. With excellent benefits, a supportive team, and a role where you can thrive while doing what you love, we’re here to help you take the next step with confidence. Join us—and discover what it means to be genuinely valued in your career.

Generous benefits package for qualified employees includes:

  • Health insurance (medical, dental, vision, and life)
  • Employer-matched 401K plan
  • Generous Paid Time Off
  • Flexible Paid Holiday Benefit

Estimated Start Date: Immediately
Location: Remote – Pacific Time Zone preferred
Job #: 2588
Job Type and Estimated Duration: W2 contract opportunity through August 31, 2027, with the potential to extend up to 18 months total, subject to performance, budget, and client discretion.
Rate: $56–$60/hour

 

Team Red Dog is committed to providing equal opportunities to everyone, regardless of race, ethnicity, gender, age, religion, sexual orientation, disability, or any other characteristic. If you need accommodation during the recruitment process, reach out to hr@teamreddog.com, and we will work to ensure an accessible experience. We strictly adhere to federal, state, and local laws to maintain a workplace free from discrimination and harassment.

We offer competitive compensation aligned with U.S. industry standards, and our final offer will reflect the candidate’s location, job-specific skills, experience, and knowledge.

  • All applicants must be authorized to work in the U.S. without the need for sponsorship.
  • Team Red Dog is an E-Verify employer.
  • Employment is contingent upon the successful completion of a reference and background check.
  • Please no solicitations from C2C or recruiting firms.

 

Similar Jobs

See more jobs