Tools support Engineer

Vantage Point Consulting

Local to Seattle, Alpharetta, or Cincinnati, GA

JOB DETAILS
SKILLS
Administrative Management, Algorithms, Amazon Web Services (AWS), Application Integration, Application Programming Interface (API), Artificial Intelligence (AI), Asset Management, Automation, Configuration Management, Cross-Functional, DHCP (Dynamic Host Configuration Protocol), DNS (Domain Name System), Database Backup, Database Technology, Ecosystems, Event Correlation, GCP (Good Clinical Practices), High Availability, Hybrid Cloud, IT Service Management (ITSM), Identify Issues, Incident Management, Linux Administration, Machine Learning, Metrics, Microsoft SQL Server, Microsoft Windows Azure, Microsoft Windows System Administration, Multiplatform/Cross-Platform, Netflow, Network Configuration Management, Network Routing, Network Switching, Network Topology, Onboarding, Operational Support, Operations Processes, Pattern Matching, Predictive Modeling, REST (Representational State Transfer), Relationship Management, Reporting Dashboards, Root Cause Analysis, SNMP (Simple Network Management Protocol), SQL Databases, Schedule Development, Splunk, Startup, Systems Administration/Management, TCP/IP (Transmission Control Protocol/Internet Protocol), Technical Support, Telemetry, Topology, WMI API (Windows Management Instrumentation)
LOCATION
Local to Seattle, Alpharetta, or Cincinnati, GA
POSTED
5 days ago

Overview: 

Observability & Enterprise Monitoring Engineer with specialized expertise in SolarWinds platform administration and broader multi-tool observability ecosystems. Working knowledge of OpenText NNMi will be an added advantage. This role will be responsible for the end-to-end administration, optimization, integration, and operational maintenance of enterprise-scale implementation of monitoring solutions (SolarWinds). Responsible for ensuring platform health, automate alert workflows, manage hybrid/cloud monitoring integrations, and collaborate closely with cross-functional infrastructure teams to maintain high availability and performance. 

 

Roles & Responsibilities: 

  1. Platform Administration & Lifecycle Management (SolarWinds) 

  • Core Module Management: Administer and optimize SolarWinds modules including NPM, NCM, NTA, SAM, and the broader Orion / SWOSH (Hybrid Cloud Observability) platform ecosystem. 

  • Upgrades & Maintenance: Perform routine and major version updates across platform components; monitor platform health using Active Diagnostics and My Deployment health checks. 

  • Polling Infrastructure: Manage, scale, and load-balance Additional Polling Engines (APEs) to ensure optimal performance across enterprise environments. 

  • Database & Backup Operations: Perform operational tasks on the underlying MS SQL Database, manage, schedule, and verify configuration and database backup jobs. 

 

  1. Network & Device Observability Operations 

  • Discovery & Asset Management: Execute network discoveries, manage node onboarding/offboarding, assign Universal Device Pollers (UnDP), and maintain custom custom attributes and group hierarchies. 

  • Configuration Management (NCM): Build and maintain NCM command templates, automate daily startup/running config backups, archive config files, and remediate compliance/transfer failures. 

  • Topology & Visualization: Create and maintain dynamic, accurate network topology maps using Network Atlas and modern visual canvases based on operational requirements. 

 

  1. Alerting, Dashboarding & ITSM Integration 

  • Signal Optimization: Design, tune, and maintain custom Alert Triggers, Actions, and Thresholds to eliminate alert noise and drive actionable alerting. 

  • Ticketing & Automation: Configure bi-directional ITSM/ticketing integrations to enable automatic ticket creation, routing, and lifecycle tracking. 

  • Reporting & Visibility: Build custom operational and executive Dashboards, Views, and Reports tailored to stakeholder requirements. 

  • Incident Support: Monitor alert channels for operational anomalies, troubleshoot lingering telemetry issues, and collaborate with domain teams to drive root cause resolution. 

 

  1. AIOps Operations 

  • Leverage AIOps, machine learning, and pattern-recognition capabilities to identify baseline anomalies, reduce event noise, and drive predictive incident management. 

  • Collaborate with cross-functional teams to integrate AI-driven event correlation models and automated remediation workflows into the central monitoring platform. 

 

  1. Integration, Vendor Coordination 

  • Manage relationships and support escalations with platform vendors. 

  • Work on REST API integrations across applications/tools as per requirements. 

 

  1. Operational Troubleshooting & Diagnostics 

  • Perform deep-dive troubleshooting and root-cause analysis for platform-level performance degradations, engine polling failures, and monitoring agent corruptions. 

  • Utilize Active Diagnostics and system telemetry to investigate and resolve complex network configuration transfer failures, polling sync latency, and data ingestion issues. 

Required Skills: 

  • Multi-tool expertise (SolarWinds, OpenText NNMi, Splunk, etc.) 

  • Protocol & Telemetry Knowledge: In-depth understanding of SNMP (v2c/v3), WMI, WinRM, Syslog, NetFlow/sFlow, and Observability (Metrics, Logs, Traces). 

  • Automation & API Integration: Good to have skills in PowerShell/Python, and API-driven automation for monitoring workflows. 

  • AIOps & Intelligent Automation: Basic understanding of AIOps concepts, machine learning algorithms for anomaly detection, automated event correlation, and predictive analytics within modern observability frameworks. 

  • Cloud & Hybrid Observability: Hands-on experience extending platform monitoring into AWS, Azure, or GCP environments. 

  • Infrastructure Knowledge 

  • System Administration: Intermediate knowledge of Windows and Linux administration. 

  • Database: Understanding of SQL/Database concepts and standard query execution. 

  • Networking: Good understanding of networking concepts including TCP/IP, DNS, DHCP, Routing and Switching. 

  • ITSM: Experience in ITSM processes and operational support. 

About the Company

V

Vantage Point Consulting