Overview:
Observability & Enterprise Monitoring Engineer with specialized expertise in SolarWinds platform administration and broader multi-tool observability ecosystems. Working knowledge of OpenText NNMi will be an added advantage. This role will be responsible for the end-to-end administration, optimization, integration, and operational maintenance of enterprise-scale implementation of monitoring solutions (SolarWinds). Responsible for ensuring platform health, automate alert workflows, manage hybrid/cloud monitoring integrations, and collaborate closely with cross-functional infrastructure teams to maintain high availability and performance.
Roles & Responsibilities:
Platform Administration & Lifecycle Management (SolarWinds)
Core Module Management: Administer and optimize SolarWinds modules including NPM, NCM, NTA, SAM, and the broader Orion / SWOSH (Hybrid Cloud Observability) platform ecosystem.
Upgrades & Maintenance: Perform routine and major version updates across platform components; monitor platform health using Active Diagnostics and My Deployment health checks.
Polling Infrastructure: Manage, scale, and load-balance Additional Polling Engines (APEs) to ensure optimal performance across enterprise environments.
Database & Backup Operations: Perform operational tasks on the underlying MS SQL Database, manage, schedule, and verify configuration and database backup jobs.
Network & Device Observability Operations
Discovery & Asset Management: Execute network discoveries, manage node onboarding/offboarding, assign Universal Device Pollers (UnDP), and maintain custom custom attributes and group hierarchies.
Configuration Management (NCM): Build and maintain NCM command templates, automate daily startup/running config backups, archive config files, and remediate compliance/transfer failures.
Topology & Visualization: Create and maintain dynamic, accurate network topology maps using Network Atlas and modern visual canvases based on operational requirements.
Alerting, Dashboarding & ITSM Integration
Signal Optimization: Design, tune, and maintain custom Alert Triggers, Actions, and Thresholds to eliminate alert noise and drive actionable alerting.
Ticketing & Automation: Configure bi-directional ITSM/ticketing integrations to enable automatic ticket creation, routing, and lifecycle tracking.
Reporting & Visibility: Build custom operational and executive Dashboards, Views, and Reports tailored to stakeholder requirements.
Incident Support: Monitor alert channels for operational anomalies, troubleshoot lingering telemetry issues, and collaborate with domain teams to drive root cause resolution.
AIOps Operations
Leverage AIOps, machine learning, and pattern-recognition capabilities to identify baseline anomalies, reduce event noise, and drive predictive incident management.
Collaborate with cross-functional teams to integrate AI-driven event correlation models and automated remediation workflows into the central monitoring platform.
Integration, Vendor Coordination
Manage relationships and support escalations with platform vendors.
Work on REST API integrations across applications/tools as per requirements.
Operational Troubleshooting & Diagnostics
Perform deep-dive troubleshooting and root-cause analysis for platform-level performance degradations, engine polling failures, and monitoring agent corruptions.
Utilize Active Diagnostics and system telemetry to investigate and resolve complex network configuration transfer failures, polling sync latency, and data ingestion issues.
Required Skills:
Multi-tool expertise (SolarWinds, OpenText NNMi, Splunk, etc.)
Protocol & Telemetry Knowledge: In-depth understanding of SNMP (v2c/v3), WMI, WinRM, Syslog, NetFlow/sFlow, and Observability (Metrics, Logs, Traces).
Automation & API Integration: Good to have skills in PowerShell/Python, and API-driven automation for monitoring workflows.
AIOps & Intelligent Automation: Basic understanding of AIOps concepts, machine learning algorithms for anomaly detection, automated event correlation, and predictive analytics within modern observability frameworks.
Cloud & Hybrid Observability: Hands-on experience extending platform monitoring into AWS, Azure, or GCP environments.
Infrastructure Knowledge
System Administration: Intermediate knowledge of Windows and Linux administration.
Database: Understanding of SQL/Database concepts and standard query execution.
Networking: Good understanding of networking concepts including TCP/IP, DNS, DHCP, Routing and Switching.
ITSM: Experience in ITSM processes and operational support.