AWS Lambda, Amazon Web Services (AWS), Apache Kafka, BGP, Budgeting, Cloud Computing, Computer Security, Data Collection, Fleet Management, GCP (Good Clinical Practices), Instrumentation, Instrumentation Engineering, Java, Metrics, Microsoft Windows Azure, Netflow, Node.js, Performance Analysis, Python Programming/Scripting Language, Receivers, Root Cause Analysis, SNMP (Simple Network Management Protocol), Simple Queue Service (SQS), Software as a Service (SaaS), Standards Development, Technical Writing, Telemetry, Vehicle Fleets, Web Client Plug-ins, World Wide Web Consortium (W3C), Writing Skills
LOCATION
Woodland Hills, CA
POSTED
27 days ago
Senior OpenTelemetry Engineer
Must Have Technical/Functional Skills
This role focuses on the development of a world-class observability platform spanning SaaS products, cloud infrastructure, network fabric, and custom applications. As a Senior OTel Instrumentation Engineer, the successful candidate will own the full signal collection layer. This includes designing and deploying Open Telemetry instrumentation across all four signal domains to ensure traces, metrics, and logs are accurate, well-attributed, and reliably directed to the AIOps Lakehouse.
The position requires close collaboration with Platform, SRE, and ML Engineering teams to define instrumentation standards, evaluate new OTel components, and deliver high-fidelity telemetry that powers anomaly detection, root cause analysis, and business observability.
Roles & Responsibilities
SaaS Instrumentation
End-to-End Instrumentation: Instrument SaaS products utilizing OTel SDKs (Python, Node.js, Java, Go) to ensure spans, metrics, and structured logs are emitted at the appropriate granularity.
Semantic Conventions: Define and enforce attribute naming conventions (e.g., service.name, tenant.id, feature.flag) aligned strictly with OTel semantic standards.
Multi-Tenant Observability: Instrument multi-tenant surfaces to guarantee robust tenant-level observability while preventing cross-tenant data leakage.
Cloud Infrastructure
Collector Fleet Management: Deploy and maintain the OTel Collector fleet across AWS, GCP, and Azure, including receiver configurations, processor pipelines, and exporter routing.
Runtime Instrumentation: Instrument serverless (Lambda, Cloud Run) and container runtimes (EKS, GKE, AKS) utilizing auto-instrumentation where feasible, and manual instrumentation when technical nuance requires it.
Metric Normalization: Collect and normalize cloud provider metrics (CloudWatch, Cloud Monitoring, Azure Monitor) via OTel receiver plugins.
Network Telemetry
Data Collection: Gather network flow data (sFlow, NetFlow/IPFIX), SNMP traps, and BGP state via OTel-native and bridged receivers.
Context Propagation: Correlate network events with application traces using consistent trace-context propagation across network boundaries.
Strategy Collaboration: Partner with NetOps to define the MELT (Metrics, Events, Logs, Traces) strategy for on-prem, SD-WAN, and cloud interconnect segments.
Application Observability
APM Ownership: Own Application Performance Monitoring (APM) instrumentation, including distributed tracing, custom span attributes, database query capture, and error fingerprinting.
SLI/SLO Definition: Define and instrument Service Level Indicators (SLIs) and Service Level Objectives (SLOs)-such as latency histograms, error-rate counters, and availability gauges-queryable directly from the Lakehouse.
Platform & Standards
Internal Libraries: Build and maintain internal instrumentation libraries and language-specific wrappers that encode organizational conventions.
Governance: Author and review instrumentation RFCs while driving the adoption of OTel semantic conventions across all engineering teams.
Data Alignment: Collaborate with the AIOps ML team to ensure telemetry schemas meet feature engineering and ML model requirements.
Pipeline Operations: Operate the Collector pipeline at scale, managing backpressure handling, sampling strategies (tail/head), and cardinality budgets.