Serve as the primary point of contact for all technical groups and customers to address support issues and requests Conduct technical analysis of incidents, perform root cause investigation and review with Engineering teams Report, monitor, and troubleshoot the pipeline, and implement necessary workarounds Debug and resolve complex issues across distributed and streaming data platforms, including latency spikes, consumer lag, and state/checkpoint failures Monitor and troubleshoot Kafka topics, partitions, consumer groups, and offset management to maintain healthy data flow Investigate and resolve Cassandra issues related to data persistence, replication, and consistency Collaborate with Engineering teams and SREs to develop, implement, and improve Incident Management and Problem Management processes Work with Project Managers and Operations on new projects and developer communications Analyze workflows, file detailed bug reports, and follow through until issues are resolved Write and maintain scripts (Python/Java) to automate diagnostics, monitoring, and remediation tasks Build and maintain observability tooling (dashboards, alerts, tracing) to proactively catch issues before they impact production5+ years of experience in debugging technical issues, with at least 2+ years in distributed systems or streaming data platforms Proficient in SQL and experience with modern streaming and big-data ecosystems (Apache Flink, Kafka, Cassandra, Spark) Excellent scripting knowledge with Python or Java Experience debugging streaming pipelines: latency analysis, lag monitoring, state management, and checkpoint/savepoint recovery in Flink Proficiency with Kafka including topics, partitioning, consumer groups, and offset management Cassandra expertise: data persistence patterns, replication strategies, and troubleshooting consistency issues Proven ability to self-start, learn, plan, prioritize, and deliver to deadlines Excellent interpersonal skills, especially the ability to filter and distill meaningful information to the right audience Self-starter with a strong sense of personal responsibility Understands priorities and doesnt compromise quality Excellent verbal and written communication skills Git repository and version control expertise Minimum BS in Computer Science, Engineering, or related field/equivalent experience. Experience implementing and administering observability tools: Splunk, Grafana/Mosaic, Sentry, and tracing systems Familiarity with HDFS and batch-to-streaming data migration patterns Knowledge of Oracle Golden Gate (OGG) Experience with CI/CD pipelines and deployment strategies across multiple environments.