Design, develop, and maintain scalable Big Data solutions using Hadoop and Spark.
Build and optimize ETL/ELT pipelines using PySpark, Hive, and Python.
Process and analyze large datasets in distributed environments.
Develop high-performance Spark jobs and optimize existing workloads.
Create and manage Hive tables, partitions, views, and complex queries.
Implement data quality, data validation, and reconciliation frameworks.
Perform code reviews and ensure adherence to coding standards and best practices.
Utilize Microsoft Copilot to accelerate development, automate code generation, troubleshooting, documentation, and testing activities.
Strong experience building both batch and real-time streaming applications with Kafka.
Collaborate with Data Architects, Data Scientists, Business Analysts, and DevOps teams.
Troubleshoot production issues and perform root cause analysis.
Design data ingestion frameworks for structured, semi-structured, and unstructured data.
Participate in Agile ceremonies including sprint planning, estimation, and retrospectives.
Mentor junior developers and provide technical leadership.