Data Engineering Execution- Build and maintain ingestion frameworks (ADF / Databricks / Spark)
- Implement Bronze -> Silver transformations aligned to architecture
- Architect data quality checks, schema validation, and contract rules
- Optimize pipelines for performance and cost
DevOps & Automation- Own CI/CD pipelines (Azure DevOps)
- Automate onboarding (repos, pipelines, policies, templates)
- Maintain IaC and scripting (Terraform, PowerShell / Python)
- Troubleshoot pipeline failures and permission issues
Data Modeling - Design and implement curated data models in Gold layer aligned to business use cases
Build:
- Dimensional models (star/snowflake schemas)
- Fact tables, dimensions, surrogate keys, SCD handling
Translate Silver datasets into:
- Analytics-ready models
- Consistent KPI definitions and business logic
Ensure:
- Consistency across domains (common dimensions, conformed models)
- Reusability and scalability of models
- Alignment with enterprise semantic layer (Power BI / Fabric / downstream tools)
Handle changes such as:
- Column derivation logic changes (not just schema changes)
- Backward compatibility and impact analysis
Partner with business / data stewards to:
- Define business rules
- Validate metrics and edge cases
- Support regression scenarios
Governance & Data Quality Implementation
Implement governance patterns defined by architecture
Enforce:
- Unity Catalog standards
- naming conventions, metadata policies
- access controls (RBAC/ABAC)
Architecture Translation
Automation first thinking and design approach.
Convert high-level architecture into:
- low-level designs
- reusable templates
- engineering playbooks
Ensure consistency across projects and teams
Team Enablement- Support delivery teams with:
- onboarding
- debugging
- design reviews
- Provide hands-on guidance
- Create reusable assets to reduce repeat questions
Required Skills
Core Platform
- Azure Data Platform (ADF, ADLS, Azure SQL)
- Databricks (DLT, Delta Lake, Spark, Workflows, Unity Catalog)
- Microsoft Fabric (preferred)
Engineering
- PySpark / Python (advanced)
- SQL (strong)
- Pipeline design (batch + streaming)
DevOps / Automation
- Azure DevOps (pipelines, repos, policies
- PowerShell automation (must-have for your setup)
- CI/CD + Infrastructure-as-Code concepts
Governance / Data Quality
- Schema enforcement, data contracts, DQ frameworks
- Metadata-driven design understanding